The two clocks that flew in opposite directions
Assumes: The clock that has to slow, and why no clock can refuse · The clock that runs slow lower down
A moving clock runs slow and a clock lower down runs slow. Both are established, both are measured, and the third rung of this ladder is about a clock that suffers both at once and whose answer depends on which way round the world it went.
The frame to do it in
The calculation is short and the difficulty is entirely in choosing where to stand.
The Earth’s surface is not an inertial frame — it is rotating — so a clock sitting on it is not at rest in any frame in which the ordinary time-dilation formula applies without corrections. The remedy is to work in a frame centred on the Earth and not rotating with it. In that frame no clock is privileged and everything is either moving or high up or both.
A clock on the ground at latitude is then moving, at , which is 402 metres a second at 30°. That is the fact the whole result turns on.
An aircraft at altitude with ground speed , taken positive eastward, is moving at . Its rate relative to the ground clock has two terms:
the first from the gravitational potential difference and the second from the difference of the squared speeds. The cross term is the one that changes sign with the direction of flight, and it is much larger than because the ground speed of the Earth’s surface is larger than the aircraft’s.
The numbers
At nine kilometres, 240 metres a second and 30° latitude, the model gives:
Height: nanoseconds an hour, either direction.
Motion, eastward: nanoseconds an hour.
Motion, westward: nanoseconds an hour — a gain, because subtracting the aircraft’s speed from the ground’s leaves the aircraft moving more slowly than the ground it left.
Summed over the flights Hafele and Keating actually made — 41.2 hours eastward and 48.6 westward — that is nanoseconds and . Their own predictions, computed from the flight logs rather than from one average altitude and speed, were and ; their measurements were and .
Three lines of arithmetic getting both signs and both magnitudes to twenty per cent is the point of the exercise. What the flight logs buy is the last twenty per cent, and what the experiment buys is the knowledge that any of it is real.
The speed at which nothing happens
The two terms cancel at one ground speed, and the figure solves for it: about 180 metres a second eastward.
An aircraft cruising at that speed at nine kilometres keeps the time of the airfield it left, exactly, for as long as it flies. It is a rather pleasing fact and it has a practical shadow: the eastward flights of the 1971 experiment were close enough to that speed for the net to be small and the sign to be uncertain, which is why the eastward prediction had a 60 per cent uncertainty while the westward one had 8.
The cancellation speed depends on the altitude, since the gravitational term does, and on the latitude, since the ground’s own speed does. At the equator the ground moves fastest and the cancellation happens at a lower airspeed; at the pole the cross term vanishes and there is no cancellation at all — a clock flown over the pole gains from height and loses from its own only, and the second is tiny.
That polar case is worth noticing because it removes the confusion the experiment is famous for. Fly a clock north to the pole, round in a small circle, and back, and the answer is the ordinary one: height gains, motion loses, no direction dependence. The asymmetry is entirely about the Earth turning.
The third term nobody mentions
There is a contribution the figure leaves out and it is not small: the Sagnac term.
A signal sent between two points on a rotating Earth takes a different time going east than going west, because the receiver moves toward or away from the sender during the flight. For a signal round the whole equator the difference is about 414 nanoseconds — larger than either relativistic effect in the flying-clock experiment.
It is not a relativistic effect in the sense the others are: it is a first-order kinematic consequence of the ground rotating, visible in the non-rotating frame as nothing more than the target having moved. But it must be corrected in exactly the same place, and any navigation or timekeeping system that transfers time between rotating stations applies it.
The same term appears in an interferometer on a turntable, where two beams sent round a loop in opposite directions come back disagreeing by . It is one of the very few effects that measure rotation absolutely, with no external reference, and its appearance in a clock-transfer calculation and in a laser gyroscope is the same physics.
Keeping it separate from the two terms in this essay is worth doing carefully, because it is a common place to double-count: the Sagnac term concerns the signal used to compare two clocks, and the terms above concern the clocks.
Two beams sent round a rotating loop arrive with a difference proportional to the enclosed area and the rate of turn, and on the scale of the Earth that is a 414-nanosecond effect between antipodal stations — larger than either of the terms the flying clocks measured. It is not a third physical effect so much as a consequence of trying to define one simultaneity over a rotating planet, and it is the reason the experiment’s arithmetic has to be done in a non-rotating frame and translated afterwards.
Where the same sum is corrected continuously
The experiment is now a routine engineering correction, applied every second in every satellite navigation system.
A satellite at 20,200 kilometres is high enough for the gravitational gain to dominate the orbital loss. The gravitational term is microseconds a day and the kinematic one , for a net gain of microseconds a day.
Since the system determines position by comparing arrival times of signals travelling at the speed of light, an error of 38.6 microseconds is a position error of 11.6 kilometres, accumulating daily. The correction is applied by setting the satellites’ clocks to run slow before launch, at a rate chosen so that they keep ground time in orbit.
That is the strongest routine confirmation of both effects there is: not a single measurement with an uncertainty but a system that would fail visibly within minutes if either term were wrong, and which has not.
The low-orbit case goes the other way. A clock on the International Space Station at 400 kilometres loses about 28 microseconds a day, because at that altitude the orbital speed’s contribution beats the height’s. The two regimes are separated by a crossing at about 3,200 kilometres, which is drawn above.
The exact time-dilation factor against distance from a mass is what a satellite system uses, and the linear approximation is what an aeroplane needs — the two differ by parts in at the altitudes involved here, which is below the clocks’ resolution and far above a satellite’s. Near the Earth’s surface the exact expression flattens into , which is the first term of this essay’s sum and is why an altitude difference does the work of a potential difference.
Why the two terms are the same term
It is worth resisting the temptation to treat these as two different effects that happen to be added.
Both are consequences of one statement: a clock measures the proper time along its own worldline, and
to the order that matters here. The gravitational term is the first bracket and the kinematic one the second, and they are two contributions to a single integral rather than two phenomena.
That unification is what makes the experiment a test of general relativity rather than of two separate theories. A clock’s reading depends on the path it took through spacetime, in exactly the way a road’s length depends on the route, and the flying clock and the ground clock simply took different routes between the same two events.
Which is the same argument as the twin problem, with the wrinkle that here neither clock’s route is a straight line — both are accelerating, one round the Earth’s axis and one round the Earth on a different circle — so there is no version of the puzzle in which one of them is the inertial one.
The kinematic effect in its bare form is a clock whose ticking is a light pulse crossing a moving gap: the pulse travels further, the tick is longer, and the factor is . That is one of the two terms in the sum, and on an aircraft it is the smaller of the pair unless the aircraft is flying east — where the ground speed adds to the Earth’s rotation and the velocity term overtakes the altitude term.
In a frame in free fall the gravitational term disappears locally, which is exactly the trouble. Both terms in this essay are statements about a comparison between two extended worldlines — a clock that stayed and a clock that went — and a comparison of that kind cannot be made inside any one falling frame. The equivalence principle removes gravity at a point and says nothing about a journey, which is why the experiment needs the full metric rather than a local argument.
The objection, and its answer
The experiment attracts a standard objection and it is worth dealing with, because the answer says something about what the theory is.
The objection runs: relativity says motion is relative, so from the aircraft’s point of view the ground was moving and the ground clock should have run slow. Both cannot be right; therefore something is wrong.
The answer is that neither clock is in an inertial frame, so neither is entitled to the naive argument. The ground clock is going round the Earth’s axis once a day and the aircraft is going round on a different circle at a different rate, and both are accelerating. What each of them measures is the proper time along its own worldline, and the two worldlines are different curves between the same pair of events — departure and return.
There is no symmetry to appeal to, and looking for one is the mistake. The comparison is made in a frame in which neither clock is at rest, precisely so that no clock’s own perspective has to be privileged, and the answer that comes out is the same whatever frame the calculation is done in as long as it is done properly.
That is worth generalising. The question “whose time dilation?” has no answer; the question “what is the proper time along each worldline?” always does. Every apparent paradox in this part of the subject dissolves on being restated in the second form, and the twin problem’s turnaround is the standard case.
What it cost to do
The experiment is famous partly for having been cheap. Hafele and Keating flew four caesium beam clocks as ordinary passengers on scheduled commercial flights, buying seats for the clocks, at a total cost of about eight thousand dollars — reportedly the least expensive test of general relativity ever performed.
The clocks were not especially good by later standards, and the largest source of error was not any of the relativity but the clocks’ own drift: each had a rate that wandered by a few nanoseconds a day, and the analysis had to model that drift and subtract it. Using four clocks and taking their ensemble was what made the result usable at all, and the treatment of the drift has been argued about since.
The modern versions are not cheap. Gravity Probe A in 1976 flew a hydrogen maser to ten thousand kilometres on a rocket and confirmed the gravitational term to 70 parts per million. Optical clocks now resolve the gravitational shift over a height difference of a centimetre, which means the effect is a routine nuisance in any laboratory comparing two clocks on the same bench.
That last development is worth pausing on. A clock good enough to see the shift across a table has to be told its own altitude to a centimetre before it can be compared with another, so relativistic geodesy — measuring height by comparing clock rates — has become a real technique rather than a thought experiment.
The reason it is a better test than either half
An experiment measuring one effect confirms one number. This one measures a sum of two with opposite signs, and that is a categorically stronger constraint.
Consider a theory that gets the gravitational shift right and the kinematic one wrong by a factor of 1.5. On the westward flight the two terms are 171 and 132 nanoseconds, both positive, and the error is 66 nanoseconds on a total of 303 — a 22 per cent discrepancy, within reach of the error bars of the day. On the eastward flight the terms are 145 and −207, and the same error makes the prediction −169 instead of −61: a factor of nearly three, and unmistakable.
That is the general virtue of measuring a difference of two large quantities rather than one small one. Where the two terms nearly cancel, the fractional sensitivity to each of them is amplified by the ratio of the individual terms to their difference — here about a factor of three eastward, and much more at the cancellation speed.
The same design principle turns up everywhere in this collection. A null result is limited only by the detector’s sensitivity and needs no calibration; a near-cancellation is the next best thing, and it is why the eastward flight, despite being the one with the larger uncertainty, is the one that carries the argument.
The same sum, measured on a bench
The experiment’s descendants have shrunk it to laboratory scale, and the smallness of the apparatus is the point rather than a convenience.
In 2010 two optical clocks based on single aluminium ions were compared over a fibre link, and one of them was raised by thirty-three centimetres. The predicted fractional shift is , which is four parts in — and it was resolved, so the gravitational term of this essay was measured across the height of a laboratory bench. The same pair measured the kinematic term by driving one ion into motion at a few metres a second, which is a shift of a few parts in .
Both terms of the flying-clock sum, in one room, at speeds a person can walk and heights a person can reach. A more recent version resolves the gravitational shift across a millimetre — the height of a cloud of cold atoms, whose top and bottom keep measurably different time.
The consequence is that altitude has stopped being a footnote for anyone comparing clocks. The second is defined on the geoid, so every national laboratory must correct its own clock for its own height above sea level before contributing to international time, and an error of a metre is a fractional error of — larger than the clocks’ own instability. Turned round, that is a surveying instrument: compare two clocks and the difference reports the gravitational potential between them, which is what a height means. A network of optical clocks measures the shape of the Earth’s potential directly, without levelling and without gravimeters.
The same sum, at a neutron star
The other direction the experiment has been pushed is to a place where the terms are not parts in but parts in a thousand.
A pulsar in a close binary orbit is a clock in an eccentric path around a companion of comparable mass. Its speed varies round the orbit and so does its distance from the companion, so both terms of this essay’s sum vary together — and the resulting modulation of the arrival times of its pulses is measurable and is a standard parameter in pulsar timing, called the Einstein delay. For the first such system found, it is about four milliseconds peak to peak, on a pulse train whose individual arrival times are known to microseconds.
Its value matters because it depends on the two masses in a different combination from the other measurable orbital parameters. Measure enough of them — the delay, the precession of the orbit, the shape and delay of the signal passing the companion, the orbital decay — and the system is over-determined: two unknown masses constrained by four or more measurements, so the theory is being tested rather than fitted. The best such system now agrees with general relativity to better than a part in ten thousand.
What is worth carrying away is that this is the same calculation. Two clocks on different worldlines between the same events, a gravitational term and a kinematic one, summed into one proper time. Hafele and Keating flew theirs as airline passengers and got twenty per cent; the same integral evaluated for a neutron star orbiting another at a fraction of the speed of light is currently the most precise test the theory has.
What the picture cannot show
The model uses one altitude, one speed and one latitude. A real flight climbs, cruises, descends and changes latitude continuously, and the published prediction integrated the flight logs. The figure’s agreement to twenty per cent is what a three-parameter model deserves, and reading it as a confirmation of the theory to twenty per cent would be a misreading — the theory does much better than that when given the actual route.
The Earth is treated as spherical and non-rotating in its gravity. The real potential includes the rotational flattening, and the surfaces of constant potential are what a clock on the ground actually follows. To the precision here it does not matter; to the precision of an optical clock it dominates.
Only the first order in is kept. That is ample — the terms are parts in — and the second order would be parts in .
And the experiment’s own uncertainty is larger than the figure suggests. The measurements carry 10 and 7 nanoseconds of error on results of 59 and 273, and the eastward one in particular has been reanalysed several times with somewhat different answers. The confidence in the physics rests on the satellite systems rather than on the aeroplanes.
The ladder from here
Later rungs on this anchor: the proper-time integral written properly, from which both terms fall out as one; relativistic geodesy, where the clock comparison is the measurement and the height is the unknown; the Sagnac term, which is a third contribution appearing whenever a signal is sent round a rotating Earth and which navigation systems correct for separately; optical clock comparisons at the level and what they can test; and the light-time and relativistic corrections in deep-space navigation, where the same integral has to be evaluated along a spacecraft’s whole trajectory.
The neighbouring ladders are the moving clock and the clock lower down, which are the two terms in isolation, and the clock that is wrong in two directions, where the same care about which frame the comparison is made in decides the answer.
Part 4 of 6
This essay is one argument about Time dilation. The others:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
Atomic clockEquivalence principleGravitational redshiftHafele keatingProper timeReference frameSatellite navigationTime dilation
- The parallelogram that will not close equivalence principle, gravitational redshift, proper time
- The wall of silence behind a rocket that never stops equivalence principle, proper time, reference frame
- Everything from an exchange of pulses proper time, time dilation
- How big now is equivalence principle, reference frame
- The clock that does not feel the turn proper time, time dilation
- The disc that cannot be spun equivalence principle, time dilation