Relativity

The two clocks that flew in opposite directions

Two caesium clocks were flown round the world in 1971, one each way, and came back disagreeing with the clock left behind — one having lost 59 nanoseconds and the other gained 273. Height alone would have made both gain. The sign flip comes from the ground already moving eastward at 400 metres a second before the aircraft took off.

Assumes: The clock that has to slow, and why no clock can refuse · The clock that runs slow lower down

A moving clock runs slow and a clock lower down runs slow. Both are established, both are measured, and the third rung of this ladder is about a clock that suffers both at once and whose answer depends on which way round the world it went.

The clock that gains going one way and loses going the other. The rate at which a flown clock gains on a clock left at 30° latitude, in nanoseconds per hour, against the aeroplane's ground speed, with east taken as positive. Two terms are drawn and then their sum. Height alone gives 3.5 nanoseconds an hour at 9 km and does not care which way the aircraft is pointed. Motion costs time, and because the ground is already moving eastward at 402 metres a second, flying east adds to that speed and flying west subtracts from it — so the kinematic term is much larger going east and can change sign going west. The sum crosses zero at 180 metres a second eastward, which is the ground speed at which an aeroplane's clock keeps the time of the airfield it left. Over the two flights Hafele and Keating actually made, this simple model gives -61 nanoseconds eastward and +304 westward, against their own predictions of -40 and +275 and their measurements of -59 and +273. The model here uses one average altitude, one average speed and one latitude, where the real prediction integrated the flight logs; getting the signs and the rough sizes out of three lines of arithmetic is the point, and the last twenty per cent is what the logs are for. What no amount of arithmetic supplies is the thing the experiment settled: that the effect is real, that it acts on a caesium clock in a passenger seat, and that a difference of a few hundred nanoseconds after two days is measurable.
Fig. 1 The rate at which a flown clock gains on one left behind, in nanoseconds per hour, against the aeroplane’s ground speed with east taken as positive. Height alone is the flat line and does not care which way the aircraft points. Motion is the curve, and it is asymmetric because the ground is already moving.

The frame to do it in

The calculation is short and the difficulty is entirely in choosing where to stand.

The Earth’s surface is not an inertial frame — it is rotating — so a clock sitting on it is not at rest in any frame in which the ordinary time-dilation formula applies without corrections. The remedy is to work in a frame centred on the Earth and not rotating with it. In that frame no clock is privileged and everything is either moving or high up or both.

A clock on the ground at latitude λ\lambda is then moving, at RΩcosλR\Omega\cos\lambda, which is 402 metres a second at 30°. That is the fact the whole result turns on.

An aircraft at altitude hh with ground speed vv, taken positive eastward, is moving at (R+h)Ωcosλ+v(R+h)\Omega\cos\lambda + v. Its rate relative to the ground clock has two terms:

Δττ=ghc22RΩcosλv+v22c2\frac{\Delta\tau}{\tau} = \frac{gh}{c^2} - \frac{2R\Omega\cos\lambda\, v + v^2}{2c^2}

the first from the gravitational potential difference and the second from the difference of the squared speeds. The cross term 2RΩcosλv2R\Omega\cos\lambda\,v is the one that changes sign with the direction of flight, and it is much larger than v2v^2 because the ground speed of the Earth’s surface is larger than the aircraft’s.

The numbers

At nine kilometres, 240 metres a second and 30° latitude, the model gives:

Height: +3.5+3.5 nanoseconds an hour, either direction.

Motion, eastward: 5.0-5.0 nanoseconds an hour.

Motion, westward: +2.7+2.7 nanoseconds an hour — a gain, because subtracting the aircraft’s speed from the ground’s leaves the aircraft moving more slowly than the ground it left.

Summed over the flights Hafele and Keating actually made — 41.2 hours eastward and 48.6 westward — that is 61-61 nanoseconds and +304+304. Their own predictions, computed from the flight logs rather than from one average altitude and speed, were 40±23-40 \pm 23 and +275±21+275 \pm 21; their measurements were 59±10-59 \pm 10 and +273±7+273 \pm 7.

Three lines of arithmetic getting both signs and both magnitudes to twenty per cent is the point of the exercise. What the flight logs buy is the last twenty per cent, and what the experiment buys is the knowledge that any of it is real.

Two corrections, opposite in sign and different in size. How fast a clock in a circular orbit runs compared with one on the ground, in microseconds a day, against the height of the orbit — with the two effects drawn apart rather than added. Being high speeds a clock up, by an amount that saturates: the potential term is bounded because there is only so much potential to climb out of. Moving slows it down, and a higher orbit is a slower one, so that term shrinks toward zero. They cancel at 3186 km — a radius of exactly 1.5 Earth radii, which follows from setting the sum to zero and contains neither G, nor the Earth's mass, nor the speed of light. At 20200 km the gravitational term is 45.7 µs a day and the speed term −7.2, leaving 38.5. Left uncorrected, that is 11.5 km of position error a day, growing without limit, from a clock that is working perfectly.
Fig. 2 The two terms drawn apart rather than summed, against altitude, with their opposite signs and different powers of the radius. The crossing is where a satellite’s clock stops losing and starts gaining, and it is the same crossing the aircraft figure shows against speed instead of height.

The speed at which nothing happens

The two terms cancel at one ground speed, and the figure solves for it: about 180 metres a second eastward.

An aircraft cruising at that speed at nine kilometres keeps the time of the airfield it left, exactly, for as long as it flies. It is a rather pleasing fact and it has a practical shadow: the eastward flights of the 1971 experiment were close enough to that speed for the net to be small and the sign to be uncertain, which is why the eastward prediction had a 60 per cent uncertainty while the westward one had 8.

The cancellation speed depends on the altitude, since the gravitational term does, and on the latitude, since the ground’s own speed does. At the equator the ground moves fastest and the cancellation happens at a lower airspeed; at the pole the cross term vanishes and there is no cancellation at all — a clock flown over the pole gains from height and loses from its own v2v^2 only, and the second is tiny.

That polar case is worth noticing because it removes the confusion the experiment is famous for. Fly a clock north to the pole, round in a small circle, and back, and the answer is the ordinary one: height gains, motion loses, no direction dependence. The asymmetry is entirely about the Earth turning.

The third term nobody mentions

There is a contribution the figure leaves out and it is not small: the Sagnac term.

A signal sent between two points on a rotating Earth takes a different time going east than going west, because the receiver moves toward or away from the sender during the flight. For a signal round the whole equator the difference is about 414 nanoseconds — larger than either relativistic effect in the flying-clock experiment.

It is not a relativistic effect in the sense the others are: it is a first-order kinematic consequence of the ground rotating, visible in the non-rotating frame as nothing more than the target having moved. But it must be corrected in exactly the same place, and any navigation or timekeeping system that transfers time between rotating stations applies it.

The same term appears in an interferometer on a turntable, where two beams sent round a loop in opposite directions come back disagreeing by 4AΩ/c24A\Omega/c^2. It is one of the very few effects that measure rotation absolutely, with no external reference, and its appearance in a clock-transfer calculation and in a laser gyroscope is the same physics.

Keeping it separate from the two terms in this essay is worth doing carefully, because it is a common place to double-count: the Sagnac term concerns the signal used to compare two clocks, and the terms above concern the clocks.

Two beams sent round a rotating loop arrive with a difference proportional to the enclosed area and the rate of turn, and on the scale of the Earth that is a 414-nanosecond effect between antipodal stations — larger than either of the terms the flying clocks measured. It is not a third physical effect so much as a consequence of trying to define one simultaneity over a rotating planet, and it is the reason the experiment’s arithmetic has to be done in a non-rotating frame and translated afterwards.

Where the same sum is corrected continuously

The experiment is now a routine engineering correction, applied every second in every satellite navigation system.

Where a clock gains, and where it loses. The rate of a clock in a circular orbit against one on the ground, in microseconds per day, plotted against altitude. Height makes it gain and speed makes it lose, and the two cancel exactly at 3186 km — where a satellite keeps the same time as the ground for two reasons that have nothing to do with each other. At 20200 km the total is 38.5 µs a day, which is about ten kilometres of position error if it is ignored.
Fig. 3 The net rate of a clock in orbit against altitude, with the two terms summed. Low orbit loses; high orbit gains; the crossing is at about three thousand kilometres. A navigation satellite at 20,200 kilometres gains 38 microseconds a day, which is 11 kilometres of position error if uncorrected.

A satellite at 20,200 kilometres is high enough for the gravitational gain to dominate the orbital loss. The gravitational term is +45.9+45.9 microseconds a day and the kinematic one 7.2-7.2, for a net gain of +38.6+38.6 microseconds a day.

Since the system determines position by comparing arrival times of signals travelling at the speed of light, an error of 38.6 microseconds is a position error of 11.6 kilometres, accumulating daily. The correction is applied by setting the satellites’ clocks to run slow before launch, at a rate chosen so that they keep ground time in orbit.

That is the strongest routine confirmation of both effects there is: not a single measurement with an uncertainty but a system that would fail visibly within minutes if either term were wrong, and which has not.

The low-orbit case goes the other way. A clock on the International Space Station at 400 kilometres loses about 28 microseconds a day, because at that altitude the orbital speed’s contribution beats the height’s. The two regimes are separated by a crossing at about 3,200 kilometres, which is drawn above.

The exact time-dilation factor against distance from a mass is what a satellite system uses, and the linear approximation is what an aeroplane needs — the two differ by parts in 101010^{10} at the altitudes involved here, which is below the clocks’ resolution and far above a satellite’s. Near the Earth’s surface the exact expression flattens into 1+gh/c21 + gh/c^2, which is the first term of this essay’s sum and is why an altitude difference does the work of a potential difference.

Why the two terms are the same term

It is worth resisting the temptation to treat these as two different effects that happen to be added.

Both are consequences of one statement: a clock measures the proper time along its own worldline, and

dτ2=(1+2Φc2)dt2dx2c2d\tau^2 = \left(1 + \frac{2\Phi}{c^2}\right)dt^2 - \frac{dx^2}{c^2}

to the order that matters here. The gravitational term is the first bracket and the kinematic one the second, and they are two contributions to a single integral rather than two phenomena.

That unification is what makes the experiment a test of general relativity rather than of two separate theories. A clock’s reading depends on the path it took through spacetime, in exactly the way a road’s length depends on the route, and the flying clock and the ground clock simply took different routes between the same two events.

Which is the same argument as the twin problem, with the wrinkle that here neither clock’s route is a straight line — both are accelerating, one round the Earth’s axis and one round the Earth on a different circle — so there is no version of the puzzle in which one of them is the inertial one.

The kinematic effect in its bare form is a clock whose ticking is a light pulse crossing a moving gap: the pulse travels further, the tick is longer, and the factor is γ\gamma. That is one of the two terms in the sum, and on an aircraft it is the smaller of the pair unless the aircraft is flying east — where the ground speed adds to the Earth’s rotation and the velocity term overtakes the altitude term.

In a frame in free fall the gravitational term disappears locally, which is exactly the trouble. Both terms in this essay are statements about a comparison between two extended worldlines — a clock that stayed and a clock that went — and a comparison of that kind cannot be made inside any one falling frame. The equivalence principle removes gravity at a point and says nothing about a journey, which is why the experiment needs the full metric rather than a local argument.

The objection, and its answer

The experiment attracts a standard objection and it is worth dealing with, because the answer says something about what the theory is.

The objection runs: relativity says motion is relative, so from the aircraft’s point of view the ground was moving and the ground clock should have run slow. Both cannot be right; therefore something is wrong.

The answer is that neither clock is in an inertial frame, so neither is entitled to the naive argument. The ground clock is going round the Earth’s axis once a day and the aircraft is going round on a different circle at a different rate, and both are accelerating. What each of them measures is the proper time along its own worldline, and the two worldlines are different curves between the same pair of events — departure and return.

There is no symmetry to appeal to, and looking for one is the mistake. The comparison is made in a frame in which neither clock is at rest, precisely so that no clock’s own perspective has to be privileged, and the answer that comes out is the same whatever frame the calculation is done in as long as it is done properly.

That is worth generalising. The question “whose time dilation?” has no answer; the question “what is the proper time along each worldline?” always does. Every apparent paradox in this part of the subject dissolves on being restated in the second form, and the twin problem’s turnaround is the standard case.

What it cost to do

The experiment is famous partly for having been cheap. Hafele and Keating flew four caesium beam clocks as ordinary passengers on scheduled commercial flights, buying seats for the clocks, at a total cost of about eight thousand dollars — reportedly the least expensive test of general relativity ever performed.

The clocks were not especially good by later standards, and the largest source of error was not any of the relativity but the clocks’ own drift: each had a rate that wandered by a few nanoseconds a day, and the analysis had to model that drift and subtract it. Using four clocks and taking their ensemble was what made the result usable at all, and the treatment of the drift has been argued about since.

The modern versions are not cheap. Gravity Probe A in 1976 flew a hydrogen maser to ten thousand kilometres on a rocket and confirmed the gravitational term to 70 parts per million. Optical clocks now resolve the gravitational shift over a height difference of a centimetre, which means the effect is a routine nuisance in any laboratory comparing two clocks on the same bench.

That last development is worth pausing on. A clock good enough to see the shift across a table has to be told its own altitude to a centimetre before it can be compared with another, so relativistic geodesy — measuring height by comparing clock rates — has become a real technique rather than a thought experiment.

A photon climbing a tower. A photon emitted at the foot of a tower 22.5 m high and received at the top. It arrives with its frequency lower by gh/c² = 2.455·10⁻¹⁵ — two and a half parts in a thousand million million. Nothing was done to the photon on the way up; the two ends of the tower disagree about how fast time passes, and the frequency is the evidence. The same fraction says a clock at the foot loses 0.21 nanoseconds a day against one at the top.
Fig. 4 Pound and Rebka’s shaft, twenty-two and a half metres of it, where the gravitational shift was first measured on Earth. The tower’s effect is two parts in 101510^{15}; an optical clock now resolves that over a centimetre, which is why altitude is a calibration rather than a curiosity.

The reason it is a better test than either half

An experiment measuring one effect confirms one number. This one measures a sum of two with opposite signs, and that is a categorically stronger constraint.

Consider a theory that gets the gravitational shift right and the kinematic one wrong by a factor of 1.5. On the westward flight the two terms are 171 and 132 nanoseconds, both positive, and the error is 66 nanoseconds on a total of 303 — a 22 per cent discrepancy, within reach of the error bars of the day. On the eastward flight the terms are 145 and −207, and the same error makes the prediction −169 instead of −61: a factor of nearly three, and unmistakable.

That is the general virtue of measuring a difference of two large quantities rather than one small one. Where the two terms nearly cancel, the fractional sensitivity to each of them is amplified by the ratio of the individual terms to their difference — here about a factor of three eastward, and much more at the cancellation speed.

The same design principle turns up everywhere in this collection. A null result is limited only by the detector’s sensitivity and needs no calibration; a near-cancellation is the next best thing, and it is why the eastward flight, despite being the one with the larger uncertainty, is the one that carries the argument.

The same sum, measured on a bench

The experiment’s descendants have shrunk it to laboratory scale, and the smallness of the apparatus is the point rather than a convenience.

In 2010 two optical clocks based on single aluminium ions were compared over a fibre link, and one of them was raised by thirty-three centimetres. The predicted fractional shift is gh/c2gh/c^2, which is four parts in 101710^{17} — and it was resolved, so the gravitational term of this essay was measured across the height of a laboratory bench. The same pair measured the kinematic term by driving one ion into motion at a few metres a second, which is a shift of a few parts in 101610^{16}.

Both terms of the flying-clock sum, in one room, at speeds a person can walk and heights a person can reach. A more recent version resolves the gravitational shift across a millimetre — the height of a cloud of cold atoms, whose top and bottom keep measurably different time.

The consequence is that altitude has stopped being a footnote for anyone comparing clocks. The second is defined on the geoid, so every national laboratory must correct its own clock for its own height above sea level before contributing to international time, and an error of a metre is a fractional error of 101610^{-16} — larger than the clocks’ own instability. Turned round, that is a surveying instrument: compare two clocks and the difference reports the gravitational potential between them, which is what a height means. A network of optical clocks measures the shape of the Earth’s potential directly, without levelling and without gravimeters.

The same sum, at a neutron star

The other direction the experiment has been pushed is to a place where the terms are not parts in 101210^{12} but parts in a thousand.

A pulsar in a close binary orbit is a clock in an eccentric path around a companion of comparable mass. Its speed varies round the orbit and so does its distance from the companion, so both terms of this essay’s sum vary together — and the resulting modulation of the arrival times of its pulses is measurable and is a standard parameter in pulsar timing, called the Einstein delay. For the first such system found, it is about four milliseconds peak to peak, on a pulse train whose individual arrival times are known to microseconds.

Its value matters because it depends on the two masses in a different combination from the other measurable orbital parameters. Measure enough of them — the delay, the precession of the orbit, the shape and delay of the signal passing the companion, the orbital decay — and the system is over-determined: two unknown masses constrained by four or more measurements, so the theory is being tested rather than fitted. The best such system now agrees with general relativity to better than a part in ten thousand.

What is worth carrying away is that this is the same calculation. Two clocks on different worldlines between the same events, a gravitational term and a kinematic one, summed into one proper time. Hafele and Keating flew theirs as airline passengers and got twenty per cent; the same integral evaluated for a neutron star orbiting another at a fraction of the speed of light is currently the most precise test the theory has.

What the picture cannot show

The model uses one altitude, one speed and one latitude. A real flight climbs, cruises, descends and changes latitude continuously, and the published prediction integrated the flight logs. The figure’s agreement to twenty per cent is what a three-parameter model deserves, and reading it as a confirmation of the theory to twenty per cent would be a misreading — the theory does much better than that when given the actual route.

The Earth is treated as spherical and non-rotating in its gravity. The real potential includes the rotational flattening, and the surfaces of constant potential are what a clock on the ground actually follows. To the precision here it does not matter; to the precision of an optical clock it dominates.

Only the first order in 1/c21/c^2 is kept. That is ample — the terms are parts in 101210^{12} — and the second order would be parts in 102410^{24}.

And the experiment’s own uncertainty is larger than the figure suggests. The measurements carry 10 and 7 nanoseconds of error on results of 59 and 273, and the eastward one in particular has been reanalysed several times with somewhat different answers. The confidence in the physics rests on the satellite systems rather than on the aeroplanes.

The ladder from here

Later rungs on this anchor: the proper-time integral written properly, from which both terms fall out as one; relativistic geodesy, where the clock comparison is the measurement and the height is the unknown; the Sagnac term, which is a third contribution appearing whenever a signal is sent round a rotating Earth and which navigation systems correct for separately; optical clock comparisons at the 101810^{-18} level and what they can test; and the light-time and relativistic corrections in deep-space navigation, where the same integral has to be evaluated along a spacecraft’s whole trajectory.

The neighbouring ladders are the moving clock and the clock lower down, which are the two terms in isolation, and the clock that is wrong in two directions, where the same care about which frame the comparison is made in decides the answer.

Part 4 of 6

This essay is one argument about Time dilation. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Atomic clockEquivalence principleGravitational redshiftHafele keatingProper timeReference frameSatellite navigationTime dilation