Astrophysics

The clocks that must all slow together

Every clock on the Earth runs slower in January than in July, by three parts in ten thousand million, because the orbit carries the planet deeper into the Sun's potential at perihelion. No clock on the Earth can see this, and that invisibility is the claim worth testing. If the redshift is a property of time rather than of clocks, two clocks built on different physics must slow by exactly the same fraction, and their ratio must not move with the seasons. A ratio that did move would mean the constants of nature depend on where they are measured.

Assumes: The clock that measures a height · The floor that cannot be told from gravity

A clock lower down runs slow, and the fraction is the difference in gravitational potential divided by the square of the speed of light. That result has been confirmed many ways, and each confirmation measures the size of the shift with one kind of clock: the Mössbauer line of iron-57 in a tower, a hydrogen maser on a rocket, an aluminium ion lifted by a third of a metre. Each experiment checks that one clock slows by the predicted amount.

There is a second claim hidden in the first, and it is the stronger one. The argument that turned the redshift into geometry — the parallelogram of clock comparisons that fails to close — needs every clock to agree. If a caesium clock slowed by one fraction and a hydrogen maser by a slightly different one, there would be no single “rate of time” at a height, only a collection of rates of different mechanisms, and the redshift would be a fact about atoms rather than about spacetime. The claim that the fraction is the same for every clock whatever it is made of is called local position invariance. It says the outcome of any local experiment that does not involve gravity is independent of where in the gravitational field it is done. It is one of the three parts of the equivalence principle, and it can be tested in a way that needs no second location, no rocket and no tower.

A potential the Earth provides for nothing

The Earth’s orbit is not quite circular. Its eccentricity — the length of the vector that points at perihelion — is 0.0167, which moves the planet about five million kilometres closer to the Sun in early January than in early July. The depth of the Sun’s potential at the Earth is GM/rc2GM_\odot/rc^2, a little under 10810^{-8} on average, and it varies as the distance does.

The potential the Earth carries its clocks through. The depth of the Sun's gravitational potential at the Earth, as a fraction of c², measured from its yearly mean of 9.873·10⁻⁹, through one year. The orbit's eccentricity of 0.0167086 takes the Earth closest to the Sun on day 3 and furthest on day 186, and the depth swings by 3.3·10⁻¹⁰ between them, found by solving Kepler's equation and checked against 2e/(1 − e²) times the mean. Every clock on the Earth runs slower in January than in July by that fraction against a clock far from the Sun — about half the 6.96·10⁻¹⁰ by which the Earth's own gravity slows a clock on its surface — and no clock on the Earth can see it by comparison with another clock on the Earth.
Fig. 1 The depth of the Sun’s potential at the Earth through a year, from its mean, found by solving Kepler’s equation. It swings by 3.3 × 10⁻¹⁰ between perihelion on day 3 and aphelion on day 186.

The swing is 2e/(1e2)2e/(1-e^2) times the mean, 3.30×10103.30 \times 10^{-10}. That is not small by the standards of the clocks involved. It is about half the 6.96×10106.96 \times 10^{-10} by which the Earth’s own gravity slows a clock at its surface, and it is hundreds of millions of times larger than the uncertainty of the best clocks now operating. Measured against a clock far out in the solar system, every clock on the Earth runs faster in July than in January by that fraction, and the effect is a routine correction in timekeeping that refers terrestrial clocks to the solar system’s centre of mass.

Yet no measurement on the Earth can see it directly. Two clocks side by side both slow by the same fraction, so their comparison shows nothing, and a clock compared with another clock anywhere else on the Earth is in the same position. The Earth is falling freely round the Sun, and in a freely falling laboratory the potential of the body it falls towards is not locally detectable — only its gradient across the laboratory is, as a tide. The annual swing is real, and the only thing the equivalence principle allows it to affect is the comparison with a clock that is not falling with the Earth.

That is the test. The principle says that every clock on the Earth slows by exactly the swing. If some kind of clock responded a little more strongly or a little less, then two different kinds of clock kept side by side would drift apart and back again once a year, in step with the orbit.

Two clocks, one ratio

Write the response of a clock of type A to a change in potential as (1+βA)ΔU/c2(1 + \beta_A)\,\Delta U/c^2, with βA=0\beta_A = 0 for a clock that obeys the principle, and similarly for type B. The frequency ratio of the two then changes by (βAβB)ΔU/c2(\beta_A - \beta_B)\,\Delta U/c^2. Everything common to the two clocks cancels: the shared slowing, the unknown potential of the laboratory, the Earth’s rotation, the variation in the Earth’s own field. What is left depends only on the difference in how the two mechanisms respond.

Both clocks move, and their ratio does not. Two years of a clock's fractional frequency against a distant clock (upper panel), and of the ratio of two unlike clocks kept side by side (lower panel). Above, both clocks slow as the Earth nears the Sun, with an amplitude of 1.65·10⁻¹⁰, and if the redshift is universal the two curves are one curve. Below, the ratio: the flat line is what universality predicts, and the sinusoid is what a clock responding to the potential 10⁻⁶ more strongly than the other would produce — an annual term of 1.65·10⁻¹⁶, a million times smaller than the shift both clocks share and within reach of clocks that compare to parts in 10¹⁷.
Fig. 2 Above, two clocks’ rates against a distant clock over two years; if the redshift is universal the curves coincide. Below, their ratio in units a million times finer: flat for a universal redshift, and an annual term of 1.65 × 10⁻¹⁶ if one clock responded 10⁻⁶ more strongly.

The picture shows the scale of the problem and why it is tractable. The shared signal is ±1.65×1010\pm 1.65 \times 10^{-10}, and it cancels exactly in the ratio under the principle. A violation at Δβ=106\Delta\beta = 10^{-6} produces a ratio signal of ±1.65×1016\pm 1.65 \times 10^{-16}, a million times smaller, with the same annual period and the same phase as the orbit. That phase is known in advance to within hours, which matters a great deal: the analysis looks for a sinusoid of known period and known phase and fits only its amplitude, and that is a far easier measurement than searching for an unknown signal.

The design is a null test, and it has the virtue all null tests share. A measurement of the redshift’s size compares a predicted number with a measured one, and every systematic error in the measurement is an error in the comparison. A null test compares zero with a measured number, and many systematic errors either cancel or show up with the wrong period. Temperature varies annually, of course, and so does humidity, and a laboratory’s seasonal cycle is the main thing such an analysis has to rule out — but it has to be ruled out only at the phase of the orbit, and perihelion falls within a couple of weeks of the northern winter solstice, which is exactly why the careful versions of the experiment are done at more than one site and in both hemispheres.

The test of size that the orbit also offers

The absolute size of the shift can be tested the same way, with a clock that carries its own changing potential.

In August 2014 two satellites of the European Galileo navigation system were placed in the wrong orbits by a fault in their launcher’s upper stage. Instead of circular orbits they went into elliptical ones, with eccentricity around 0.16, and although the orbits were later raised they stayed noticeably eccentric. For navigation that was a loss. For relativity it was an opportunity, because each satellite carries a passive hydrogen maser and a rubidium clock and is tracked continuously, and an eccentric orbit swings a clock between deep and shallow potential, and between fast and slow, twice a day.

The orbit that tests the shift by itself. The periodic time offset a clock in an eccentric orbit accumulates against the average rate, from the depth of the potential and the speed together, over two orbits. For an eccentric navigation satellite with semi-major axis 27978 km and eccentricity 0.162, the period is 12.94 h and the offset swings by ±381 ns; for an ordinary navigation satellite with semi-major axis 26560 km and eccentricity 0.01, the period is 11.97 h and the offset swings by ±23 ns. The curve is integrated from the rate and checked against −(2/c²)√(GMa)·e·sin E. The offset is 17 times larger for the first orbit, and a clock on it tests the size of the redshift against its own orbit without a second clock sent anywhere.
Fig. 3 The periodic time offset of a clock on an eccentric orbit, from depth and speed together, over two orbits. At eccentricity 0.162 it swings by ±381 nanoseconds; at 0.01, typical of a working navigation satellite, by ±23.

Near perigee the satellite is deeper in the potential and moving faster, and both effects slow its clock, so the offset accumulates; near apogee both reverse. The two do not cancel over the orbit the way they do in comparing orbits of different heights. The periodic part integrates to (2/c2)GMa  esinE-(2/c^2)\sqrt{GMa}\;e\sin E, where EE is the eccentric anomaly, and the same Kepler-orbit averaging that makes every orbit of one period lose the same total time is what leaves this periodic term behind as the only thing that depends on the shape. For the eccentric Galileo orbit it is ±381 nanoseconds. A navigation clock is compared with ground clocks at the level of a fraction of a nanosecond, so the signal is enormous.

Two independent analyses of about three years of data, published in 2018, confirmed the predicted amplitude to a few parts in 10510^5, improving on Gravity Probe A’s 1976 result by a factor of a few. It is worth being clear about what that measures. It tests the size of the shift for hydrogen masers — that β\beta for that kind of clock is zero to a few parts in 10510^5. It says nothing about whether a caesium clock would have shown the same, and that is the question the null test answers much more precisely.

What a record of nothing is worth

A null result is a number with an uncertainty, and the useful thing to understand is how the uncertainty is set.

Suppose two unlike clocks are compared once a week for three years, each comparison with a random error of 5×10165 \times 10^{-16}. Real clocks also drift, slowly and steadily, as they age, so the fit has three terms: a constant, a linear drift, and the known annual shape multiplied by an unknown Δβ\Delta\beta. The figure below is a synthetic record made that way, with no violation in it, to show what the analysis sees.

A synthetic record with nothing in it. A synthetic record, not a measurement: 156 comparisons of two unlike clocks over 3 years, each with a random error of 5·10⁻¹⁶, plus a steady drift of 2·10⁻¹⁶ a year and no violation of universality. The drift is fitted and removed along with the annual template, and each dot is the mean of 13 consecutive comparisons, with its error bar, in units of 10⁻¹⁶. The fitted Δβ is -2.55·10⁻⁷ ± 3.43·10⁻⁷, consistent with zero; the dashed curves are the annual term a violation of twice that uncertainty, 6.86·10⁻⁷, would put in the ratio. The fitted uncertainty agrees with σ divided by the amplitude and the root of half the number of points to better than a tenth of a per cent, and an injected violation of 6.9·10⁻⁶ is recovered by the same fit.
Fig. 4 A synthetic record, not a measurement: 156 weekly comparisons with errors of 5 × 10⁻¹⁶, drift fitted out, shown as quarterly means. The fit returns Δβ = (−2.6 ± 3.4) × 10⁻⁷; the dashed curves are what a violation of 6.9 × 10⁻⁷ would look like.

The fitted value, (2.6±3.4)×107(-2.6 \pm 3.4) \times 10^{-7}, is consistent with zero, as it must be for data that contain no violation. The dashed curves show the size of annual term that a violation twice the uncertainty would produce, and the quarterly means scatter across them by more than their own amplitude. That is the usual state of affairs in a null test: no individual stretch of data rules the signal out, and the constraint comes from all 156 comparisons sharing one known phase. The figure’s own check is that an injected violation of twenty times the uncertainty is recovered by the same fit, which is the only way to know that a fit returning zero would have returned something else.

The uncertainty follows a simple rule. A sinusoid of amplitude AA sampled NN times with random error σ\sigma has its amplitude determined to about σ/(AN/2)\sigma/(A\sqrt{N/2}), and the drift term costs almost nothing as long as the record covers whole years. With A=1.65×1010A = 1.65 \times 10^{-10}:

What a year of comparisons is worth. The one-standard-deviation bound on a difference Δβ in how two clocks respond to the Sun's potential, against the random error of a single daily comparison, for 1, 3, 10 years of daily comparisons. The bound is the error divided by the annual amplitude 1.65·10⁻¹⁰ and by the root of half the number of comparisons, checked against a full fit that also removes a linear drift. At 10⁻¹⁵ per day, 1 yr gives 4.5·10⁻⁷, 3 yr gives 2.6·10⁻⁷, 10 yr gives 1.4·10⁻⁷; at 10⁻¹⁷ per day, 1 yr gives 4.5·10⁻⁹, 3 yr gives 2.6·10⁻⁹, 10 yr gives 1.4·10⁻⁹. Every factor of ten in clock quality is a factor of ten in the bound; every factor of ten in duration is only a factor of three.
Fig. 5 The one-standard-deviation bound on Δβ against the random error of a daily comparison, for 1, 3 and 10 years of daily records. At 10⁻¹⁵ a day, a year gives 4.5 × 10⁻⁷; at 10⁻¹⁷, 4.5 × 10⁻⁹.

The two slopes make an argument about where effort should go. The bound improves in direct proportion to the quality of a single comparison, and only as the square root of the number of comparisons. Ten years of daily comparisons at 101510^{-15} reach 1.4×1071.4 \times 10^{-7}. One year at 101710^{-17} reaches 4.5×1094.5 \times 10^{-9}, thirty times better. Microwave clocks — hydrogen masers, caesium fountains and rubidium clocks — compared over several years have constrained the difference in this way at around 10610^{-6} and below; optical clocks, whose single comparisons are a hundred times better, push the bound down by the corresponding factor in a fraction of the time. The orbit sets the size of the lever, and a laboratory can do nothing about that. Everything else is clock quality.

Why unlike clocks, and how unlike

The test only works for clocks that would respond differently to whatever might be varying, and “different kinds of clock” can be made precise.

An atomic clock counts the frequency of a transition, and that frequency depends on the constants of nature in a way set by the physics of the transition. Optical transitions are fixed mainly by the electrostatic energy of an electron in an atom, which scales with the Rydberg energy, times a factor that depends on the fine-structure constant α\alpha through relativistic corrections. Those corrections are small in light atoms and large in heavy ones, and their sign depends on the structure of the states involved. Microwave hyperfine clocks count the interaction between an electron’s magnetic moment and the nucleus, which brings in α2\alpha^2 along with nuclear magnetic moments and the ratio of electron to proton mass. The sensitivity coefficient KK of a clock is the fractional change in its frequency per fractional change in α\alpha, relative to the Rydberg.

How hard each clock leans on the fine-structure constant. The sensitivity coefficient of each clock's frequency to the fine-structure constant α — the fractional change in frequency per fractional change in α, relative to an atomic unit of energy — from atomic-structure calculations. caesium 2.83, rubidium 2.34, hydrogen maser 2, ytterbium ion, quadrupole 1, ytterbium lattice 0.31, strontium lattice 0.06, aluminium ion 0.008, mercury ion -2.94, ytterbium ion, octupole -5.95. Two clocks with the same coefficient cannot see a change in α at all, so a comparison is only as sensitive as the difference between its two coefficients. The widest pair in the list is caesium against ytterbium ion, octupole, a difference of 8.78. The microwave clocks' frequencies also depend on the ratio of electron to proton mass and on nuclear magnetic moments, which is a separate lever and not drawn.
Fig. 6 The sensitivity of nine clocks’ frequencies to the fine-structure constant, from atomic-structure calculations, from the octupole transition of the ytterbium ion at −5.95 to caesium’s hyperfine transition at 2.83. The widest pair differs by 8.78.

The spread is the point. The aluminium ion clock, at 0.008, is almost blind to α\alpha; strontium, at 0.06, nearly so. The mercury ion transition moves against α\alpha with a coefficient of −2.94, and the octupole transition of the ytterbium ion at −5.95. A comparison of two clocks with the same coefficient cannot see a change in α\alpha however good it is, and a comparison of two with very different coefficients amplifies one. The two transitions of a single ytterbium ion, at 1.00 and −5.95, are nearly seven apart and share an ion, a trap and most of their systematic errors — which is why that comparison has become one of the sharpest instruments of its kind.

This converts the abstract Δβ\Delta\beta into a question about nature. If the fine-structure constant depended on the gravitational potential as Δα/α=kαΔU/c2\Delta\alpha/\alpha = k_\alpha\,\Delta U/c^2, then two clocks with coefficients KAK_A and KBK_B would show Δβ=(KAKB)kα\Delta\beta = (K_A - K_B)\,k_\alpha. A bound on Δβ\Delta\beta from a pair with a large difference in KK is a bound on kαk_\alpha divided by that difference. Pairs of clocks with different dependence on the electron-to-proton mass ratio test the analogous coupling for that constant, and the microwave clocks, which depend on it strongly, are the natural instruments there.

Why anyone expects the constants to care

General relativity gives β=0\beta = 0 exactly, for every clock, because in it the laws of non-gravitational physics are the laws of special relativity in every freely falling frame, with the same constants everywhere. The reason to test it at the level of 10810^{-8} is that almost every attempt to join gravity to the rest of physics produces a small violation.

The mechanism is generic. Theories that unify the forces tend to contain additional fields — scalar fields, in the simplest versions — whose values set the strengths of interactions, and which are themselves sourced by mass. Near a large mass such a field takes a slightly different value, and so the coupling constants do too. A fine-structure constant that depends on distance from the Sun is the natural prediction of that kind of theory, with a coupling whose size the theory does not fix. The null test measures that coupling or bounds it.

There is also a reason the three parts of the equivalence principle are unlikely to fail separately. Suppose the energy levels of atoms depend on the gravitational potential. Then the rest energy of an atom — which includes the binding energy of its electrons and, far more importantly, of its nucleus — depends on position, so an atom feels a force from the gradient of that dependence in addition to its ordinary weight. The extra force is proportional to the binding energy fraction, which differs between elements. So a violation of local position invariance implies, through energy conservation alone, a violation of the universality of free fall. The argument, which goes back to Robert Dicke and Kenneth Nordtvedt, is why clock comparisons and torsion balances are testing the same theories from different directions, and why a result from one is a constraint on the other.

A clean sinusoid that no real record contains

The figures show the signal a violation would make, and they show it more cleanly than any laboratory can. The annual curve in the ratio is drawn as a perfect sinusoid at a known phase on a flat line; a real ratio record has gaps where one clock was down for maintenance, steps where a component was replaced, a drift that is not quite linear, and seasonal changes in the laboratory whose phase happens to sit within two weeks of perihelion. The synthetic record has random noise and a straight drift and nothing else, which is what makes its uncertainty follow the simple rule — and a real analysis spends most of its effort on everything that record omits. The figures show what the test is looking for and how its sensitivity scales. What they cannot show is the work of convincing anyone that a nonzero amplitude, if one appeared, came from the Sun rather than from the building.

Where the comparison stops being clean

The annual signal has competitors with the same period. Laboratory temperature, humidity and pressure all have yearly cycles, and a clock’s frequency responds to each at some level. The orbital phase peaks within two weeks of the northern midwinter, so a seasonal systematic in a northern laboratory has almost the same phase as the signal. Comparisons at several sites, and especially in the southern hemisphere, where the seasons are reversed and the orbit is not, are what separate the two.

The lever is small, and cannot be made larger on the Earth. The swing of 3.3×10103.3 \times 10^{-10} is fixed by the orbit. A clock carried closer to the Sun would see a much larger change in potential, and missions have been proposed to fly an optical clock on an elliptical orbit reaching inside the orbit of Mercury, where the change would be more than a hundred times larger. Until then the terrestrial test gains only by making the clocks better.

The clock model is a parametrisation, not a theory. Writing the response as (1+β)(1+\beta) assumes the violation is linear in the potential and the same for all changes in potential. A theory could produce effects that depend on the potential in other ways, or on its gradient, or on the Earth’s motion through a cosmological field, and those would appear in clock comparisons with other periods. Searches for daily and sidereal variations in clock ratios, and for oscillations at frequencies set by hypothetical light dark-matter fields, are the same instrument pointed at those alternatives.

And the sensitivity coefficients are calculated. The KK values come from relativistic atomic-structure calculations. For the heavy ions with large coefficients those calculations are good to a few per cent, which is ample for a bound but would matter for interpreting a detection.

Still open: whether some clock has a different time

The redshift has now been followed from a photon climbing a tower, through the proof that the measurement forbids flat spacetime and the use of a clock as a surveying instrument, to the test of whether the thing being measured is time at all or only a property of particular clocks. The size of the shift is known for hydrogen masers to a few parts in 10510^5; its universality across different kinds of clock is known to parts in 10810^8 or better, which is the stronger statement and the one the geometric picture depends on.

The habit worth carrying away is how to turn an effect into a null test. When a quantity is predicted to be the same for every instrument, compare two unlike instruments and look for the variation with a known period; the shared effect cancels, and what is left is the claim itself. The annual swing of the Sun’s potential is invisible to every clock on the Earth taken one at a time, and becomes a clean experiment the moment two are compared. The same arrangement appears in the comparison of clocks flown in opposite directions, where two clocks share everything except the direction of travel.

What is not known is whether the bound will simply keep improving or whether a violation is waiting a few orders of magnitude further down. Theories that unify the forces do not predict its size, and the best clocks have improved by roughly a factor of ten a decade for half a century. Each factor tightens the null result by the same amount, and a clock flown towards the Sun would add a factor of a hundred or more at once. A nonzero result anywhere in that range would mean that the constants of nature, and with them every clock, depend on where in the universe the measurement is made.

Part 4 of 4

This essay is one argument about Gravitational redshift. The others:

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

ClocksEccentricityEquivalence principleFine structure constantGeneral relativityGravitational redshiftLocal position invarianceNull testPotentialPrecision