The clocks that must all slow together
Assumes: The clock that measures a height · The floor that cannot be told from gravity
A clock lower down runs slow, and the fraction is the difference in gravitational potential divided by the square of the speed of light. That result has been confirmed many ways, and each confirmation measures the size of the shift with one kind of clock: the Mössbauer line of iron-57 in a tower, a hydrogen maser on a rocket, an aluminium ion lifted by a third of a metre. Each experiment checks that one clock slows by the predicted amount.
There is a second claim hidden in the first, and it is the stronger one. The argument that turned the redshift into geometry — the parallelogram of clock comparisons that fails to close — needs every clock to agree. If a caesium clock slowed by one fraction and a hydrogen maser by a slightly different one, there would be no single “rate of time” at a height, only a collection of rates of different mechanisms, and the redshift would be a fact about atoms rather than about spacetime. The claim that the fraction is the same for every clock whatever it is made of is called local position invariance. It says the outcome of any local experiment that does not involve gravity is independent of where in the gravitational field it is done. It is one of the three parts of the equivalence principle, and it can be tested in a way that needs no second location, no rocket and no tower.
A potential the Earth provides for nothing
The Earth’s orbit is not quite circular. Its eccentricity — the length of the vector that points at perihelion — is 0.0167, which moves the planet about five million kilometres closer to the Sun in early January than in early July. The depth of the Sun’s potential at the Earth is , a little under on average, and it varies as the distance does.
The swing is times the mean, . That is not small by the standards of the clocks involved. It is about half the by which the Earth’s own gravity slows a clock at its surface, and it is hundreds of millions of times larger than the uncertainty of the best clocks now operating. Measured against a clock far out in the solar system, every clock on the Earth runs faster in July than in January by that fraction, and the effect is a routine correction in timekeeping that refers terrestrial clocks to the solar system’s centre of mass.
Yet no measurement on the Earth can see it directly. Two clocks side by side both slow by the same fraction, so their comparison shows nothing, and a clock compared with another clock anywhere else on the Earth is in the same position. The Earth is falling freely round the Sun, and in a freely falling laboratory the potential of the body it falls towards is not locally detectable — only its gradient across the laboratory is, as a tide. The annual swing is real, and the only thing the equivalence principle allows it to affect is the comparison with a clock that is not falling with the Earth.
That is the test. The principle says that every clock on the Earth slows by exactly the swing. If some kind of clock responded a little more strongly or a little less, then two different kinds of clock kept side by side would drift apart and back again once a year, in step with the orbit.
Two clocks, one ratio
Write the response of a clock of type A to a change in potential as , with for a clock that obeys the principle, and similarly for type B. The frequency ratio of the two then changes by . Everything common to the two clocks cancels: the shared slowing, the unknown potential of the laboratory, the Earth’s rotation, the variation in the Earth’s own field. What is left depends only on the difference in how the two mechanisms respond.
The picture shows the scale of the problem and why it is tractable. The shared signal is , and it cancels exactly in the ratio under the principle. A violation at produces a ratio signal of , a million times smaller, with the same annual period and the same phase as the orbit. That phase is known in advance to within hours, which matters a great deal: the analysis looks for a sinusoid of known period and known phase and fits only its amplitude, and that is a far easier measurement than searching for an unknown signal.
The design is a null test, and it has the virtue all null tests share. A measurement of the redshift’s size compares a predicted number with a measured one, and every systematic error in the measurement is an error in the comparison. A null test compares zero with a measured number, and many systematic errors either cancel or show up with the wrong period. Temperature varies annually, of course, and so does humidity, and a laboratory’s seasonal cycle is the main thing such an analysis has to rule out — but it has to be ruled out only at the phase of the orbit, and perihelion falls within a couple of weeks of the northern winter solstice, which is exactly why the careful versions of the experiment are done at more than one site and in both hemispheres.
The test of size that the orbit also offers
The absolute size of the shift can be tested the same way, with a clock that carries its own changing potential.
In August 2014 two satellites of the European Galileo navigation system were placed in the wrong orbits by a fault in their launcher’s upper stage. Instead of circular orbits they went into elliptical ones, with eccentricity around 0.16, and although the orbits were later raised they stayed noticeably eccentric. For navigation that was a loss. For relativity it was an opportunity, because each satellite carries a passive hydrogen maser and a rubidium clock and is tracked continuously, and an eccentric orbit swings a clock between deep and shallow potential, and between fast and slow, twice a day.
Near perigee the satellite is deeper in the potential and moving faster, and both effects slow its clock, so the offset accumulates; near apogee both reverse. The two do not cancel over the orbit the way they do in comparing orbits of different heights. The periodic part integrates to , where is the eccentric anomaly, and the same Kepler-orbit averaging that makes every orbit of one period lose the same total time is what leaves this periodic term behind as the only thing that depends on the shape. For the eccentric Galileo orbit it is ±381 nanoseconds. A navigation clock is compared with ground clocks at the level of a fraction of a nanosecond, so the signal is enormous.
Two independent analyses of about three years of data, published in 2018, confirmed the predicted amplitude to a few parts in , improving on Gravity Probe A’s 1976 result by a factor of a few. It is worth being clear about what that measures. It tests the size of the shift for hydrogen masers — that for that kind of clock is zero to a few parts in . It says nothing about whether a caesium clock would have shown the same, and that is the question the null test answers much more precisely.
What a record of nothing is worth
A null result is a number with an uncertainty, and the useful thing to understand is how the uncertainty is set.
Suppose two unlike clocks are compared once a week for three years, each comparison with a random error of . Real clocks also drift, slowly and steadily, as they age, so the fit has three terms: a constant, a linear drift, and the known annual shape multiplied by an unknown . The figure below is a synthetic record made that way, with no violation in it, to show what the analysis sees.
The fitted value, , is consistent with zero, as it must be for data that contain no violation. The dashed curves show the size of annual term that a violation twice the uncertainty would produce, and the quarterly means scatter across them by more than their own amplitude. That is the usual state of affairs in a null test: no individual stretch of data rules the signal out, and the constraint comes from all 156 comparisons sharing one known phase. The figure’s own check is that an injected violation of twenty times the uncertainty is recovered by the same fit, which is the only way to know that a fit returning zero would have returned something else.
The uncertainty follows a simple rule. A sinusoid of amplitude sampled times with random error has its amplitude determined to about , and the drift term costs almost nothing as long as the record covers whole years. With :
The two slopes make an argument about where effort should go. The bound improves in direct proportion to the quality of a single comparison, and only as the square root of the number of comparisons. Ten years of daily comparisons at reach . One year at reaches , thirty times better. Microwave clocks — hydrogen masers, caesium fountains and rubidium clocks — compared over several years have constrained the difference in this way at around and below; optical clocks, whose single comparisons are a hundred times better, push the bound down by the corresponding factor in a fraction of the time. The orbit sets the size of the lever, and a laboratory can do nothing about that. Everything else is clock quality.
Why unlike clocks, and how unlike
The test only works for clocks that would respond differently to whatever might be varying, and “different kinds of clock” can be made precise.
An atomic clock counts the frequency of a transition, and that frequency depends on the constants of nature in a way set by the physics of the transition. Optical transitions are fixed mainly by the electrostatic energy of an electron in an atom, which scales with the Rydberg energy, times a factor that depends on the fine-structure constant through relativistic corrections. Those corrections are small in light atoms and large in heavy ones, and their sign depends on the structure of the states involved. Microwave hyperfine clocks count the interaction between an electron’s magnetic moment and the nucleus, which brings in along with nuclear magnetic moments and the ratio of electron to proton mass. The sensitivity coefficient of a clock is the fractional change in its frequency per fractional change in , relative to the Rydberg.
The spread is the point. The aluminium ion clock, at 0.008, is almost blind to ; strontium, at 0.06, nearly so. The mercury ion transition moves against with a coefficient of −2.94, and the octupole transition of the ytterbium ion at −5.95. A comparison of two clocks with the same coefficient cannot see a change in however good it is, and a comparison of two with very different coefficients amplifies one. The two transitions of a single ytterbium ion, at 1.00 and −5.95, are nearly seven apart and share an ion, a trap and most of their systematic errors — which is why that comparison has become one of the sharpest instruments of its kind.
This converts the abstract into a question about nature. If the fine-structure constant depended on the gravitational potential as , then two clocks with coefficients and would show . A bound on from a pair with a large difference in is a bound on divided by that difference. Pairs of clocks with different dependence on the electron-to-proton mass ratio test the analogous coupling for that constant, and the microwave clocks, which depend on it strongly, are the natural instruments there.
Why anyone expects the constants to care
General relativity gives exactly, for every clock, because in it the laws of non-gravitational physics are the laws of special relativity in every freely falling frame, with the same constants everywhere. The reason to test it at the level of is that almost every attempt to join gravity to the rest of physics produces a small violation.
The mechanism is generic. Theories that unify the forces tend to contain additional fields — scalar fields, in the simplest versions — whose values set the strengths of interactions, and which are themselves sourced by mass. Near a large mass such a field takes a slightly different value, and so the coupling constants do too. A fine-structure constant that depends on distance from the Sun is the natural prediction of that kind of theory, with a coupling whose size the theory does not fix. The null test measures that coupling or bounds it.
There is also a reason the three parts of the equivalence principle are unlikely to fail separately. Suppose the energy levels of atoms depend on the gravitational potential. Then the rest energy of an atom — which includes the binding energy of its electrons and, far more importantly, of its nucleus — depends on position, so an atom feels a force from the gradient of that dependence in addition to its ordinary weight. The extra force is proportional to the binding energy fraction, which differs between elements. So a violation of local position invariance implies, through energy conservation alone, a violation of the universality of free fall. The argument, which goes back to Robert Dicke and Kenneth Nordtvedt, is why clock comparisons and torsion balances are testing the same theories from different directions, and why a result from one is a constraint on the other.
A clean sinusoid that no real record contains
The figures show the signal a violation would make, and they show it more cleanly than any laboratory can. The annual curve in the ratio is drawn as a perfect sinusoid at a known phase on a flat line; a real ratio record has gaps where one clock was down for maintenance, steps where a component was replaced, a drift that is not quite linear, and seasonal changes in the laboratory whose phase happens to sit within two weeks of perihelion. The synthetic record has random noise and a straight drift and nothing else, which is what makes its uncertainty follow the simple rule — and a real analysis spends most of its effort on everything that record omits. The figures show what the test is looking for and how its sensitivity scales. What they cannot show is the work of convincing anyone that a nonzero amplitude, if one appeared, came from the Sun rather than from the building.
Where the comparison stops being clean
The annual signal has competitors with the same period. Laboratory temperature, humidity and pressure all have yearly cycles, and a clock’s frequency responds to each at some level. The orbital phase peaks within two weeks of the northern midwinter, so a seasonal systematic in a northern laboratory has almost the same phase as the signal. Comparisons at several sites, and especially in the southern hemisphere, where the seasons are reversed and the orbit is not, are what separate the two.
The lever is small, and cannot be made larger on the Earth. The swing of is fixed by the orbit. A clock carried closer to the Sun would see a much larger change in potential, and missions have been proposed to fly an optical clock on an elliptical orbit reaching inside the orbit of Mercury, where the change would be more than a hundred times larger. Until then the terrestrial test gains only by making the clocks better.
The clock model is a parametrisation, not a theory. Writing the response as assumes the violation is linear in the potential and the same for all changes in potential. A theory could produce effects that depend on the potential in other ways, or on its gradient, or on the Earth’s motion through a cosmological field, and those would appear in clock comparisons with other periods. Searches for daily and sidereal variations in clock ratios, and for oscillations at frequencies set by hypothetical light dark-matter fields, are the same instrument pointed at those alternatives.
And the sensitivity coefficients are calculated. The values come from relativistic atomic-structure calculations. For the heavy ions with large coefficients those calculations are good to a few per cent, which is ample for a bound but would matter for interpreting a detection.
Still open: whether some clock has a different time
The redshift has now been followed from a photon climbing a tower, through the proof that the measurement forbids flat spacetime and the use of a clock as a surveying instrument, to the test of whether the thing being measured is time at all or only a property of particular clocks. The size of the shift is known for hydrogen masers to a few parts in ; its universality across different kinds of clock is known to parts in or better, which is the stronger statement and the one the geometric picture depends on.
The habit worth carrying away is how to turn an effect into a null test. When a quantity is predicted to be the same for every instrument, compare two unlike instruments and look for the variation with a known period; the shared effect cancels, and what is left is the claim itself. The annual swing of the Sun’s potential is invisible to every clock on the Earth taken one at a time, and becomes a clean experiment the moment two are compared. The same arrangement appears in the comparison of clocks flown in opposite directions, where two clocks share everything except the direction of travel.
What is not known is whether the bound will simply keep improving or whether a violation is waiting a few orders of magnitude further down. Theories that unify the forces do not predict its size, and the best clocks have improved by roughly a factor of ten a decade for half a century. Each factor tightens the null result by the same amount, and a clock flown towards the Sun would add a factor of a hundred or more at once. A nonzero result anywhere in that range would mean that the constants of nature, and with them every clock, depend on where in the universe the measurement is made.
Part 4 of 4
This essay is one argument about Gravitational redshift. The others:
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
ClocksEccentricityEquivalence principleFine structure constantGeneral relativityGravitational redshiftLocal position invarianceNull testPotentialPrecision
- The delay that is not a bend general relativity, gravitational redshift
- The disc that cannot be spun equivalence principle, general relativity
- The fall that does not depend on what is falling equivalence principle, precision
- The longest way round is the shortest clock equivalence principle, gravitational redshift