The fall that does not depend on what is falling
Assumes: The floor that cannot be told from gravity · The pendulum, and the small lie that makes it simple
Two objects of different composition, released together, fall together. That is the observation the whole of general relativity is built on: if gravity affects everything the same way, it can be described as a property of spacetime rather than as a force, and the floor that cannot be told from gravity is where that step is taken.
The observation is therefore worth checking harder than almost anything else in physics, and it has been. The quantity measured is
the fractional difference in the accelerations of two materials in the same field, and the whole history of the subject is a sequence of upper limits on it.
Why dropping things stops working
The obvious experiment is to drop two objects and watch. Done carefully, in a vacuum, it reaches perhaps a part in a hundred — good enough to settle the question Galileo was asking and nowhere near good enough for anything since.
The trouble is that it measures two large quantities and subtracts them. Both objects accelerate at 9.8 m/s²; the difference sought is smaller than that by a factor to be determined, and every source of error — the release mechanism, the residual air, the timing — enters at full size rather than at the size of the effect.
A modern drop test does better than a part in a hundred, but not by much and not for the reason people expect. A free-fall interferometer can time a falling corner cube to a few parts in , which sounds promising — and the limit on η from comparing two such drops is nothing like that, because the two objects cannot be dropped at the same instant from the same place. Every difference between the two drops enters the answer, and the vibration of the floor between them is larger than the effect.
Every advance since has come from making the instrument report the difference directly. That is what a null experiment is, and this subject is where the technique was perfected.
Counting swings
The first serious improvement was Newton’s, and it is still an instructive design. Two pendulums of identical length, with bobs of different materials: if the gravitational and inertial masses differ by η, the period differs by η/2, and the difference can be accumulated by counting.
Newton reported no difference at the part-in-a-thousand level. Bessel, working in the 1830s with far better apparatus, reached a few parts in a hundred thousand — which by the figure means counting through the better part of a day.
The pendulum is a well-understood instrument by this point in the ladder — the pendulum and its small lie sets out what the small-angle formula assumes, the period that depends on the swing computes the correction exactly, and the length nobody has to measure is the reversible design that removes the hardest systematic of all. Every one of those refinements was made in the service of measuring , and all of them apply here unchanged, because a composition test with pendulums is a comparison of two values of inferred from two different bobs.
Past that the method stops, and not for want of patience. A day is enough time for the temperature to change, the amplitude to decay, the air density to shift and the knife edge to wear, and each of those moves the period by more than the effect being looked for. The instrument had to change, not the observer.
The signal the Earth’s spin provides
Eötvös’s answer was a torsion balance: two masses of different material on the ends of a horizontal beam hung from a fine fibre. The question is what would twist it.
A plumb line hangs along the sum of gravity and the centrifugal effect. If the two masses respond to those two differently — one being a gravitational pull and the other an inertial reaction — their preferred directions differ slightly, and the beam feels a horizontal torque that reverses when the apparatus is turned round.
The available acceleration is about a thousandth of gravity, and giving up that factor of a thousand buys four orders of magnitude, because the instrument measures the difference and nothing else. A quartz fibre twisting by a microradian is a force a pendulum could not have noticed at all. Eötvös reached five parts in , and it stood as the best measurement for forty years.
Turning the apparatus is the other half of the design. A torque that reverses when the instrument is rotated by 180° is distinguishable from every drift that does not, and modulating the signal that way is what makes the measurement possible rather than merely sensitive.
Choosing what to fall towards
Dicke’s improvement in the 1960s was to change the source.
The Sun’s pull at the Earth is 5.9 mm/s² — a third of what the spin makes available from the Earth itself, and therefore a step backwards in size. What it buys is a frequency. The Earth’s rotation carries the apparatus round once a day, so the Sun’s direction relative to the instrument sweeps through a full circle every twenty-four hours, and a composition-dependent effect would appear as a signal at exactly that period with a known phase.
That is worth more than the factor of three lost, because the instrument’s own drifts do not respect a sidereal day. Dicke reached a part in and Braginsky a part in , both using the Sun, and modern ground experiments use a rotating turntable to put the signal wherever the apparatus happens to be quietest.
What a null instrument buys
The word “null” is doing a lot of work above, and it is worth unpacking because the principle is general.
An instrument that measures a quantity and looks for a small change in it must have a full-scale accuracy comparable to the change: to see one part in of , everything about the instrument’s calibration, linearity and drift must be good to that level across its whole range. An instrument that measures the difference between two nearly equal quantities has to be that good only over the difference, which is small — so its calibration can be poor, its linearity irrelevant, and its zero can drift as long as the drift is slower than the modulation.
A torsion balance is the extreme case. Its fibre never twists by more than a microradian, and everything about the instrument is characterised at that amplitude. Its absolute calibration matters only in converting a final null into a limit, and a twenty per cent error there costs twenty per cent of a limit that is being quoted to one significant figure.
The same reasoning is why almost every high-precision measurement in physics is a comparison rather than a determination. A frequency is compared with a frequency, a mass with a mass, a length with a wavelength — and the quantities that are known best are the ratios, which is a fact about instruments rather than about nature.
Where the last factor came from
The satellite went right, not down. Its accelerometer resolves a differential acceleration within an order of magnitude of what the best torsion balance manages; what it has is a drive of eight metres per second squared instead of seventeen millimetres, because a body in orbit is in free fall in the whole of the Earth’s field rather than in the small horizontal residue the spin leaves behind.
There is a second advantage to orbit that the figure does not draw, and it is nearly as large. On the ground, a torsion balance is suspended, and the suspension has to hold the whole weight of the masses while permitting a microradian of twist. In orbit nothing has to be held up at all, so the only forces on the test masses are the ones the experiment applies deliberately — which removes the fibre, its thermal noise, its creep, and the whole family of problems that comes with hanging a precision instrument in a gravitational field. The same argument is what makes a drag-free satellite the natural home for any experiment whose signal is an acceleration, and it is the reason gravitational-wave detection is moving the same way, as what the instrument actually hears sets out for a different measurement.
MICROSCOPE flew two pairs of concentric cylindrical test masses — one pair of platinum and titanium, the other of platinum and platinum as a control — each held at the centre of the satellite by electrostatic forces. The force needed to keep the two coaxial is the measurement: if they fell differently, one would drift and the servo would have to push. The satellite rotated to put that signal at a chosen frequency, well away from the orbital period and from anything thermal. The final limit is about a part in .
The lesson is one about instrument design rather than about gravity. When a measurement is a ratio, look at the denominator before improving the numerator — and here the denominator had been fixed at a thousandth of gravity by the accident of doing the experiment on a rotating planet.
The one systematic that will not go away
Every experiment of this kind ends up limited by the same thing, and naming it explains a good deal of the design of all of them.
A composition test compares two masses that are not at the same place. Whatever they are made of, they sit at slightly different positions, and the gravitational field is not uniform: a gradient across the separation produces a differential acceleration that has nothing to do with composition. Worse, the gradient comes partly from the apparatus itself — the vacuum chamber, the pump, the experimenter — and it varies as those move.
The countermeasures are geometric. A torsion balance is built so that the two masses sit at the same distance from the axis and the beam is balanced against every low-order moment of the field, not merely the first. MICROSCOPE’s test masses are concentric cylinders sharing a centre, which cancels the gradient exactly to first order and is the reason for a shape that is otherwise inconvenient. And the whole thing is rotated, so a gradient effect fixed in the laboratory appears at a different frequency from the signal.
Even so, gravity gradients are the leading systematic in every one of these measurements, and the published limits are dominated by how well they were modelled rather than by how quiet the instrument was. That the field the experiment is testing is also its worst noise source is not a coincidence — it is a property of any experiment sensitive enough to be worth doing.
What a null result excludes
A sequence of zeros is not a sequence of failures, and it is worth being specific about what each one buys.
Any long-range force that couples to something other than mass-energy — to the number of baryons, say, or to the number of neutrons minus protons — would produce a composition-dependent acceleration, because different materials have different ratios of those to their mass. Platinum and titanium differ in neutron fraction by a few per cent, so a limit on η is a limit on the strength of such a force divided by gravity’s.
Many extensions of the standard model contain a light scalar that would do exactly this, and the limits from these experiments are the strongest constraints on several of them — stronger than anything an accelerator provides, because the effect grows with the number of particles rather than with the energy.
So the experiment is a search for new physics, conducted by looking as hard as possible for nothing. The comparison mass is chosen to maximise the contrast in whatever charge the hypothetical force might couple to, which is why the pairs are exotic: beryllium and titanium, platinum and titanium, and in one experiment a test mass made of material from a meteorite.
The reanalysis that nearly found something
In 1986 a group re-examined Eötvös’s published data and reported that the residuals correlated with the ratio of baryon number to mass across his sample of materials — a signature of exactly the kind of fifth force the limits are meant to exclude. The claimed effect was at the level of a part in of gravity with a range of a few hundred metres.
It set off a decade of experiments. New torsion balances were built, mine shafts and television towers were used as sources, submarines measured gravity at depth. Nothing was found, and the reanalysis is now generally attributed to a systematic in the original data rather than to physics — most plausibly a gravity gradient from the building the balance was in, which correlated with the sample materials because of the order they were measured in.
The episode is worth remembering for two reasons. It is the clearest case in this subject of a real effect being claimed from a real dataset by a careful reanalysis, and being wrong; and the response to it — build the apparatus again, differently, in several places — is exactly the response that a claim of this kind should get. The limits improved by two orders of magnitude in the process.
Where the model stops
The principle is not a theorem. Nothing in mechanics requires that the mass appearing in and the mass appearing in the gravitational force be the same quantity; that they are is an experimental fact, and it is the fact that lets gravity be described as geometry. Every step from the floor that cannot be told from gravity onwards depends on it, which is why the limit is worth pushing.
Only the weak principle is tested here. That bodies fall alike says nothing directly about whether a body’s own gravitational binding energy falls the same way, which is the strong equivalence principle and is tested by lunar laser ranging — the Earth and the Moon have different fractions of their mass in gravitational binding, so a violation would distort the Moon’s orbit.
The test masses are macroscopic and unpolarised. A separate family of experiments asks whether a spinning body falls differently, or whether antimatter does, and neither is covered by anything here. The antimatter question was open until very recently and is now answered at the per-cent level, which is where the composition tests were in 1687.
The satellite’s limit is systematics-limited, not statistics-limited. More flight time would not have improved it; the residual is a thermal effect on the test masses whose size had to be modelled rather than measured, which is the usual condition of a mature precision experiment.
And η is not the only parameter. A violation could depend on the distance to the source as well as on the composition, so a limit obtained with the Earth as the source does not directly bound a force with a range of a few metres. Those are constrained by their own, much weaker, short-range experiments.
What the pictures cannot show
The history figure draws each limit as a point, and a limit is a statement with a confidence level and a set of assumed systematic corrections behind it. Several of the points involved reanalyses that moved them, and one of the historical results — Eötvös’s own, reanalysed in 1986 — was claimed for a while to show a positive signal correlated with baryon number. That claim did not survive, but it is the reason the field spent a decade doing the experiment again with new apparatus.
The budget figure treats each experiment as a single resolution and a single drive, and both are averages over a measurement whose noise is a spectrum. What matters in practice is the noise at the modulation frequency, which is why the choice of that frequency — the rotation rate of a turntable, the spin of a satellite — is as much of the design as the accelerometer.
Where the ladder goes next
The equivalence-principle ladder began with the floor that cannot be told from gravity, which states the principle and uses it, and continued to the term free fall cannot remove, where the tidal difference survives every choice of frame and is what curvature actually is. This rung asks how well the first of those is known, and finds a number with fifteen digits of zero behind it.
The rung after it is the strong principle: whether a body’s own gravitational energy falls at the same rate as the rest of it, which cannot be tested with laboratory masses because their binding energy is a part in of their mass, and which is measured instead by watching the Moon. The habit worth carrying is the one this rung is built on: when a quantity has been measured to be zero for three hundred years, the interesting question is what each new digit rules out — because a limit with nothing behind it is a limit nobody should have paid for.
Part 3 of 4
This essay is one argument about Equivalence principle. The others:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
Centrifugal forceEquivalence principleFree fallGravitational massInertial massMeasurementModulationNull experimentPendulumPrecisionSystematic errorTorsion balance
- The clock that measures a height equivalence principle, measurement, precision
- Whether a charge on a table glows equivalence principle, free fall, measurement
- How big now is equivalence principle, measurement
- The clocks that must all slow together equivalence principle, precision
- The horizon that nothing marks equivalence principle, free fall
- The longest way round is the shortest clock equivalence principle, free fall