Astrophysics

The fall that does not depend on what is falling

Everything falls at the same rate, and the statement has been tested for three hundred years by people looking for the exception. Twelve orders of magnitude have been added to the limit and every measurement has returned zero. The last three orders came not from a better instrument but from finding something bigger to fall towards.

Assumes: The floor that cannot be told from gravity · The pendulum, and the small lie that makes it simple

Two objects of different composition, released together, fall together. That is the observation the whole of general relativity is built on: if gravity affects everything the same way, it can be described as a property of spacetime rather than as a force, and the floor that cannot be told from gravity is where that step is taken.

The observation is therefore worth checking harder than almost anything else in physics, and it has been. The quantity measured is

η=2(a1a2)a1+a2\eta = \frac{2(a_1 - a_2)}{a_1 + a_2}

the fractional difference in the accelerations of two materials in the same field, and the whole history of the subject is a sequence of upper limits on it.

Three centuries of finding nothing. The upper limit on η against the year it was set, on a logarithmic scale. Newton (1687) reached 1e-3; Bessel (1832) reached 2e-5; Eötvös (1922) reached 5e-9; Dicke (1964) reached 1e-11; Braginsky (1972) reached 1e-12; Eöt-Wash (2008) reached 2e-13; MICROSCOPE (2022) reached 1e-15. That is 12 orders of magnitude in 335 years, and every one of those measurements returned zero. A sequence of null results is not a sequence of failures. Each one is a statement that a principle assumed by every theory of gravity holds to a new level, and each new level excludes a class of theories that would have shown a departure there — a long-range force coupling to something other than mass-energy, a scalar partner to the graviton, a violation arising at some energy scale. The measurement is worth making again precisely because it has always come out the same way, which is what makes any departure decisive.
Fig. 1 The upper limit on η against the year it was set, over three and a half centuries and twelve orders of magnitude. Every one of those measurements returned zero.

Why dropping things stops working

The obvious experiment is to drop two objects and watch. Done carefully, in a vacuum, it reaches perhaps a part in a hundred — good enough to settle the question Galileo was asking and nowhere near good enough for anything since.

The trouble is that it measures two large quantities and subtracts them. Both objects accelerate at 9.8 m/s²; the difference sought is smaller than that by a factor to be determined, and every source of error — the release mechanism, the residual air, the timing — enters at full size rather than at the size of the effect.

A modern drop test does better than a part in a hundred, but not by much and not for the reason people expect. A free-fall interferometer can time a falling corner cube to a few parts in 10910^9, which sounds promising — and the limit on η from comparing two such drops is nothing like that, because the two objects cannot be dropped at the same instant from the same place. Every difference between the two drops enters the answer, and the vibration of the floor between them is larger than the effect.

Every advance since has come from making the instrument report the difference directly. That is what a null experiment is, and this subject is where the technique was perfected.

Counting swings

The first serious improvement was Newton’s, and it is still an instructive design. Two pendulums of identical length, with bobs of different materials: if the gravitational and inertial masses differ by η, the period differs by η/2, and the difference can be accumulated by counting.

How long a pendulum has to swing. The number of swings a pair of pendulums with different bobs must be counted through before a composition-dependent difference of η would show up as a whole period of disagreement. The period goes as the square root of the ratio of gravitational to inertial mass, so a fractional difference in period is half of η, and counting is the only amplification available. To reach η = 1e-3 takes 2e+3 swings; To reach η = 1e-5 takes 2e+5 swings; To reach η = 1e-8 takes 2e+8 swings. Newton did this with pendulums of equal length carrying boxes of different materials and reported no difference at the part-in-a-thousand level; Bessel repeated it far more carefully in the 1830s and reached a few parts in a hundred thousand. Past that the method stops. A day is about eighty thousand swings of a one-second pendulum, and the amplitude, the temperature, the air and the knife edge all change over a day by more than the effect being looked for. The next factor of a thousand needed a different instrument rather than a more patient observer.
Fig. 2 The number of swings that must be counted before a composition-dependent difference of η shows up as one whole period of disagreement. The period goes as the square root of the mass ratio, so a fractional period difference is half of η, and counting is the only amplification available.

Newton reported no difference at the part-in-a-thousand level. Bessel, working in the 1830s with far better apparatus, reached a few parts in a hundred thousand — which by the figure means counting through the better part of a day.

The pendulum is a well-understood instrument by this point in the ladder — the pendulum and its small lie sets out what the small-angle formula assumes, the period that depends on the swing computes the correction exactly, and the length nobody has to measure is the reversible design that removes the hardest systematic of all. Every one of those refinements was made in the service of measuring gg, and all of them apply here unchanged, because a composition test with pendulums is a comparison of two values of gg inferred from two different bobs.

Past that the method stops, and not for want of patience. A day is enough time for the temperature to change, the amplitude to decay, the air density to shift and the knife edge to wear, and each of those moves the period by more than the effect being looked for. The instrument had to change, not the observer.

The signal the Earth’s spin provides

Eötvös’s answer was a torsion balance: two masses of different material on the ends of a horizontal beam hung from a fine fibre. The question is what would twist it.

The sideways push the Earth's spin provides. The horizontal acceleration available to a torsion balance, against latitude. Gravity points towards the centre of the Earth and the centrifugal effect points away from its axis, and those two directions differ everywhere except at the equator and the poles. A plumb line hangs along their sum; two masses that respond differently to the two would hang along slightly different directions, and the difference is a horizontal torque on a suspended beam. The signal peaks at 45° at 1.70e-2 m/s², which is 1.7e-3 of gravity. At 20° it is 1.09e-2 m/s²; At 45° it is 1.70e-2 m/s²; At 70° it is 1.09e-2 m/s². That factor of a thousand down on gravity is the price of the method and the whole of its advantage. The instrument is a null device: it does not measure a large force and look for a small difference, it measures the difference directly, and a fibre that twists by a microradian under a torque a pendulum could not notice is what buys the four orders of magnitude Eötvös gained over Bessel.
Fig. 3 The horizontal acceleration available to a torsion balance, against latitude. Gravity points at the centre of the Earth and the centrifugal effect points away from its axis; those directions differ everywhere except at the equator and the poles, so the signal vanishes at both and peaks at 45°.

A plumb line hangs along the sum of gravity and the centrifugal effect. If the two masses respond to those two differently — one being a gravitational pull and the other an inertial reaction — their preferred directions differ slightly, and the beam feels a horizontal torque that reverses when the apparatus is turned round.

The available acceleration is about a thousandth of gravity, and giving up that factor of a thousand buys four orders of magnitude, because the instrument measures the difference and nothing else. A quartz fibre twisting by a microradian is a force a pendulum could not have noticed at all. Eötvös reached five parts in 10910^9, and it stood as the best measurement for forty years.

Turning the apparatus is the other half of the design. A torque that reverses when the instrument is rotated by 180° is distinguishable from every drift that does not, and modulating the signal that way is what makes the measurement possible rather than merely sensitive.

Choosing what to fall towards

Dicke’s improvement in the 1960s was to change the source.

What is available to pull on the two masses. The acceleration each possible source offers a composition test, on a logarithmic scale spanning ten orders of magnitude. the Galaxy: 2.0e-10 m/s²; the Sun: 5.9e-3 m/s²; the Earth, sideways: 1.7e-2 m/s²; the Earth, all of it: 8.0e+0 m/s². The choice matters because η is a ratio: the differential acceleration an instrument can resolve is fixed by the instrument, so the reach in η is that resolution divided by whatever is doing the pulling. A ground-based torsion balance can only use the horizontal part of the Earth's field, which the spin makes available and which is a thousandth of gravity. Dicke's improvement was to use the Sun instead — a smaller acceleration but one that is fully horizontal twice a day and therefore modulated at a clean, known frequency, which is worth more than its size. The whole of the Earth's field is available only to something in free fall, and using it is the single reason a satellite test beats the best ground instrument.
Fig. 4 The acceleration each possible source offers, on a logarithmic scale spanning ten orders of magnitude. η is a ratio: the differential acceleration an instrument can resolve is fixed by the instrument, and the reach in η is that resolution divided by whatever is pulling.

The Sun’s pull at the Earth is 5.9 mm/s² — a third of what the spin makes available from the Earth itself, and therefore a step backwards in size. What it buys is a frequency. The Earth’s rotation carries the apparatus round once a day, so the Sun’s direction relative to the instrument sweeps through a full circle every twenty-four hours, and a composition-dependent effect would appear as a signal at exactly that period with a known phase.

That is worth more than the factor of three lost, because the instrument’s own drifts do not respect a sidereal day. Dicke reached a part in 101110^{11} and Braginsky a part in 101210^{12}, both using the Sun, and modern ground experiments use a rotating turntable to put the signal wherever the apparatus happens to be quietest.

What a null instrument buys

The word “null” is doing a lot of work above, and it is worth unpacking because the principle is general.

An instrument that measures a quantity XX and looks for a small change in it must have a full-scale accuracy comparable to the change: to see one part in 10910^9 of XX, everything about the instrument’s calibration, linearity and drift must be good to that level across its whole range. An instrument that measures the difference between two nearly equal quantities has to be that good only over the difference, which is small — so its calibration can be poor, its linearity irrelevant, and its zero can drift as long as the drift is slower than the modulation.

A torsion balance is the extreme case. Its fibre never twists by more than a microradian, and everything about the instrument is characterised at that amplitude. Its absolute calibration matters only in converting a final null into a limit, and a twenty per cent error there costs twenty per cent of a limit that is being quoted to one significant figure.

The same reasoning is why almost every high-precision measurement in physics is a comparison rather than a determination. A frequency is compared with a frequency, a mass with a mass, a length with a wavelength — and the quantities that are known best are the ratios, which is a fact about instruments rather than about nature.

Where the last factor came from

Where the last three orders of magnitude came from. Each experiment placed by the differential acceleration it could resolve and by the acceleration it had to work with. The diagonal lines are constant η, since η is the ratio of the two: a point moves down the plot by building a better accelerometer and to the right by finding a bigger source. Eötvös: resolution 8.5e-11 m/s², drive 1.7e-2, η 5e-9; Braginsky: resolution 5.9e-15 m/s², drive 5.9e-3, η 1e-12; Eöt-Wash: resolution 3.4e-15 m/s², drive 1.7e-2, η 2e-13; MICROSCOPE: resolution 1.2e-14 m/s², drive 8.0e+0, η 1e-15. The point worth taking is that the satellite's accelerometer is not much better than the best torsion balance — the two resolutions are within an order of magnitude of one another. What the satellite has is a drive five hundred times larger, because a body in free fall around the Earth is being pulled by the whole of the Earth's field rather than by the thousandth of it that the spin leaves pointing sideways. The improvement was bought by changing the numerator's denominator rather than by improving the instrument, which is worth knowing before proposing to improve the instrument.
Fig. 5 Each experiment placed by the differential acceleration it can resolve and by the acceleration it has to work with, with lines of constant η running diagonally. A point moves down by building a better accelerometer and to the right by finding a bigger source.

The satellite went right, not down. Its accelerometer resolves a differential acceleration within an order of magnitude of what the best torsion balance manages; what it has is a drive of eight metres per second squared instead of seventeen millimetres, because a body in orbit is in free fall in the whole of the Earth’s field rather than in the small horizontal residue the spin leaves behind.

There is a second advantage to orbit that the figure does not draw, and it is nearly as large. On the ground, a torsion balance is suspended, and the suspension has to hold the whole weight of the masses while permitting a microradian of twist. In orbit nothing has to be held up at all, so the only forces on the test masses are the ones the experiment applies deliberately — which removes the fibre, its thermal noise, its creep, and the whole family of problems that comes with hanging a precision instrument in a gravitational field. The same argument is what makes a drag-free satellite the natural home for any experiment whose signal is an acceleration, and it is the reason gravitational-wave detection is moving the same way, as what the instrument actually hears sets out for a different measurement.

MICROSCOPE flew two pairs of concentric cylindrical test masses — one pair of platinum and titanium, the other of platinum and platinum as a control — each held at the centre of the satellite by electrostatic forces. The force needed to keep the two coaxial is the measurement: if they fell differently, one would drift and the servo would have to push. The satellite rotated to put that signal at a chosen frequency, well away from the orbital period and from anything thermal. The final limit is about a part in 101510^{15}.

The lesson is one about instrument design rather than about gravity. When a measurement is a ratio, look at the denominator before improving the numerator — and here the denominator had been fixed at a thousandth of gravity by the accident of doing the experiment on a rotating planet.

The one systematic that will not go away

Every experiment of this kind ends up limited by the same thing, and naming it explains a good deal of the design of all of them.

A composition test compares two masses that are not at the same place. Whatever they are made of, they sit at slightly different positions, and the gravitational field is not uniform: a gradient across the separation produces a differential acceleration that has nothing to do with composition. Worse, the gradient comes partly from the apparatus itself — the vacuum chamber, the pump, the experimenter — and it varies as those move.

The countermeasures are geometric. A torsion balance is built so that the two masses sit at the same distance from the axis and the beam is balanced against every low-order moment of the field, not merely the first. MICROSCOPE’s test masses are concentric cylinders sharing a centre, which cancels the gradient exactly to first order and is the reason for a shape that is otherwise inconvenient. And the whole thing is rotated, so a gradient effect fixed in the laboratory appears at a different frequency from the signal.

Even so, gravity gradients are the leading systematic in every one of these measurements, and the published limits are dominated by how well they were modelled rather than by how quiet the instrument was. That the field the experiment is testing is also its worst noise source is not a coincidence — it is a property of any experiment sensitive enough to be worth doing.

What a null result excludes

A sequence of zeros is not a sequence of failures, and it is worth being specific about what each one buys.

Any long-range force that couples to something other than mass-energy — to the number of baryons, say, or to the number of neutrons minus protons — would produce a composition-dependent acceleration, because different materials have different ratios of those to their mass. Platinum and titanium differ in neutron fraction by a few per cent, so a limit on η is a limit on the strength of such a force divided by gravity’s.

Many extensions of the standard model contain a light scalar that would do exactly this, and the limits from these experiments are the strongest constraints on several of them — stronger than anything an accelerator provides, because the effect grows with the number of particles rather than with the energy.

So the experiment is a search for new physics, conducted by looking as hard as possible for nothing. The comparison mass is chosen to maximise the contrast in whatever charge the hypothetical force might couple to, which is why the pairs are exotic: beryllium and titanium, platinum and titanium, and in one experiment a test mass made of material from a meteorite.

The reanalysis that nearly found something

In 1986 a group re-examined Eötvös’s published data and reported that the residuals correlated with the ratio of baryon number to mass across his sample of materials — a signature of exactly the kind of fifth force the limits are meant to exclude. The claimed effect was at the level of a part in 101110^{11} of gravity with a range of a few hundred metres.

It set off a decade of experiments. New torsion balances were built, mine shafts and television towers were used as sources, submarines measured gravity at depth. Nothing was found, and the reanalysis is now generally attributed to a systematic in the original data rather than to physics — most plausibly a gravity gradient from the building the balance was in, which correlated with the sample materials because of the order they were measured in.

The episode is worth remembering for two reasons. It is the clearest case in this subject of a real effect being claimed from a real dataset by a careful reanalysis, and being wrong; and the response to it — build the apparatus again, differently, in several places — is exactly the response that a claim of this kind should get. The limits improved by two orders of magnitude in the process.

Where the model stops

The principle is not a theorem. Nothing in mechanics requires that the mass appearing in F=maF = ma and the mass appearing in the gravitational force be the same quantity; that they are is an experimental fact, and it is the fact that lets gravity be described as geometry. Every step from the floor that cannot be told from gravity onwards depends on it, which is why the limit is worth pushing.

Only the weak principle is tested here. That bodies fall alike says nothing directly about whether a body’s own gravitational binding energy falls the same way, which is the strong equivalence principle and is tested by lunar laser ranging — the Earth and the Moon have different fractions of their mass in gravitational binding, so a violation would distort the Moon’s orbit.

The test masses are macroscopic and unpolarised. A separate family of experiments asks whether a spinning body falls differently, or whether antimatter does, and neither is covered by anything here. The antimatter question was open until very recently and is now answered at the per-cent level, which is where the composition tests were in 1687.

The satellite’s limit is systematics-limited, not statistics-limited. More flight time would not have improved it; the residual is a thermal effect on the test masses whose size had to be modelled rather than measured, which is the usual condition of a mature precision experiment.

And η is not the only parameter. A violation could depend on the distance to the source as well as on the composition, so a limit obtained with the Earth as the source does not directly bound a force with a range of a few metres. Those are constrained by their own, much weaker, short-range experiments.

What the pictures cannot show

The history figure draws each limit as a point, and a limit is a statement with a confidence level and a set of assumed systematic corrections behind it. Several of the points involved reanalyses that moved them, and one of the historical results — Eötvös’s own, reanalysed in 1986 — was claimed for a while to show a positive signal correlated with baryon number. That claim did not survive, but it is the reason the field spent a decade doing the experiment again with new apparatus.

The budget figure treats each experiment as a single resolution and a single drive, and both are averages over a measurement whose noise is a spectrum. What matters in practice is the noise at the modulation frequency, which is why the choice of that frequency — the rotation rate of a turntable, the spin of a satellite — is as much of the design as the accelerometer.

Where the ladder goes next

The equivalence-principle ladder began with the floor that cannot be told from gravity, which states the principle and uses it, and continued to the term free fall cannot remove, where the tidal difference survives every choice of frame and is what curvature actually is. This rung asks how well the first of those is known, and finds a number with fifteen digits of zero behind it.

The rung after it is the strong principle: whether a body’s own gravitational energy falls at the same rate as the rest of it, which cannot be tested with laboratory masses because their binding energy is a part in 102510^{25} of their mass, and which is measured instead by watching the Moon. The habit worth carrying is the one this rung is built on: when a quantity has been measured to be zero for three hundred years, the interesting question is what each new digit rules out — because a limit with nothing behind it is a limit nobody should have paid for.

Part 3 of 4

This essay is one argument about Equivalence principle. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Centrifugal forceEquivalence principleFree fallGravitational massInertial massMeasurementModulationNull experimentPendulumPrecisionSystematic errorTorsion balance