The bend Newton got half right
Assumes: The floor that cannot be told from gravity · The bend at the boundary, and what it is really about
Two theories, one prediction each, and the answers differ by a factor of exactly two. That is an unusually clean situation in physics, where disagreements are more often about the third decimal place, and it is worth understanding why the gap is so wide and so precise.
The lower prediction is old. Newton himself asked, in the Opticks, whether bodies act upon light at a distance and bend its rays. Cavendish worked out an answer in the 1780s and did not publish it; Johann Georg von Soldner published one in 1801. The upper prediction is Einstein’s, from 1915, and it replaced his own earlier answer, which was the lower one.
The half that comes from falling
Start with the sealed box, because the equivalence principle supplies a complete derivation of the first half with no field equations in it.
The Newtonian half fits in one picture. Light crossing an accelerating box lands lower than it left by , and the equivalence principle says a box at rest in a gravitational field must behave the same way — so light falls. That argument is complete, it needs no field equations, and Einstein had it in 1911. It gives exactly half the observed deflection.
A photon passing a mass at impact parameter spends roughly a time in the region where the transverse gravitational acceleration is of order . Multiply: the transverse velocity picked up is about , and dividing by gives a deflection of order . Doing the integral properly — the transverse acceleration integrated along a straight-line path — gives
which for the Sun’s limb is 0.87 arcseconds. This is the answer Einstein published in 1911, four years before the complete theory, and it is the answer the equivalence principle can reach on its own.
The integral is worth seeing once, because the estimate above lands on the exact coefficient and that is not luck. Take the ray along the axis at height , so that . The transverse component of the acceleration at each point is times , and integrating along the whole path,
with the factor of two arriving from the integral rather than from any physical argument. Dividing by gives the angle. The whole deflection is accumulated within a region a few across, which is why the result depends on the closest approach and not on how long the ray has been travelling.
Notice what has been assumed. The photon has been treated as a small body moving at speed through a flat space, subject to a gravitational acceleration. The acceleration is legitimate; the flat space is the mistake.
There is also a curiosity in the Newtonian version that is usually passed over. A corpuscle theory has no business predicting a deflection independent of the corpuscle’s mass — but it does, for the same reason everything falls at the same rate, and it also has no business assuming the corpuscle travels at , since in Newtonian mechanics a body passing a mass speeds up. Soldner’s 1801 calculation includes that speeding up, and it is the reason his published number is not quite what the modern textbook version of the Newtonian answer gives.
The half that comes from space
In general relativity the metric outside a spherical mass differs from flat space in two places at once: the time part is stretched by and the radial space part by its reciprocal. The first is what makes a clock run slow; the second has no Newtonian analogue at all.
The other half comes from space rather than from falling, and the two factors are worth separating. A clock’s rate falls as and a radial ruler stretches as the reciprocal. Newtonian gravity is entirely the first of those — the gravitational potential is a statement about time — and the second has no Newtonian counterpart at all. The 1911 argument used the first and knew nothing of the second, which is precisely why it came out at half.
The size of each contribution depends on speed. A body moving at spends its “motion budget” mostly on the time direction and a little on the space directions, in the ratio ; a photon spends it all on space. Carrying the calculation through, the space term contributes to the deflection in proportion to relative to the time term. For a planet at 30 km/s that is one part in and utterly negligible. For a photon it is one, and the two terms are equal.
So the factor of two is a statement about which parts of the geometry a fast thing samples. Anything slow enough to have been part of Newtonian astronomy could not have revealed the discrepancy — which is exactly why it took light to find it.
The version that reads as optics
There is a second route to the relativistic answer, and for anybody who has met Snell’s law it is the more intuitive one: treat the region round the mass as a medium with a refractive index.
Light in a medium of varying index bends toward the region of higher index — exactly the behaviour a ray shows crossing into glass — and running the ray-tracing integral with the above reproduces precisely.
Read as optics it is ordinary. A ray crossing into a slower medium bends toward the normal by an amount fixed by the ratio of speeds, and the gravitational case is that with the discontinuity smoothed into a gradient — an effective refractive index that rises towards the mass. The analogy is exact enough to be computed with, and it is how gravitational lensing is usually simulated.
The refractive-index picture is a genuine calculational device rather than an analogy, and it is used to trace rays through mass distributions numerically. But it has a hard boundary. An index describes light slowed by a medium, and light is not slowed here: it travels at locally everywhere, as it must. What the index encodes is a coordinate speed — how much coordinate distance is covered per unit coordinate time — and coordinates are a bookkeeping choice. Push the analogy and it will claim things about a medium that no measurement can support.
Why it is a lens with no focus
If a mass deflects light, it images. What it does not do is image the way any manufactured optic does, and the difference is one line of algebra.
A glass lens is shaped so that the deflection grows in proportion to the distance from the axis, which is what puts every parallel ray through one point.
An ordinary lens deflects in proportion to height above the axis, which is what puts every ray through one focus and gives an image plane. A mass deflects as instead — more strongly close in — so there is no focus and no image plane, only a caustic surface where rays happen to pile up. That is why a gravitational lens produces arcs and multiple images rather than a picture, and why the word “lens” is a courtesy.
The consequences follow immediately, and they are the reason the subject looks nothing like ordinary optics. A point source directly behind a point mass is imaged not as a point but as a ring, because every azimuth around the mass gives an equally valid path. Move the source off-axis and the ring breaks into two images of unequal brightness on opposite sides. Extended masses that are not spherical produce arcs, multiple images and, in the right circumstances, an odd number of them.
The scale of the ring is worth one line of algebra, because it shows how the geometry converts an angle into a size. A source at distance , a mass at , and a deflection give a ring of angular radius
in which the deflection’s has become a square root — because the ring’s own radius is what sets the impact parameter, so the geometry has to solve for itself. Everything about the strangeness of the subject is in that self-consistency: the deflection depends on where the ray passes, and where the ray passes depends on the deflection.
Where a mass does act as a mass — as a device for weighing what cannot be weighed any other way — is a subject with a very large literature, and it belongs to the collection that measures things in the sky rather than to this one. What is being argued here is only the physics of a single deflection: an angle, computed two ways, differing by two.
What it costs, and where the model stops
The formula is a weak-field, small-angle result. Both derivations assume the ray travels in a nearly straight line and picks up a small transverse velocity, so the deflection is calculated along the unperturbed path. Close to a compact object that fails: at the exact treatment gives a photon that orbits, and there is no expansion in which that is a small correction.
The impact parameter is not the distance of closest approach. In the exact treatment the two differ, and papers that quote one when they mean the other differ by terms of the same order as the effects being measured.
A real ray does not pass a vacuum. The Sun’s corona is a plasma, and a plasma has a genuine, frequency-dependent refractive index — so a measurement at radio wavelengths carries a coronal bending that must be removed, and the removal is done by measuring at two frequencies and exploiting the fact that gravitational deflection is achromatic while plasma deflection is not. That achromaticity is the sharpest observational signature the effect has.
Nothing in the derivations distinguishes light from anything else fast. A neutrino at nearly is deflected by the same , and so is a proton in a cosmic-ray beam at high enough energy. The formula is about speed, not about being electromagnetic.
The measurement, and how thin the margin was
The 1919 eclipse expeditions are the most retold episode in the history of this subject, and the retelling is usually wrong in one respect worth correcting: the question was never whether starlight bends. It was whether the bend was 0.87 or 1.75 arcseconds, a distinction of under an arcsecond in photographs taken through cloud, with instruments that had been shipped to Príncipe and Sobral and reassembled in the field, on plates whose scale depended on the temperature of the telescope.
Eddington’s Príncipe plates were poor and few; the Sobral astrographic plates showed evidence that the telescope’s focus had shifted; the Sobral four-inch plates were good and gave . The published conclusion favoured the relativistic value, and the analysis has been argued about ever since — the fair modern summary is that the data supported the larger prediction but not as decisively as the announcement suggested.
The difficulty was not conceptual, it was photographic. A star’s image on a plate had to be measured against the same star’s image on a plate taken months earlier, when the Sun was elsewhere in the sky, and the two plates had to share a scale to a part in a hundred thousand. Everything conspires against that: the telescope’s mirror is a different temperature in Príncipe in May than in Oxford in January, the plate emulsion shrinks, and the field of stars near a totally eclipsed Sun is not the field anybody would have chosen. Half the labour of the expeditions was in the scale factor rather than in the positions.
What settled it was repetition with better instruments and, eventually, a different technique entirely: radio interferometry, which observes quasars passing near the Sun without needing an eclipse at all, and which has confirmed the coefficient to a few parts in . The deflection is now among the best-tested predictions in gravitation, and the modern test is quoted not as an angle but as a dimensionless parameter measuring how much space curvature a theory produces per unit mass — which is precisely the half that Newton’s version is missing.
The expeditions that failed, and the escape that followed
The 1919 result is remembered because it succeeded. Three earlier attempts did not, and what they would have measured is the most instructive fact in the whole history of this prediction.
Einstein published the equivalence-principle calculation in 1911 and immediately began urging astronomers to test it. The value he gave them was the lower one — 0.83 arcseconds, the number this essay attributes to falling alone — because the complete theory did not exist yet and would not for another four years.
Three expeditions went out on the strength of it. An Argentine party observed the eclipse of October 1912 from Brazil and was rained out completely. A German party led by Freundlich reached the Crimea for the eclipse of 21 August 1914, and war was declared while they were setting up; the astronomers were interned by Russia as enemy aliens, their instruments confiscated, and they were released some weeks later in a prisoner exchange, having seen nothing. An American attempt at the 1918 eclipse in Washington state produced plates whose analysis was never brought to a firm conclusion.
Einstein corrected the prediction to 1.75 arcseconds in November 1915, when the field equations gave him the space-curvature term the equivalence principle could not see.
So every failed expedition was sent to measure a prediction that was wrong by a factor of two, by the man who had made it, and each of them would — had the weather and the war allowed — have returned a measurement disagreeing with the published theory at roughly twice its stated value. The consequences of that are not knowable, and the usual speculation is worth resisting; what is knowable is that the theory would have entered its most difficult years already carrying a refuted prediction, and that the refutation would have been of the right theory’s provisional half.
The episode is worth carrying for what it says about the status of a partial derivation. The 1911 calculation was not sloppy and it was not wrong in its own terms — it correctly computes what the equivalence principle implies, and the equivalence principle is exactly true. What it missed is that a local statement about a freely falling box has nothing to say about the curvature of space at large, and the deflection of light samples both. A derivation that is rigorous within its assumptions can still be half an answer, and nothing inside the derivation announces which half.
The focus that starts at 550 astronomical units
The absence of a focal plane was stated above as a curiosity of the geometry. It has a concrete consequence: the Sun is a lens, its focus can be located, and the number is worth working out because it explains both why the idea is tempting and why nobody has used it.
Rays grazing the limb are deflected by 1.75 arcseconds, which is radians. They cross the axis at a distance equal to the solar radius divided by that angle — metres, or about 548 astronomical units. Rays passing further out bend less and cross further away, so the focus is a line beginning at 548 AU and running outward without end.
The gain available there is extraordinary. A lens concentrates light in proportion to the area it collects divided by the area it concentrates into, and for a point source seen exactly on axis the concentration is limited only by diffraction. The figure that comes out at optical wavelengths is of order a hundred thousand million, which would make an ordinary telescope placed at that focus capable of resolving surface features on a planet around another star.
Four things stand between that arithmetic and an instrument, and all four follow from the geometry rather than from engineering.
There is no focal plane, so the instrument does not sit at a place and look at a field. It sits on the focal line of one target and sees only that target; observing a different star means moving to a different point on a sphere 548 AU in radius, which is not a manoeuvre.
The gain applies to a source exactly on axis, and the image is a ring. An extended source maps to a set of overlapping rings, so recovering a picture means deconvolving a point-spread function that is an annulus — which is possible and is not the same as taking a photograph.
The light passes through the corona, which is a plasma with a refractive index of its own, varying and frequency-dependent. That is the noise source, and it is the reason such a mission would work best well beyond the minimum distance, where the rays pass further from the Sun.
And 548 AU is fourteen times the distance to Neptune. Voyager 1, the most distant object ever launched, has taken forty-eight years to reach about a third of it.
The idea is nonetheless taken seriously enough to be studied, and what makes it worth mentioning here is what it demonstrates about the law. Every strange feature of the proposal — the line instead of a plane, the ring instead of a point, the one target per position — is a direct reading of the fact that a mass bends the closest rays most, which is the opposite of what glass does.
The number the modern test actually quotes
Nobody now reports a deflection in arcseconds. The result is quoted as a single dimensionless parameter, and understanding what it is makes the factor of two sharper than any measurement of an angle.
Write a general metric around a mass with an adjustable amount of space curvature per unit mass, and call that amount . Newtonian gravity is : all of the effect in the time part, none in the space part. General relativity is . The deflection then comes out as
which is the whole of this essay in one line: the first term is falling, the second is space curvature, and the ratio of the two predictions is .
That parametrisation is what makes the experiment a measurement rather than a yes-or-no test, and it is why the modern numbers are so impressive. Radio interferometry of quasars near the Sun and the Cassini spacecraft’s radio tracking give to within about . There is no gap left in which a theory could hide half a factor; what is being constrained now is a departure in the fifth decimal place.
The same appears in the Shapiro delay and in the precession of an orbit, which is the point of writing it that way: one number, measured in three unrelated experiments, all of which have to agree.
The ladder from here
Later rungs on this anchor: the Shapiro delay, which is the same metric read as a time of flight and is measured to a part in by bouncing radar off a planet; the exact deflection near a compact object, where the weak-field formula fails and photon orbits exist; deflection as an achromatic effect and how that separates it from plasma; the bending of massive particles, where the coefficient runs continuously with speed from the Newtonian value to twice it; and the geometry of multiple imaging, which is what a lens with no focal plane produces.
The neighbouring ladders are refraction, whose Fermat-style argument reappears here almost word for word, and the horizon, which is where the small-angle expansion on this page finally breaks and the deflection becomes a capture.
Part 1 of 3
This essay is one argument about Light deflection. The others:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
Deflection angleEquivalence principleFermat's principleGeodesicImpact parameterRefractive indexSpacetime curvature
- The ray that bends without a surface deflection angle, fermat's principle, refractive index
- The horizon that nothing marks equivalence principle, spacetime curvature
- The longest way round is the shortest clock equivalence principle, geodesic
- The path that does not change fermat's principle, refractive index
- The principle that fixes the energy instead of the clock fermat's principle, refractive index