Astrophysics

The lens with no focal length

A glass lens bends a ray by an angle proportional to how far off the axis it passes, which is precisely the condition for every ray to arrive at one point. Gravity bends by an angle that falls with distance off the axis, so every ray has its own focus and there is no image plane anywhere — only a half-line of foci, beginning 548 astronomical units from the Sun and running outward for ever.
18 min read 6 figures The shape decidesFields, not forces

Assumes: The bend Newton got half right · What a lens is doing, and why three rays are enough

Light passing the edge of the Sun is deflected by 1.75 arcseconds. That is a bend, and a bend brings rays together, so somewhere behind the Sun there is a place where the light from a distant star is concentrated. There is — but it is not a focus, it is not a plane, and calling the arrangement a lens gets almost everything about it wrong.

A lens with no focal length. Where a ray crosses the axis, against how far off the axis it passed the deflecting body, for 1 solar mass of radius 1 solar radius. A glass lens deflects a ray by an angle proportional to its distance off axis, which is precisely the condition for every ray to arrive at one point — the flat dashed line. Gravity deflects by 4GM/c²b, which grows smaller further out, so the crossing distance goes as b² and each ray has its own focus. The grazing ray crosses at 548 astronomical units and a ray passing at 12 radii crosses at 78857; the square law is verified on the drawn curve to 1.5e-16. So there is no image plane at all, only a half-line of foci beginning at the first of those and running outward for ever. Anything placed on that line sees not an image but a ring, and moving along it does not refocus anything — it selects which rays are being seen.
Fig. 1 Where a ray crosses the axis, against how far off the axis it passed the Sun. A glass lens would give the flat dashed line — every ray arriving at one place, which is what a focal plane means. Gravity gives a curve rising as the square of the impact parameter: the grazing ray crosses at 548 astronomical units and a ray at twelve solar radii at 78,857.

What a lens is doing that gravity is not

The condition for a focus is a specific one and it is easy to state.

Three construction rays locate an image, and the reason they meet is that a lens deflects each of them by an angle proportional to its height above the axis. That proportionality is the whole of what makes a focus: rays further out are bent more, in exactly the ratio needed to bring them to the same place. A mass does not do that, and everything in this essay follows from the one difference.

That proportionality is the whole of imaging. The thin-lens law is nothing more than its consequence.

Every feature of the thin-lens relation — the divergence at the focal length, the sign change beyond it, the asymptote at large distances — depends on there being a focal length. A gravitational lens has none, so none of those features has a counterpart: there is no distance at which the image runs to infinity, no boundary between real and virtual, and no plane where an image lives.

Gravity’s deflection is

α=4GMc2b,\alpha = \frac{4GM}{c^2 b},

which falls as the impact parameter bb grows. A ray twice as far out is bent half as hard, so it crosses the axis four times further away.

Two predictions, a factor of two apart. The deflection of light passing a mass, against impact parameter, on logarithmic axes. The lower line is what a Newtonian photon does — it falls while it crosses, and comes out bent by 2GM/bc². The upper line is what a geodesic does in curved spacetime, which is exactly twice that. At the surface of a body of 1.99·10³⁰ kg the two are 0.88″ and 1.75″. Both are straight lines of slope minus one, so the ratio is two everywhere and the measurement is a choice between two theories rather than a fit.
Fig. 2 The deflection against impact parameter, with the Newtonian half-value beneath. Newton’s calculation gives exactly half and both fall as 1/b, so the qualitative statement in this essay was already true of the wrong theory: the failure to focus is a property of the inverse law rather than of general relativity.

The half-line of foci

Setting the crossing distance D=b/α=b2c2/4GMD = b/\alpha = b^2c^2/4GM and putting in the Sun’s mass gives 548 astronomical units for a ray that grazes the surface. Rays passing further out cross further out, without limit.

Rays passing a mass. Light passing a body of 1.99·10³⁰ kg and radius 696,000 km at four impact parameters, deflected by 4GM/bc². At the surface that is 1.75 arcseconds and it falls off as 1/b, so the ray passing at five radii bends by 0.35. The angles are drawn 2.6·10⁴ times their true size; at the true size every ray on this canvas would be straight to within a hundredth of a pixel.
Fig. 3 Rays traced past the body at a range of impact parameters. They do not converge on a point; the innermost cross first and the outermost last, and the locus of crossings is a line along the axis rather than a plane across it. Nothing has been drawn imprecisely — this is what the deflection law produces.

The practical statement is that an observer anywhere on that line beyond 548 AU sees a ring rather than a point, and moving further out along the line does not sharpen the ring. It selects a different impact parameter — a different annulus of the Sun’s limb — and therefore a different ring.

The gain is real and enormous. A telescope at 550 AU looking back at the Sun would see light from a source directly behind it concentrated into a thin annulus with a brightness gain of order 101110^{11}, which is the basis of the proposals for a solar gravitational lens mission. The gain is a brightness gain and not an increase in resolution, and no arrangement of optics can convert one into the other.

What is at 548 astronomical units, and what is not

The number deserves scrutiny, because it is often quoted as though it named a destination.

It is the crossing distance of a ray that grazes the Sun’s photosphere. Rays that graze closer do not exist, because there is a Sun in the way; rays that graze further out cross further out. So 548 AU is the beginning of the focal line and not a point on it, and a spacecraft at 550 AU is at the inner end of a region that extends indefinitely.

Two things spoil the inner end. The corona is a plasma, and a plasma refracts radio waves in the opposite sense to gravity and with a strong frequency dependence, so the innermost rays are corrupted for anything but the shortest wavelengths. And the Sun is not spherical: its rotation makes it very slightly oblate, and the quadrupole term in its field adds a deflection that varies round the limb.

A gravitational deflection falls as 1/b1/b rather than 1/b21/b^2, one power more slowly than a field does — because it is an integral of the field along the whole path rather than the field at a point. That single exponent is what prevents a focus: a lens needs deflection rising with height, and this one falls.

The distance is also very large. Voyager 1, the most distant object ever launched, is at about 165 AU after nearly fifty years, so reaching the inner end of the focal line is a mission of a century at current speeds. Nothing about that is a physical obstacle and all of it is why the arrangement is discussed rather than used.

Two images, always

For a source not exactly behind the lens, the geometry gives an equation rather than a construction. With angles measured in units of the Einstein radius, the lens equation is β=θ1/θ\beta = \theta - 1/\theta, a quadratic with two roots.

Two images, always, and a ring when they merge. The positions of the two images a point mass makes, and their total brightness, against how far the source lies from perfect alignment — all in units of the Einstein radius. The lens equation β = θ − θE²/θ is a quadratic in θ, so there are exactly two solutions and never one or three: one image outside the Einstein radius and one inside it, on the opposite side. Their positions multiply to −1 at every alignment, checked here to 6.7e-16, so knowing one gives the other, and their magnifications differ by exactly one at every alignment, to 1.8e-15. Reading off: at 0.1 Einstein radii off, the images sit at 1.05 and -0.95 and the total brightness is 10.04 times the unlensed source; at 0.3 Einstein radii off, the images sit at 1.16 and -0.86 and the total brightness is 3.44 times the unlensed source; at 1 Einstein radius off, the images sit at 1.62 and -0.62 and the total brightness is 1.34 times the unlensed source; at 2 Einstein radii off, the images sit at 2.41 and -0.41 and the total brightness is 1.06 times the unlensed source; at 3 Einstein radii off, the images sit at 3.30 and -0.30 and the total brightness is 1.02 times the unlensed source. The total is greater than one for every alignment — lensing never dims anything — and it runs away at perfect alignment, where the two images become a ring. That divergence is the model's and not the world's: a source of any finite size averages the magnification over its own face, and the answer is large and finite.
Fig. 4 The two image positions and the total brightness, against how far the source lies from alignment. There are exactly two solutions and never one or three: one image outside the Einstein radius and one inside it, on the opposite side, with positions multiplying to −1 at every alignment to a part in 10¹⁵.

Three features are worth reading off directly.

The product is fixed. θ+θ=1\theta_+\theta_- = -1, so measuring one image gives the other, and the geometric mean of the two image separations is the Einstein radius — which is a mass measurement, since the Einstein radius depends on the mass and the distances and nothing else.

Lensing never dims. The total magnification is greater than one at every alignment, approaching one only as the source moves far from the lens. That is a consequence of the deflection being an attraction, and it is why a survey looking for lensed objects looks for brightenings.

And the magnification runs away at alignment. Two rays arriving from opposite sides of the lens are then joined by a whole circle of them, and an infinite ring of contributions arriving at one point is what the divergence is counting. The point-mass model gives an infinite total brightness for a perfectly aligned point source. That divergence is the model announcing that it has been used past its range, and the two things that cure it are the two things every real system has: a source with a finite angular size and a lens with internal structure.

An odd number of images

The claim that there are always exactly two images is true of a point mass and is the exception rather than the rule, and the general statement is more elegant.

For any lens whose mass is spread out smoothly — no singular points, finite density everywhere — the number of images of a background point source is always odd. The result is topological rather than computational: the lens map from source plane to image plane is continuous, and counting how many times it covers a given point, with signs for orientation, gives an odd total.

So a galaxy lensing a quasar produces three images or five, never two or four. What is observed is usually two or four, and the missing one is always the same one: a central image, formed by light passing close to the lens’s centre, demagnified enormously by the steep central density and often hidden behind the lens galaxy itself.

The point-mass case is where the odd number appears to fail, and the reason is instructive. A point mass is infinitely dense at its centre, so the central image is demagnified not merely severely but to exactly zero and pushed to exactly the centre. The third image exists in the counting and has no brightness, which is what “the point-mass model gives two images” means.

That makes the central image a diagnostic. Its brightness depends on how steeply the lens’s density rises toward the centre — a shallower core gives a brighter central image — so detecting one, or failing to detect one at a stated sensitivity, constrains the innermost mass distribution of a galaxy in a way nothing else does. Several have now been found, at brightnesses of a fraction of a per cent of the outer images.

The curve that is the whole observation

For the stellar case, where the images can never be separated, everything observable is a single curve, and its shape is fixed.

As a lens drifts across the line of sight to a background star, the alignment improves and then worsens, and the total brightness rises and falls. The shape of that rise and fall is determined entirely by the geometry: it depends on how close the alignment gets and on how long the crossing takes, and on nothing else at all. It is symmetric in time, and it is the same at every wavelength — which is the signature that distinguishes it from every kind of intrinsic stellar variability, all of which is asymmetric, coloured, or both.

The timescale carries the physics. It is the Einstein radius divided by the relative angular speed of lens and source, so it combines the lens’s mass, the distances and the transverse velocity into one number. Toward the centre of the Galaxy, for ordinary stellar lenses, it comes out at a few tens of days — which is why such surveys monitor millions of stars nightly for years and detect a few thousand events.

And a companion to the lens breaks the symmetry in a way that is the whole point of the technique. A planet orbiting the lensing star adds its own small Einstein radius, and if a source image happens to sweep past it the light curve acquires a brief anomaly on top of the smooth one. The anomaly’s duration is the main event’s timescale multiplied by the square root of the mass ratio — so a Jupiter gives a feature lasting a day and an Earth a few hours, on a curve that spans a month.

Which makes the method sensitive to a range of planets nothing else reaches: cold, low-mass, far from their stars, and detected by their mass alone rather than by any light they emit or block. The cost is that each event happens once and cannot be revisited, so a candidate has to be recognised and followed while it is happening.

Why a magnification is a brightness and not a picture

The insistence that lensing magnification is not resolution deserves its reason, because the reason is a theorem rather than a practical limitation.

Gravitational deflection conserves surface brightness. A bundle of rays leaving a source with a certain brightness per unit solid angle arrives with the same brightness per unit solid angle, because nothing along the way emits or absorbs and the deflection is a change of direction rather than of energy per mode. That is the same conservation that makes it impossible for any passive optical system to make an image brighter per unit area than its source.

What lensing changes is the solid angle the source subtends. A magnified image is a source spread over a larger patch of sky at the same surface brightness, so it collects more photons in a telescope — which is what an increase in total flux means, and which is genuinely valuable when the source is faint.

But the resolution is untouched. Two features of the source separated by some angle are magnified to a larger separation, which does help; and the telescope’s own resolution limit is unchanged, so what is gained is exactly the magnification factor and no more, in a distorted geometry that has to be modelled before anything can be read. Lensing is a light bucket that also stretches the picture, and calling the stretch a magnification invites the reading that detail has been created where none was resolved.

The distinction matters most for the solar-lens proposals. A gain of 101110^{11} in collected light is enormous and real; the accompanying claim that it would resolve continents on a distant planet needs the stretch rather than the gain, and the stretch is heavily distorted and applies only along one direction at a time.

Why the failure to focus is not a defect

There is a temptation to file all of this under aberration — a real lens does not focus perfectly either, and this is a lens that focuses very badly.

Spherical aberration is a defect — the surface has the wrong shape and a parabola cures it. The failure to focus here is not of that kind: no shape of mass distribution produces a deflection proportional to impact parameter, because the deflection is fixed by the mass enclosed and the geometry. There is nothing to correct, which is why “aberration” is the wrong word for it.

The distinction matters because it decides what can be improved. A spherical mirror can be re-figured. A gravitational lens cannot be, because the deflection is set by the mass distribution and the geometry of spacetime around it.

The deflection formula used throughout is the weak-field one, valid where the impact parameter is far larger than the Schwarzschild radius — which for every body in this essay it is, by many orders of magnitude. The strong-field case is a different calculation with a different answer, and it is where the photon sphere and the shadow live rather than anywhere a lens is being discussed.

The other thing gravity does to light on the way past

Deflection is not the only effect, and separating the two is a standing requirement.

A radar echo past the Sun, delayed by 233 microseconds. The extra time a round-trip radar signal takes when its path passes close to the Sun, against how close, for a reflector 0.723 AU away. Grazing the Sun's limb the delay is 233 microseconds — about 70 kilometres of light travel, on a path of hundreds of millions — and it falls only as the logarithm of the impact parameter, so the effect is still tens of microseconds ten solar radii out. That slow falloff is what makes the measurement possible: the delay can be watched building and fading as the geometry changes, rather than having to be caught at one instant. The dashed curves are the delay the solar corona's plasma adds at three radio frequencies. It is the competing effect, it is larger than the gravitational one close in, and it falls as the square of the frequency while the gravitational delay does not depend on frequency at all — which is how the two are separated, and why the sharpest measurement of this was made with a spacecraft carrying three radio links instead of one.
Fig. 5 The extra travel time for a signal passing near a mass — a delay rather than a bend, and measurable independently. In a lensing system the two images correspond to paths of different length and different depth in the potential, so a source that varies is seen to vary at different times in its different images.

That time delay is one of the few clean routes to a distance in cosmology, because it depends on the geometry of the whole path rather than on any property of the source, and measuring it in a variable lensed object gives a length scale directly.

The numbers for a lens that is not the Sun

The Einstein radius is the natural scale of every lensing configuration, and its size decides which observations are possible.

θE=4GMc2DlsDlDs.\theta_E = \sqrt{\frac{4GM}{c^2}\cdot\frac{D_{ls}}{D_l D_s}} .

A star in the Galaxy lensing a star behind it gives about a milliarcsecond for a solar mass at a few kiloparsecs. That is far below any telescope’s resolution, so the two images are never separated and the only observable is the total brightness — which is why the field is called microlensing and why its data are light curves rather than pictures.

A galaxy lensing a quasar gives about an arcsecond for 101110^{11} solar masses at cosmological distances, which is resolvable, and this is the regime that produces the photographs of multiple quasar images.

A cluster gives ten to thirty arcseconds and produces the arcs that are the visual signature of the subject.

Across those three the mass varies by fifteen orders of magnitude and the angle by four, because the angle goes as the square root of the mass. That square root is worth noticing: it means the Einstein radius is a remarkably insensitive probe of mass, and a factor of four error in a mass estimate is a factor of two in an angle.

Resolution is set by a wavelength over an aperture, and a gravitational lens has no aperture in any ordinary sense — so the criterion has to be applied to the arrangement rather than to an instrument. What is being resolved is two images of one source, and whether they can be told apart depends on the geometry of the alignment rather than on anything about the telescope.

What was seen, and when

The observational history is short and unusually decisive.

The deflection was measured in 1919 at the Sun’s limb, which is the rung at the bottom of this ladder. The idea that a star could act as a lens for another star was worked out by Einstein in 1936 — at the prompting of an amateur, and published with a note that the effect had no chance of being observed. He was right about stars and wrong about galaxies, which Zwicky pointed out the following year.

The first lensed object was found in 1979: two quasars a few arcseconds apart with identical spectra and identical redshifts, which is the signature nothing else produces. The first Einstein ring followed in 1988, and the first microlensing events — brightenings of stars in the Magellanic Clouds by unseen masses in the Galactic halo — in 1993.

Where a light-travel delay is picked up. The fraction of the one-way Shapiro delay accumulated by the time the signal has got a given distance from closest approach, for a ray grazing at 1 solar radius, on a logarithmic distance axis. The curve is a straight line over most of its range, which is the whole point: the integrand falls as 1/r, so every factor of ten in distance contributes the same amount, and half of the delay is picked up beyond 22 solar radii — a region where the field is thousands of times weaker than at the limb. The delay is not a local event at closest approach. It is a logarithm, and a logarithm has no scale, which is also why the total depends on where the two endpoints are and not only on how close the path came.
Fig. 6 The coordinate speed of light near a mass, which is what a distant observer computes and is not what anybody measures locally. The bending, the delay and the lensing are all consequences of that variation, and every one of them is a statement about coordinates in a curved geometry rather than about light behaving unusually.

What Einstein got wrong is instructive. He judged the effect unobservable because he was thinking of resolving the images, and every use the subject has since found — microlensing light curves, the statistical distortion of background galaxies, the time delay between quasar images — measures something other than a separation. The instrument turned out to be useful in ways that did not require it to work as a lens, which is the essay’s point restated as a piece of history.

Where the model stops

The weak-field approximation is assumed throughout. Close to a horizon the deflection is not 4GM/c²b at all; it diverges at the photon sphere, where light orbits, and rays passing nearer than that are captured. A body inside its own Schwarzschild radius produces an infinite series of images from rays that have gone round once, twice and more, and none of that is in the expression used here.

The lens is a point. Real lenses are galaxies and clusters with extended mass distributions, and an extended lens has a deflection law that is not 1/b1/b: inside the mass distribution the enclosed mass grows with radius, so the deflection can rise rather than fall, and then a genuine focus and a genuine caustic surface become possible. Everything in this essay is the point-mass case, which is the simplest and the least representative.

Only one deflection has been counted. Light passing through a cluster is deflected repeatedly, and the weak-lensing regime — small distortions of many background galaxies — needs a statistical treatment rather than a ray construction.

The alignment is treated as static. Lens, source and observer all move, so a real lensing event is a light curve rather than a configuration, and the observable is a brightening and fading over days to months rather than a pair of images. Whether the images can be resolved at all depends on the Einstein radius against the telescope’s resolution, and for a stellar-mass lens in the Galaxy it cannot.

And the geometry is Euclidean here and is not in the sky. The distances that enter the Einstein radius are angular diameter distances in an expanding universe, which do not add, so the “distance from lens to source” is not the difference of two distances. Every practical lensing calculation begins by getting that right, and nothing in this essay does.

What the pictures cannot show

The crossing-distance figure plots a single curve, which implies a single ray at each impact parameter. Light arrives from a source in every direction at once, and every annulus of impact parameter contributes; the figure is a slice through a family, and the family is what makes a ring rather than a point.

Nor can any of these figures show the ring. Everything here is drawn in a plane containing the axis, and the whole of the interesting behaviour is what happens when that plane is rotated about the axis — the two images become a circle, the two brightnesses become one brightness, and the caustic point becomes a caustic line. A two-dimensional drawing of an axially symmetric situation shows one meridian and hides the symmetry that is doing the work.

Where this ladder goes next

Three rungs establish, in order: that light is deflected and by twice the Newtonian amount; that it is also delayed, which is a separate and independently measurable effect; and now that the deflection law forbids the one thing the word lens promises. The last is the rung that turns a curiosity about starlight into an instrument, because knowing that the arrangement produces rings and pairs rather than images is what makes the pairs usable.

The habit worth carrying away is to read a law’s exponent before its magnitude. What a deflection law does to an image is decided by how it depends on position, not by how strong it is. A weak deflection proportional to height focuses perfectly; a strong one falling as 1/b does not focus at all. The same reading applies elsewhere in this collection: an inverse-square field’s flux depends only on what is enclosed for the same reason, and a force law with any other exponent would not have that property. The exponent is where the structure lives, and the coefficient only sets the scale.

What is left on this ladder is the extended lens: what happens when the mass is spread out, where the caustics and the odd-numbered image counts come from, and why a real galaxy produces four images of a quasar rather than two.

Part 3 of 3

This essay is one argument about Light deflection. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

CausticEinstein radiusFocal lengthGeneral relativityGravitational lensImagingImpact parameterLight deflectionMagnificationSchwarzschild radius