Optics

The surface every ray is normal to

Light from a point, bent by any number of lenses and mirrors, comes out as a tangle of rays that no longer meet anywhere. Hidden in the tangle is a surface every ray crosses at a right angle: the set of points the light reaches at the same optical time. That a surface always exists is a theorem two centuries old, and it is the reason a modern wavefront sensor works — an array of tiny lenses measures only the direction each patch of light is travelling, and the measurements can be added up into a shape only because they are the slopes of one.

Assumes: The path that does not change · The surface that images one point exactly

The path that does not change reduced reflection, refraction and the rainbow to one condition: light between two points takes a path whose optical length — distance times refractive index, added up along the way — is stationary. The surface that images one point exactly turned that condition into a design rule. For every ray from an object point to reach the same image point, every ray must have the same optical length, and the surface that makes it so is a Cartesian oval.

Both essays read Fermat’s principle one ray at a time. Read across a whole bundle of rays from one source, it says something else, which is less obvious and in practice more useful. Mark, along every ray leaving a point, the place the light has reached after a given optical length. Those marks lie on a surface, and the theorem is that every ray crosses that surface at a right angle — before any lens, after any lens, after any number of reflections and refractions, however badly the optics are made. The surface is the wavefront, and the theorem guarantees that there always is one.

A surface at right angles to everything

In empty space around a point source the statement is obvious: the rays are radii and the surfaces of equal path are spheres centred on the source. After a lens it is not obvious at all. A real lens bends rays at different heights by slightly the wrong amounts, the rays no longer converge to one point, and the picture is a tangle.

Rays and the surfaces of equal optical path. Rays from a point source 20 mm in front of a curved glass surface (radius 4 mm, index 1.5), traced exactly, and the surfaces on which the optical path from the source — distance times index — takes the same value (curves, every 2.5 mm of path). In air the surfaces are spheres round the source. In the glass the rays from the edge of the surface cross the axis at 12.8 mm, well short of the paraxial focus at 20.0 mm, and the equal-path surfaces are no longer spheres. They are still at right angles to every ray: measured on a fan of 641 rays over 10 surfaces in the glass, the largest departure from 90° is 0.005°. Further in, the surfaces fold into cusps where neighbouring rays cross, and there a surface has no single normal.
Fig. 1 Rays from a point source 20 mm in front of a glass surface of radius 4 mm and index 1.5, traced exactly, with the surfaces of equal optical path every 2.5 mm. Edge rays cross the axis at 12.8 mm, far short of the 20 mm paraxial focus. The equal-path surfaces are not spheres in the glass, and every ray still crosses them at a right angle — 90° to within 0.005°, measured on a fan of 641 rays.

The figure traces one curved glass surface with a strong aberration: rays near its edge are bent so much more than rays near the axis that they cross the axis eight millimetres short of where the central rays do. The surfaces of equal optical path, drawn every two and a half millimetres of path, are spheres in the air and something else in the glass — flattened at the edges, where the edge rays are ahead. They are still at right angles to every ray. Measured directly on the traced points, over a fan of 641 rays, the largest departure from a right angle is five thousandths of a degree, which is the precision of the drawing, not a property of the light.

The property fails in one place, and the figure shows where. Further into the glass, neighbouring rays cross each other on a curved envelope, the caustic, and the equal-path surfaces fold over into cusps there. At a cusp a surface has no single normal, and the theorem says nothing. Everywhere else, every ray meets every equal-path surface square on.

Why it must be so

The proof is Fermat’s principle used twice. Call S(r)S(\mathbf r) the optical length of the ray from the source to the point r\mathbf r. Move the end point a small distance along the surface S=S = constant. The new ray’s optical length is the same, by the definition of the surface. But by Fermat’s principle the optical length of the real ray to the new point differs from the length of the old ray, extended by the small step, only at second order — and the extension adds nn times the step’s component along the ray. For the two to agree to first order, that component must vanish. The step along the surface is at right angles to the ray.

The argument used nothing about lenses. It holds in a medium whose index varies smoothly, where there is no surface to refract at and rays bend continuously — the air over a hot road, the graded glass of a fibre — and there too the rays are the normals of the equal-path surfaces. The surfaces simply crowd together where the index is high and the light is slow, and a ray, staying perpendicular to them, turns towards the crowding. Snell’s law, the mirage and the curved rays of a graded-index lens are all the one statement that rays run square across the surfaces of equal optical time.

In the language of calculus, the gradient of SS is nn times the unit vector along the ray:

∇S=n s^.\nabla S = n\,\hat{\mathbf s}.

That is the eikonal equation, and it says that rays are the gradient lines of a single function. Étienne-Louis Malus proved the first case of the theorem in 1808, for a single reflection or refraction, in the course of the work that led him to discover polarisation by reflection; Charles Dupin and others extended it over the following two decades to any number of reflections and refractions. William Rowan Hamilton built his whole optics on the function SS in 1828 — he called it the characteristic function — and then carried it over to mechanics, where the same function becomes the action that knows where every path ends, with momentum in place of ns^n\hat{\mathbf s}.

What an aberration is

With the theorem in hand, the tangle of rays has a one-number description at every point of the beam: how far the real wavefront lies from the ideal one. For a lens that is supposed to focus to a point, the ideal wavefront is a sphere centred on that point, and the wavefront aberration is the distance, in wavelengths, between the real surface and the sphere.

How far the wavefront is from a sphere. The departure of the equal-path surface from a sphere centred on the paraxial focus, in wavelengths of 550 nm, against the height at which the ray met the glass, for the same single surface. It is computed from the traced optical path to the reference sphere. Near the axis it grows as the fourth power of the height — the dashed curve is −1.7 h⁴ waves with h in millimetres — and at the edge of the 2.8 mm aperture it reaches −205 waves, where higher powers have joined in. This single function of position is what a lens designer calls the aberration. It exists only because the equal-path surfaces exist.
Fig. 2 How far the equal-path surface lies from a sphere centred on the paraxial focus, in wavelengths of 550 nm, against the height at which each ray met the glass, from the traced optical paths. It starts as the fourth power of height — dashed, −1.7 h4-1.7\,h^4 waves with hh in mm — and reaches 205 waves at the 2.8 mm edge, where higher powers join in.

For the single surface in the figure the aberration starts as the fourth power of the height at which a ray meets the glass — the signature of spherical aberration, the defect a spherical mirror cannot escape — and reaches two hundred wavelengths at the edge of a three-millimetre aperture. That is a terrible lens, chosen so the effect can be seen. A good camera lens holds the aberration to a fraction of a wavelength over its whole aperture, and how accurate a mirror has to be found the tolerance: a quarter of a wavelength of error, spread over the aperture, costs a fifth of the light in the central spot.

Because the aberration is a single function over the aperture, it can be broken into standard pieces, much as a sound is broken into tones. Ludwig von Seidel did it in 1857 for the lowest order, naming the five aberrations every lens designer still lists — spherical aberration, coma, astigmatism, field curvature and distortion — each a particular way the wavefront can depart from a sphere as the aperture and the angle to the axis grow. The condition a lens must meet to be free of coma is, in this language, a condition on how the wavefront tilts for points just off the axis. Frits Zernike gave the pieces their modern form in 1934, as a set of polynomials over a circle that are independent of one another, so that the wavefront error a sensor measures can be read off as so much of each: so many waves of defocus, so many of astigmatism, so many of coma.

This is why optical designers speak of wavefronts rather than rays. Every ray-tracing program computes rays, because rays are what Snell’s law moves. Every specification is written in wavefront error, because it is a single function over the aperture that decides the image, and because it is additive: the aberrations of two lenses in a row add, surface by surface, in a way that the rays’ scattered crossing points do not. The addition works because each lens maps one wavefront to another, which the theorem guarantees exists.

Measuring the slope

The wavefront is a surface a wavelength or so from flat, and it cannot be seen directly. What can be measured is the direction the light travels, and the theorem says that direction is the wavefront’s normal — so measuring directions is measuring the wavefront’s slope.

Johannes Hartmann did this with a sheet of card in 1900. To test a large telescope objective he covered it with a mask pierced by a grid of small holes and photographed the pencils of light coming through them on either side of the focus. Each pencil’s position revealed the direction of light from that part of the lens, and from the directions he worked out where the lens was wrong. Roland Shack and Ben Platt replaced the holes with an array of tiny lenses around 1971, which wastes no light and focuses each patch of the beam to a spot. The Shack–Hartmann sensor is now the commonest wavefront sensor there is.

What a wavefront sensor sees: one spot per small lens. A Shack–Hartmann sensor: an array of 12 × 12 small lenses across a circular beam, each focusing its patch of the wavefront to a spot. A flat wavefront puts every spot at its lens's centre (crosses); a tilted patch moves its spot in proportion to the local slope (dots, displacement exaggerated). The wavefront here is 1.2 waves of spherical aberration and 0.5 of coma. Integrating the measured slopes across the grid by least squares recovers it with an error of 0.087 waves rms against 0.53 waves rms in the wavefront itself; the remainder is detail finer than one lenslet, and it shrinks as the square of the lenslet size. The integration is legitimate only because the slopes are the gradient of one surface, which is what Malus and Dupin proved rays from a point always are.
Fig. 3 A Shack–Hartmann sensor: 112 small lenses across a circular beam, each focusing its patch to a spot. A flat wavefront would put every spot at its lens’s centre (crosses); the measured spots (dots, displacement exaggerated) are moved by the local slope of a wavefront carrying 1.2 waves of spherical aberration and 0.5 of coma. Rebuilt from the slopes by least squares, the wavefront comes back to 0.087 waves rms against 0.53.

Each lenslet reports two numbers, the shift of its spot across and up, which are the average slope of the wavefront over its patch in the two directions, multiplied by the lenslet’s focal length. That is all the sensor measures. The wavefront itself comes from adding the slopes up across the grid, and the figure does that addition by least squares: of all surfaces, find the one whose slopes best match the measured ones. The rebuilt surface matches the true one to 0.087 waves rms, where the wavefront itself is 0.53 waves rms. The remainder is detail too fine for twelve lenslets across to see, and it shrinks as the square of the lenslet size — to 0.051 waves with sixteen across and 0.022 with twenty-four.

Why adding up the slopes is allowed

From slopes back to the wavefront. A cut through the same wavefront along one diameter. Below: the average slope each of 12 lenslets reports across its own width (bars). Above: the true wavefront (curve) and the running sum of slope times width from the left edge (dots), which lands on the curve at every lenslet boundary. Along a line, adding up slopes always works. Across a whole pupil it works only if the slopes measured round any closed loop add up to nothing — if the slope field has no curl — and that is the condition rays from a single point satisfy.
Fig. 4 A cut through the same wavefront along one diameter. Below: the average slope each of twelve lenslets reports across its width. Above: the true wavefront (curve) and the running sum of slope times width from the left edge (dots), which lands on the curve at every lenslet boundary.

Along a single line, adding up slopes always gives back the curve — the sum of a function’s increments is the function. Across a two-dimensional pupil it need not. There are many paths from one lenslet to another, and if the slopes summed along two different paths gave different answers, there would be no single surface to rebuild, and the least-squares fit would be finding a compromise between contradictions.

The two paths agree exactly when the slope field has no curl: when the slopes summed round any closed loop give zero. A field of gradients always has that property — a walk round a loop on a hillside ends at the height it started from. And the theorem says that the ray directions from a point source are the gradient of the optical path. So the slopes a Shack–Hartmann sensor measures on light from one point are guaranteed to be integrable, and the sensor’s arithmetic is guaranteed to have an answer.

Adding the slopes round a loop. The sum of the measured wavefront slope along a circle round the centre of the beam, in wavelengths, against the circle's radius, for three bundles of rays. For rays from a point source — here with 1.2 waves of spherical aberration and 0.5 of coma — it is zero at every radius: the slopes are the gradient of one surface. For a twisted bundle of skew lines, which no surface is normal to, it grows as the area enclosed, reaching 8.5 waves at the edge; no arrangement of lenses and mirrors fed from one point can produce it. For a vortex beam it is exactly one wavelength at every radius: the wavefront is a helical ramp that climbs one wavelength per turn, normal to every ray but not a single-valued surface, which is why its axis is dark.
Fig. 5 The slope summed round a circle in the beam, in wavelengths, against the circle’s radius. For rays from a point source, aberrated or not, it is zero at every radius. For a twisted bundle of skew lines it grows as the area enclosed, reaching 8.5 waves at the edge. For a vortex beam it is exactly one wavelength at every radius.

The figure sums the slopes round circles of increasing radius for three bundles. For rays from a point, through the same aberrations as before, the sum is zero at every radius to the precision of the arithmetic. For a bundle of skew rays — straight lines that twist round the axis like the rulings of a cooling tower, none of them crossing the axis — the sum grows with the area of the loop. No surface is at right angles to all of those lines, so no arrangement of lenses and mirrors fed from a single point can produce them. A sensor looking at such a bundle would find slopes that disagree with each other, and the size of the disagreement round a loop measures how far the light is from having come from one point.

The surface that winds

The third curve in the last figure is the exception that sharpens the rule. A vortex beam — light made by passing an ordinary beam through a plate whose thickness rises steadily round its centre, a spiral staircase of glass — has a wavefront shaped like a helical ramp, climbing one wavelength for each turn round the axis. Every ray is still at right angles to the ramp. But the ramp is not a single-valued surface: go once round the axis and it comes back one wavelength higher, so the slopes summed round any loop enclosing the axis give one wavelength, whatever the loop’s size.

Locally the rays are the gradient of something; globally the something is defined only up to whole wavelengths, which for a wave is all that matters, since a phase that differs by a whole cycle is the same phase. At the axis itself the wavefront’s height has no value at all, and there the light must be dark. These dark lines, which a wavefront can wind round, are where the ray picture of this essay hands over to the wave picture of the dark lines a wave cannot avoid, and they are why a real Shack–Hartmann sensor looking through strong turbulence, where such points appear spontaneously, sometimes finds slopes that cannot be added up.

Where it is used

The arithmetic in the lenslet figure is done hundreds or thousands of times a second in adaptive optics. A telescope’s image is blurred by the turbulent atmosphere, which puts a wavefront error of several wavelengths on starlight and changes it in milliseconds. A Shack–Hartmann sensor measures the slopes, a computer adds them up into a wavefront, and a mirror whose surface is pushed by hundreds of actuators takes on the opposite shape, so that the light leaving it is flat again. The same loop corrects lasers that must stay focused through hot air, and, run once rather than continuously, measures the aberrations of a human eye before laser surgery reshapes its cornea.

In each case the sensor’s output is trusted because of a theorem about rays from a point, applied to light that has come through a lens, an atmosphere or an eye. Starlight is the cleanest case: a star is a point, so whatever the air does to its light, the rays reaching the telescope are normal to a surface, and the surface is what the mirror must undo.

What the pictures cannot show

The ray figure is two-dimensional, a slice through the beam, and in a slice any family of curves crossing a set of lines can be fitted with perpendicular curves; the content of the theorem is that the same holds in three dimensions, where most bundles of lines have no perpendicular surfaces at all. The skew bundle in the last figure is such a bundle, and it can only be shown through what its slopes do round a loop.

The lenslet figure uses a smooth wavefront and perfect lenslets. A real sensor’s slopes carry noise from photon counting, the spots blur when the wavefront curves within a single lenslet, and a pupil edge that cuts a lenslet in half gives it a biased slope. Rebuilding the wavefront from noisy slopes is a problem in estimation, and the least-squares answer is one of several in use.

The domain of the theorem is geometrical optics: wavelengths much smaller than any feature of the optics, light from a single point, and no scattering. Inside it, rays are always the normals of the equal-path surfaces. At caustics, at foci and near the dark axis of a vortex, where rays cross or wavefronts wind, the ray picture itself fails and the wave picture takes over.

Still open: measuring a wavefront that tears

A Shack–Hartmann sensor assumes the wavefront it sees is smooth over each lenslet and integrable across the pupil. Light that has come through strong turbulence — over a long horizontal path near the ground, or through the atmosphere at a low angle — breaks both assumptions. Its intensity drops to zero at scattered points, the wavefront winds round them, and the measured slopes contain loops that sum to whole wavelengths. Methods exist to find the winding points and handle them, either by locating them in the slope data or by using sensors that measure phase directly by interference, and none yet works reliably at the speed adaptive optics needs when such points are numerous. How to correct light whose wavefront is no longer a single surface is a live engineering question for free-space laser communication and for the next generation of ground-based telescopes.

The theorem that makes the ordinary case work is old and short. Mark along every ray from a point the place it reaches after a fixed optical path, and the marks lie on a surface every ray crosses at a right angle — 90° to within 0.005° on rays traced through a glass surface with 205 waves of aberration — because the gradient of the optical path is n times the ray direction; so the slopes a lenslet array measures can be added up into a wavefront, their sum round any loop is zero, and a bundle whose loop sums are not zero cannot have come from a point. Every wavefront sensor in every adaptive-optics system rests on that, and every vortex beam is the case where it holds everywhere but on one line.

Part 7 of 7

This essay is one argument about Fermat. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

AberrationAdaptive opticsFermat's principleGradientOptical path lengthOptical vortexRay opticsWavefront