Optics

The mirror that cannot focus, and the shape that can

A perfect sphere does not bring parallel light to a point. The blur is not a manufacturing defect — it is what the shape does, and the shape is used anyway, for a reason worth knowing.

The lens equation relates three distances and says nothing about which rays it applies to. That silence conceals an approximation, and this rung is about what happens when the approximation is removed and the rays are traced honestly.

The answer is that a spherical surface — the shape of nearly every lens and mirror ever made — does not have a focus. It has a region.

A sphere does not have a focusParallel rays reflected off a spherical mirror. Rays striking further from the axis cross it nearer the mirror, so there is no single point where all of them meet — the blur is spherical aberration, and a perfect sphere has it inherently.paraxial focuscrossings spread over 16 units
Fig. 1 Parallel rays reflected from a spherical mirror, each turned by the exact law of reflection about the local normal. The rays cross the axis at visibly different places, spreading over sixteen units of the drawing, and the paraxial focus is only where the innermost of them go.

What the construction assumed

Every step of the standard imaging construction contains the same substitution, and it is worth locating it precisely.

A ray striking a curved surface at height hh from the axis meets it at an angle set by the local normal, and the geometry involves sin\sin and tan\tan of that angle. Every derivation of the mirror and lens equations replaces those by the angle itself. It is the same substitution that makes a pendulum simple, applied to the same function, and it fails in the same way — quadratically at first, then decisively. The exact pendulum period is the same story told with time instead of distance.

Keeping only the first term is called the paraxial approximation, and what it produces is the familiar result: all parallel rays cross the axis at R/2R/2, so the focal length is half the radius of curvature and the shape has a focus. Keeping the next term produces a focal distance that depends on hh:

f(h)R2h24R,f(h) \approx \frac{R}{2} - \frac{h^2}{4R},

and the h2h^2 is the whole of the problem. Rays further from the axis cross nearer the mirror, and the crossing points spread over a length proportional to the square of the aperture.

The figure is a direct check on that formula. With a radius of curvature of 460 units and rays out to 166, the predicted spread is h2/4R=15h^2/4R = 15 units, and the traced rays spread over 16. Nothing in the tracing knows about the formula — it reflects each ray about the local normal and records where it crosses — so the agreement is between two independent routes to the same number.

Why it is a shape problem, not a quality problem

The point to be clear about is that nothing is wrong with the mirror. It is a perfect sphere; the aberration is what a perfect sphere does.

A sphere does not have a focusParallel rays reflected off a spherical mirror. Rays striking further from the axis cross it nearer the mirror, so there is no single point where all of them meet — the blur is spherical aberration, and a perfect sphere has it inherently.paraxial focuscrossings spread over 4 units
Fig. 2 The same mirror stopped down to half the aperture. The spread collapses to about a quarter of its previous value, because it goes as the square of the height — which is the standard remedy, and it costs three-quarters of the light.

Stopping down works, and it works quadratically: halving the aperture quarters the blur. That is why a pinhole camera has no aberration worth mentioning and why cheap lenses look much better at small apertures. It is also why the remedy is so unattractive, since the light collected falls as the square of the aperture at exactly the same rate. Every unit of sharpness bought this way is paid for in brightness, one for one.

The alternative is to change the shape, and the shape that works is a parabola.

A parabola brings every ray to one pointParallel rays reflected off a parabolic mirror, each by the exact law of reflection about the local normal. Every ray crosses the axis at the same place, which is the defining property of the shape.paraxial focusevery ray crosses at one point
Fig. 3 The same rays on a parabolic mirror of the same paraxial focal length. Every ray crosses the axis at the same point, exactly — which the figure states by checking the spread of the crossings rather than by asserting it in the caption.

A parabola brings all parallel rays to one point, with no approximation anywhere, and that is essentially its definition: the locus of points equidistant from a focus and a line. Light arriving parallel to the axis and reflecting to the focus travels the same total distance whatever height it struck at, so every path arrives in phase and the convergence is exact.

The generator computes the two cases with the same reflection code and the same ray heights, and prints the spread of the crossings for each. The sphere gives sixteen units; the parabola gives zero. That difference is the figure’s assertion, and it would fail loudly if the surface normals were computed wrongly — which, during development, they were.

The paraxial answer is still the reference

It would be easy to read all this as showing that the imaging equation is wrong. It is not, and the distinction is worth care.

A concave mirror forming a real imageAn object 2.50 focal lengths from a concave mirror. The image forms where the construction rays cross, inverted and magnified -0.67×.FCobjectimagethe mirror equation is the lens equation
Fig. 4 The standard mirror construction, with the image located by the paraxial equation. Every ray in it is a paraxial ray, and within that restriction the construction is exact — the equation is not an approximation to the behaviour of paraxial rays, it is their behaviour.

1/u+1/v=1/f1/u + 1/v = 1/f is exactly true for rays close enough to the axis, and the aberrations are defined as departures from what it predicts. That is not a circular definition but a working one: the paraxial result supplies the ideal image position and magnification, and every real ray’s miss distance is measured against it.

Image distance against object distanceImage distance in focal lengths against object distance in focal lengths. At exactly one focal length the image runs off to infinity; inside it the image distance goes negative, which means virtual.0.511.522.533.54-6-4-20246object distance (focal lengths)object at the focusvirtual imagesreal images
Fig. 5 Image distance against object distance, in units of the focal length. This curve is the paraxial prediction, and every aberration in every optical system is quoted as a departure from a point on it.

The practical consequence is that a designer needs both. The paraxial calculation lays out the system — where the elements go, what powers they need, where the image lands, what the magnification is — in a few lines of arithmetic that can be done by hand and reasoned about. The exact ray trace then evaluates how badly that layout performs, and the design is adjusted. First-order optics decides the architecture; higher-order optics decides the quality; and no amount of exact tracing tells anyone where to put the lenses in the first place.

That division of labour is why the approximation survived being known to be wrong for three hundred years, and it is the general answer to the question of what a superseded model is for. A model that gives the right structure with the wrong details is not replaced by one that gives the right details — it is used to set up the problem that the second one then solves.

So why is anything spherical

Given that a parabola works and a sphere does not, the fact that nearly every optical surface ever manufactured is spherical needs explaining, and the explanation is entirely about how surfaces are made.

A sphere is the only shape with no preferred point on it. Rub two surfaces together with abrasive between them, with random relative motion, and the high spots wear preferentially; the process converges, on its own, to a pair of matching spherical surfaces, because a sphere is the only shape that can slide over its mate in every direction and every orientation while staying in contact. Grinding a sphere requires no measurement of where on the surface the tool is. That is a very large advantage.

Every other shape has to be figured: measured, corrected locally, measured again. Until the late twentieth century that meant hand work by a small number of skilled people, at a cost that scaled with area, and the results were tested against reference surfaces that themselves had to be made. Aspheric surfaces are now moulded and diamond-turned in enormous numbers, which is why phone cameras contain them and why a modern lens has fewer elements than its equivalent of thirty years ago — but the economics were the other way round for three centuries, and the entire vocabulary of lens design is built around correcting the aberrations of spheres rather than avoiding them.

The correcting is done by combination. Aberration has a sign that depends on the surface’s orientation and curvature, so a positive element and a negative one can be arranged to cancel each other’s spherical aberration while their focusing powers do not cancel. That is the same trick as the achromatic doublet, applied to a different defect, and it is why a good lens has six or ten elements when one would form an image.

The parabola’s own failure

The parabola solves the problem on the axis, and only there. The figures above draw light arriving parallel to the axis, which is a star directly ahead, and a star slightly off to the side is a different problem.

Rays arriving at an angle to a parabola’s axis do not converge to a point. They form an asymmetric flare — a small comet-shaped smear with a bright head and a fan behind it, which is the aberration called coma, and it grows linearly with the angle off-axis. A parabolic telescope therefore has a small usable field: perfect at the centre, visibly comatic a few arcminutes out. Newton’s reflector, and every simple Newtonian since, has this.

The escape is to give up the idea of one surface doing the whole job. A Ritchey–Chrétien telescope uses two hyperbolic mirrors chosen so that the coma of one cancels the coma of the other, giving a field an order of magnitude larger at the cost of two surfaces neither of which images well alone. Nearly every large research telescope is one, Hubble included.

This is a pattern worth extracting, because it recurs whenever a design is pushed. A single element optimised for one condition performs perfectly there and degrades away from it; a combination optimised jointly performs slightly worse at the optimum and far better over a range. Choosing between them is choosing what the instrument is for.

The same defect in a lens, and the free half of the cure

Everything above was drawn with mirrors because reflection is easier to trace, and lenses have the identical problem for the identical reason: two spherical refracting surfaces, each bending by whatever Snell’s law requires at the local angle, with the outer rays over-bent.

A converging lens making a real imageAn object 2.00 focal lengths from a thin converging lens. The image sits where the construction rays cross, at 2.00 focal lengths, magnified -1.00×.FFobjectimageu = 2.00 fv = 2.00 fmagnification -1.00
Fig. 6 A lens forming an image with the paraxial construction. The outer rays of a real bundle cross nearer the lens than these, and the image is a small bright core inside a halo rather than a point.

There is one piece of good news, and it costs nothing. The aberration of a lens depends on how the total bending is shared between its two surfaces. A ray that is bent a little at each surface accumulates less error than one bent a lot at one and not at all at the other, because the error grows faster than linearly with the angle at each surface.

That has a directly usable consequence. A plano-convex lens used to focus parallel light has about four times less spherical aberration with its curved side facing the incoming parallel beam than with the flat side facing it. The lens is the same lens; only its orientation changed. Turning it round moves it from splitting the bending between the two surfaces to doing all of it at one, and the penalty is a factor of four in a defect that goes as the cube of the aperture.

This is one of the few pieces of optical design available to anyone with a lens in a drawer, and it generalises to a design rule: for a lens of given power, there is a shape — a particular ratio of the two curvatures — that minimises the spherical aberration, and it depends on the refractive index and on where the object is. Working out that shape is the “bending” of a lens, and it is the first thing done to any element before more expensive corrections are considered.

The equivalent statement for a system rather than an element is to split the power across more elements. Two lenses of half the power each aberrate far less than one of the full power, because each bends the ray half as much and the defect grows faster than linearly. That is why long-focus systems are easier to correct than short ones, and why fast wide-angle lenses are the hardest objects in the catalogue.

What the aberration costs, measured

Two numbers make the practical scale of this concrete, and one of them is the most expensive optical error ever made.

The first is ordinary. A photographic lens at full aperture is usually aberration-limited, and stopping it down improves it until diffraction takes over. The crossing point is around f/5.6f/5.6 for a general-purpose lens, and the residual blur there is a few microns — comparable to a sensor pixel, which is not a coincidence, since sensors are specified against what the lenses in front of them can deliver.

The second is Hubble. Its 2.4-metre primary mirror was ground to a hyperboloid, tested against a null corrector — an auxiliary optic that converts the expected wavefront into a flat one so that any departure shows as a fringe. A spacer inside that corrector was positioned 1.3 millimetres wrong, because a technician used a field-fitted cap rather than the specified assembly, and the reflection came off a painted surface rather than the intended one. The corrector therefore reported the wrong shape as correct. The mirror was ground, extremely accurately, to a figure whose edge was about 2.2 microns too flat.

Two microns on a 2.4-metre mirror is a relative error of one part in a million, and it was enough to spread half the light of a star into a halo several arcseconds across, which is roughly the blur the atmosphere gives from the ground — the very thing an aperture that large had been put in orbit to beat. The telescope’s entire purpose was defeated by an error smaller than a wavelength of light multiplied by four.

The repair is the part worth remembering. The mirror could not be replaced, so instruments were built with corrective optics of exactly the opposite error — small mirrors figured with the same wrong conic constant, mounted so the light bounced off them on its way to each detector. An aberration cancelled by an equal and opposite aberration is the same move as the achromatic doublet, executed in orbit with a robotic arm, and it worked to specification. There were also two independent test instruments that had reported the mirror as wrong, and both were dismissed as less trustworthy than the null corrector.

Where the model stops

The tracing above is exact geometry, and geometry is not the whole story.

It is still a ray model. Even the parabola’s perfect point is not a point: the finite aperture makes an Airy disc, and the honest question about any optical system is whether its aberration blur is smaller or larger than its diffraction blur. A system whose aberrations are below that threshold is called diffraction-limited, and the word means that further improvement in the surfaces would change nothing.

Only one aberration is drawn. Spherical aberration is one of five classical monochromatic defects, and the others — coma, astigmatism, field curvature and distortion — have their own dependencies on aperture and field angle. A real design balances all five simultaneously, plus the chromatic ones, which is why lens design is an optimisation problem rather than a construction.

The mirror is a curve on a page. These are two-dimensional sections; a real mirror is a surface of revolution and the crossings form a caustic — a three-dimensional envelope surface, the same object visible as the bright cusped curve on the bottom of a mug of tea, which is spherical aberration in a coffee cup and the most-photographed aberration in existence.

And the source is at infinity. All rays here arrive parallel. An object at finite distance strikes the surface with a different set of angles, so the aberration of a given surface depends on where the object is — which is why a lens designed for landscape work performs poorly close up and why macro lenses are separate designs rather than the same ones focused nearer.

The ladder from here

Later rungs: the other four Seidel aberrations, each with its characteristic star image. The aplanatic condition, and lens shapes that minimise spherical aberration for a given power. The Schmidt corrector plate, which is an aspheric plate at the centre of curvature of a spherical mirror and gives a wide field with no coma at all. Ritchey–Chrétien and other two-mirror forms. Ray-tracing as an optimisation, and the merit functions real designs are scored against. Wavefront error rather than ray error, and the quarter-wave criterion that says when a system is diffraction-limited. Adaptive optics, which corrects a wavefront that changes a thousand times a second. And the caustic treated as a mathematical object in its own right, where it turns out that only a small number of stable caustic shapes exist, and that they are the same ones classified in the fold of the rainbow’s deviation curve.