Optics

The mirror that cannot focus, and the shape that can

A perfect sphere does not bring parallel light to a point. The blur is not a manufacturing defect — it is what the shape does, and the shape is used anyway, for a reason worth knowing.

Assumes: What a lens is doing, and why three rays are enough

The lens equation relates three distances and says nothing about which rays it applies to. That silence conceals an approximation, and this rung is about what happens when the approximation is removed and the rays are traced honestly.

The answer is that a spherical surface — the shape of nearly every lens and mirror ever made — does not have a focus. It has a region.

A sphere does not have a focus. Parallel rays reflected off a spherical mirror. Rays striking further from the axis cross it nearer the mirror, so there is no single point where all of them meet — the blur is spherical aberration, and a perfect sphere has it inherently.
Fig. 1 Parallel rays reflected from a spherical mirror, each turned by the exact law of reflection about the local normal. The rays cross the axis at visibly different places, spreading over sixteen units of the drawing, and the paraxial focus is only where the innermost of them go.

What the construction assumed

Every step of the standard imaging construction contains the same substitution, and it is worth locating it precisely.

A ray striking a curved surface at height hh from the axis meets it at an angle set by the local normal, and the geometry involves sin\sin and tan\tan of that angle. Every derivation of the mirror and lens equations replaces those by the angle itself. It is the same substitution that makes a pendulum simple, applied to the same function, and it fails in the same way — quadratically at first, then decisively. The exact pendulum period is the same story told with time instead of distance.

Keeping only the first term is called the paraxial approximation, and what it produces is the familiar result: all parallel rays cross the axis at R/2R/2, so the focal length is half the radius of curvature and the shape has a focus. Keeping the next term produces a focal distance that depends on hh:

f(h)R2h24R,f(h) \approx \frac{R}{2} - \frac{h^2}{4R},

and the h2h^2 is the whole of the problem. Rays further from the axis cross nearer the mirror, and the crossing points spread over a length proportional to the square of the aperture.

The figure is a direct check on that formula. With a radius of curvature of 460 units and rays out to 166, the predicted spread is h2/4R=15h^2/4R = 15 units, and the traced rays spread over 16. Nothing in the tracing knows about the formula — it reflects each ray about the local normal and records where it crosses — so the agreement is between two independent routes to the same number.

Why it is a shape problem, not a quality problem

The point to be clear about is that nothing is wrong with the mirror. It is a perfect sphere; the aberration is what a perfect sphere does.

A sphere does not have a focus. Parallel rays reflected off a spherical mirror. Rays striking further from the axis cross it nearer the mirror, so there is no single point where all of them meet — the blur is spherical aberration, and a perfect sphere has it inherently.
Fig. 2 The same mirror stopped down to half the aperture. The spread collapses to about a quarter of its previous value, because it goes as the square of the height — which is the standard remedy, and it costs three-quarters of the light.

Stopping down works, and it works quadratically: halving the aperture quarters the blur. That is why a pinhole camera has no aberration worth mentioning and why cheap lenses look much better at small apertures. It is also why the remedy is so unattractive, since the light collected falls as the square of the aperture at exactly the same rate. Every unit of sharpness bought this way is paid for in brightness, one for one.

The alternative is to change the shape, and the shape that works is a parabola.

A parabola brings every ray to one point. Parallel rays reflected off a parabolic mirror, each by the exact law of reflection about the local normal. Every ray crosses the axis at the same place, which is the defining property of the shape.
Fig. 3 The same rays on a parabolic mirror of the same paraxial focal length. Every ray crosses the axis at the same point, exactly — which the figure states by checking the spread of the crossings rather than by asserting it in the caption.

A parabola brings all parallel rays to one point, with no approximation anywhere, and that is essentially its definition: the locus of points equidistant from a focus and a line. Light arriving parallel to the axis and reflecting to the focus travels the same total distance whatever height it struck at, so every path arrives in phase and the convergence is exact.

The generator computes the two cases with the same reflection code and the same ray heights, and prints the spread of the crossings for each. The sphere gives sixteen units; the parabola gives zero. That difference is the figure’s assertion, and it would fail loudly if the surface normals were computed wrongly — which, during development, they were.

The paraxial answer is still the reference

It would be easy to read all this as showing that the imaging equation is wrong. It is not, and the distinction is worth care.

The standard construction locates an image with the paraxial rays, and that answer is still the right one — it is where the centre of the blur sits, and it is what the mirror equation computes. What the construction cannot say is how large the blur is, because it has already assumed the rays it would need in order to find out. Every statement in this essay lives in the difference between that point and the patch of light actually there.

1/u+1/v=1/f1/u + 1/v = 1/f is exactly true for rays close enough to the axis, and the aberrations are defined as departures from what it predicts. That is not a circular definition but a working one: the paraxial result supplies the ideal image position and magnification, and every real ray’s miss distance is measured against it.

Plotting image distance against object distance in units of the focal length shows a relation with no aperture in it anywhere. That is the honest limitation: the paraxial relation is exact in the limit it names and says nothing about departures from it, so a mirror’s position of focus is a well-defined quantity and its quality of focus is not a question the relation can be asked.

The practical consequence is that a designer needs both. The paraxial calculation lays out the system — where the elements go, what powers they need, where the image lands, what the magnification is — in a few lines of arithmetic that can be done by hand and reasoned about. The exact ray trace then evaluates how badly that layout performs, and the design is adjusted. First-order optics decides the architecture; higher-order optics decides the quality; and no amount of exact tracing tells anyone where to put the lenses in the first place.

That division of labour is why the approximation survived being known to be wrong for three hundred years, and it is the general answer to the question of what a superseded model is for. A model that gives the right structure with the wrong details is not replaced by one that gives the right details — it is used to set up the problem that the second one then solves.

A concave mirror forming a real image. An object 2.50 focal lengths from a concave mirror. The image forms where the construction rays cross, inverted and magnified -0.67×.
Fig. 4 The reference the failure is measured against: the paraxial construction for a concave mirror, an object at 2.50 focal lengths giving an inverted image magnified 0.67×-0.67\times. Every ray here is drawn under the assumption the rest of this page is about — that the angles are small enough for sinθ\sin\theta and θ\theta to be the same number. Within that assumption a mirror has a focus and the construction is exact. The aberration on this page is entirely the difference between sinθ\sin\theta and θ\theta, which is why it is a shape problem rather than a manufacturing one.

So why is anything spherical

Given that a parabola works and a sphere does not, the fact that nearly every optical surface ever manufactured is spherical needs explaining, and the explanation is entirely about how surfaces are made.

A sphere is the only shape with no preferred point on it. Rub two surfaces together with abrasive between them, with random relative motion, and the high spots wear preferentially; the process converges, on its own, to a pair of matching spherical surfaces, because a sphere is the only shape that can slide over its mate in every direction and every orientation while staying in contact. Grinding a sphere requires no measurement of where on the surface the tool is. That is a very large advantage.

Every other shape has to be figured: measured, corrected locally, measured again. Until the late twentieth century that meant hand work by a small number of skilled people, at a cost that scaled with area, and the results were tested against reference surfaces that themselves had to be made. Aspheric surfaces are now moulded and diamond-turned in enormous numbers, which is why phone cameras contain them and why a modern lens has fewer elements than its equivalent of thirty years ago — but the economics were the other way round for three centuries, and the entire vocabulary of lens design is built around correcting the aberrations of spheres rather than avoiding them.

The correcting is done by combination. Aberration has a sign that depends on the surface’s orientation and curvature, so a positive element and a negative one can be arranged to cancel each other’s spherical aberration while their focusing powers do not cancel. That is the same trick as the achromatic doublet, applied to a different defect, and it is why a good lens has six or ten elements when one would form an image.

The parabola’s own failure

The parabola solves the problem on the axis, and only there. The figures above draw light arriving parallel to the axis, which is a star directly ahead, and a star slightly off to the side is a different problem.

Rays arriving at an angle to a parabola’s axis do not converge to a point. They form an asymmetric flare — a small comet-shaped smear with a bright head and a fan behind it, which is the aberration called coma, and it grows linearly with the angle off-axis. A parabolic telescope therefore has a small usable field: perfect at the centre, visibly comatic a few arcminutes out. Newton’s reflector, and every simple Newtonian since, has this.

The escape is to give up the idea of one surface doing the whole job. A Ritchey–Chrétien telescope uses two hyperbolic mirrors chosen so that the coma of one cancels the coma of the other, giving a field an order of magnitude larger at the cost of two surfaces neither of which images well alone. Nearly every large research telescope is one, Hubble included.

This is a pattern worth extracting, because it recurs whenever a design is pushed. A single element optimised for one condition performs perfectly there and degrades away from it; a combination optimised jointly performs slightly worse at the optimum and far better over a range. Choosing between them is choosing what the instrument is for.

The same defect in a lens, and the free half of the cure

Everything above was drawn with mirrors because reflection is easier to trace, and lenses have the identical problem for the identical reason: two spherical refracting surfaces, each bending by whatever Snell’s law requires at the local angle, with the outer rays over-bent.

The same defect appears in a lens, and by the same argument: the outer rays of a real bundle cross nearer the lens than the paraxial construction puts them, so the image is a small bright core inside a halo. A lens has a second problem a mirror does not — its focal length depends on wavelength — so the two defects compound, which is why the correction of a refracting telescope is harder than the correction of a reflecting one.

There is one piece of good news, and it costs nothing. The aberration of a lens depends on how the total bending is shared between its two surfaces. A ray that is bent a little at each surface accumulates less error than one bent a lot at one and not at all at the other, because the error grows faster than linearly with the angle at each surface.

That has a directly usable consequence. A plano-convex lens used to focus parallel light has about four times less spherical aberration with its curved side facing the incoming parallel beam than with the flat side facing it. The lens is the same lens; only its orientation changed. Turning it round moves it from splitting the bending between the two surfaces to doing all of it at one, and the penalty is a factor of four in a defect that goes as the cube of the aperture.

This is one of the few pieces of optical design available to anyone with a lens in a drawer, and it generalises to a design rule: for a lens of given power, there is a shape — a particular ratio of the two curvatures — that minimises the spherical aberration, and it depends on the refractive index and on where the object is. Working out that shape is the “bending” of a lens, and it is the first thing done to any element before more expensive corrections are considered.

The equivalent statement for a system rather than an element is to split the power across more elements. Two lenses of half the power each aberrate far less than one of the full power, because each bends the ray half as much and the defect grows faster than linearly. That is why long-focus systems are easier to correct than short ones, and why fast wide-angle lenses are the hardest objects in the catalogue.

What the aberration costs, measured

Two numbers make the practical scale of this concrete, and one of them is the most expensive optical error ever made.

The first is ordinary. A photographic lens at full aperture is usually aberration-limited, and stopping it down improves it until diffraction takes over. The crossing point is around f/5.6f/5.6 for a general-purpose lens, and the residual blur there is a few microns — comparable to a sensor pixel, which is not a coincidence, since sensors are specified against what the lenses in front of them can deliver.

The second is Hubble. Its 2.4-metre primary mirror was ground to a hyperboloid, tested against a null corrector — an auxiliary optic that converts the expected wavefront into a flat one so that any departure shows as a fringe. A spacer inside that corrector was positioned 1.3 millimetres wrong, because a technician used a field-fitted cap rather than the specified assembly, and the reflection came off a painted surface rather than the intended one. The corrector therefore reported the wrong shape as correct. The mirror was ground, extremely accurately, to a figure whose edge was about 2.2 microns too flat.

Two microns on a 2.4-metre mirror is a relative error of one part in a million, and it was enough to spread half the light of a star into a halo several arcseconds across, which is roughly the blur the atmosphere gives from the ground — the very thing an aperture that large had been put in orbit to beat. The telescope’s entire purpose was defeated by an error smaller than a wavelength of light multiplied by four.

The repair is the part worth remembering. The mirror could not be replaced, so instruments were built with corrective optics of exactly the opposite error — small mirrors figured with the same wrong conic constant, mounted so the light bounced off them on its way to each detector. An aberration cancelled by an equal and opposite aberration is the same move as the achromatic doublet, executed in orbit with a robotic arm, and it worked to specification. There were also two independent test instruments that had reported the mirror as wrong, and both were dismissed as less trustworthy than the null corrector.

The lens that is not allowed to be corrected

Optical spherical aberration is a defect of a particular shape, and grinding a different shape removes it. There is a case where a theorem forbids that escape, and it held up an entire instrument for sixty years.

An electron microscope focuses with magnetic fields rather than glass, and its lenses are round coils. Scherzer proved in 1936 that any electron lens which is rotationally symmetric, static, free of space charge and forms a real image must have positive spherical aberration. There is no shape of coil that avoids it, because the constraint is on the field the coil can produce rather than on the surface it presents.

That is a much stronger statement than the one on this page. A spherical mirror is a bad choice among available choices; a round electron lens is the only choice, and it is bad. The resolution of electron microscopes accordingly sat well short of what the electron’s wavelength allowed for six decades — the wavelength offered picometres and the aberration delivered ångströms.

The correction, when it came in the 1990s, had to break one of Scherzer’s premises, and the one that could be spared was the symmetry. A stack of non-round multipole elements introduces a negative spherical aberration that cancels the round lenses’, at the cost of a great many alignment degrees of freedom and a computer to hold them. That is the same trade as the aspheric surface, made against a theorem rather than against a grinding cost.

Where the model stops

The tracing above is exact geometry, and geometry is not the whole story.

It is still a ray model. Even the parabola’s perfect point is not a point: the finite aperture makes an Airy disc, and the honest question about any optical system is whether its aberration blur is smaller or larger than its diffraction blur. A system whose aberrations are below that threshold is called diffraction-limited, and the word means that further improvement in the surfaces would change nothing.

Only one aberration is drawn. Spherical aberration is one of five classical monochromatic defects, and the others — coma, astigmatism, field curvature and distortion — have their own dependencies on aperture and field angle. A real design balances all five simultaneously, plus the chromatic ones, which is why lens design is an optimisation problem rather than a construction.

The mirror is a curve on a page. These are two-dimensional sections; a real mirror is a surface of revolution and the crossings form a caustic — a three-dimensional envelope surface, the same object visible as the bright cusped curve on the bottom of a mug of tea, which is spherical aberration in a coffee cup and the most-photographed aberration in existence.

And the source is at infinity. All rays here arrive parallel. An object at finite distance strikes the surface with a different set of angles, so the aberration of a given surface depends on where the object is — which is why a lens designed for landscape work performs poorly close up and why macro lenses are separate designs rather than the same ones focused nearer.

The ladder from here

Later rungs: the other four Seidel aberrations, each with its characteristic star image. The aplanatic condition, and lens shapes that minimise spherical aberration for a given power. The Schmidt corrector plate, which is an aspheric plate at the centre of curvature of a spherical mirror and gives a wide field with no coma at all. Ritchey–Chrétien and other two-mirror forms. Ray-tracing as an optimisation, and the merit functions real designs are scored against. Wavefront error rather than ray error, and the quarter-wave criterion that says when a system is diffraction-limited. Adaptive optics, which corrects a wavefront that changes a thousand times a second. And the caustic treated as a mathematical object in its own right, where it turns out that only a small number of stable caustic shapes exist, and that they are the same ones classified in the fold of the rainbow’s deviation curve.

Part 2 of 6

This essay is one argument about Imaging. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Aspheric surfaceCausticConic sectionFocal lengthThe paraxial approximationSpherical aberration