Optics

The condition a lens must meet

A paraboloid brings every parallel ray to exactly one point. Move the source a fifth of a degree off axis and the image is a fan rather than a point, and the reason is a condition Abbe wrote down that has nothing to do with the axis — perfection at one point buys nothing at the next one along.

Assumes: What a lens is doing, and why three rays are enough · The mirror that cannot focus, and the shape that can

A spherical mirror cannot focus because rays striking further from the axis cross it nearer the mirror. The paraboloid cures that exactly: every ray from infinity crosses at one point, by the definition of the shape, and there is nothing left to improve.

Then move the source a fifth of a degree off the axis and the image falls apart.

The blur a mirror makes off its own axis. The size of the image blur at the paraxial focal plane, against how far off axis the source is, for parabola and sphere mirrors of the same focal length, traced with the exact law of reflection at every ray. On axis the paraboloid is perfect, by construction, and the sphere is not. A fifth of a degree off axis the paraboloid has a blur of 2.4 units at 1.2° for the parabola, 30.3 units at 1.2° for the sphere. The paraboloid's grows very nearly in proportion to the field angle, which is the signature of coma and is exactly what the offence against the sine condition predicts. This is the trade every reflecting telescope makes and the reason two mirrors are usually better than one: a single paraboloid buys a perfect axis at the price of a field a few minutes of arc wide, and the Ritchey–Chrétien pairing gives up the perfect axis to satisfy the sine condition and gets a usable field in exchange. The blur here is geometric only. Whether it matters depends on the diffraction limit sitting underneath it, which is the previous rung's subject, and on a large instrument the two cross at a field angle worth computing.
Fig. 1 The blur at the focal plane against how far off axis the source is, for a paraboloid and a sphere of the same focal length, traced with the exact law of reflection at every ray. The paraboloid is perfect at zero and its blur climbs very nearly in proportion to the field angle from there.

The blur that appears is coma: an asymmetric flare, sharp at one end and spread at the other, that grows linearly with the distance off axis and quadratically with the aperture. It is the aberration that limits a fast telescope, and it is present in a mirror that has no spherical aberration whatever.

The condition

Abbe’s statement is about the relationship between where a ray enters and how steeply it converges.

Consider a system that images an axial point perfectly — every ray from the object arrives at the image. Now ask for the image of a point a small distance η\eta off axis to be sharp as well. Fermat’s principle requires that all routes from the off-axis object to its image have the same optical path length, and expanding that requirement to first order in η\eta gives

hsinu=constant\frac{h}{\sin u} = \text{constant}

where hh is the ray’s height in the entrance pupil and uu the angle it makes with the axis on the way to the image. For an object at infinity the constant is the focal length, so the requirement is h=fsinuh = f\sin u.

That is the sine condition, and it is worth noticing what it is not. The paraxial relation between height and angle is h=ftanuh = f\tan u, and the two agree only to first order. A system built so that its rays obey the tangent relation exactly — which is what a paraboloid does — violates the sine condition by the difference between the two functions.

Abbe's ratio, measured on the traced rays. The quantity h divided by the sine of the angle at which the ray converges on the focus, in units of the paraxial focal length, against how far up the aperture the ray entered. Abbe's sine condition says that a system already free of spherical aberration images a small region round the axis faithfully only if this ratio is the same for every ray. A horizontal line means the condition is met. The parabola departs by 12.96 per cent across the aperture; The sphere departs by 7.18 per cent across the aperture. The paraboloid is the interesting case, because it is exactly stigmatic on axis — every ray from infinity crosses at one point, which is the definition of the shape — and it still fails this test. Perfection at one point buys nothing at the next one along. What the departure predicts is coma, a blur that grows linearly with the distance off axis and quadratically with the aperture, and the offaxis figure measures exactly that blur on the same surfaces. The condition is not a design rule invented for telescopes: it follows from requiring that the same optical path length join object and image for every route, and any instrument that images a field rather than a point has to meet it.
Fig. 2 The ratio h/(fsinu)h/(f\sin u) measured on the traced rays, against how far up the aperture the ray entered. A horizontal line means the condition is met. The paraboloid departs by thirteen per cent across the aperture — while being exactly stigmatic on axis — and the departure is what the coma in the previous figure is made of.

A system meeting both requirements at once, stigmatic on axis and obeying the sine condition, is called aplanatic, and the word is worth having because such systems are the ones that image a field rather than a point.

Why one surface cannot do it

The reason a paraboloid fails is a counting argument, and it is the useful way to think about the whole subject.

A single reflecting surface has one shape function to choose. Demanding stigmatism on axis uses it up completely: the requirement that all path lengths from infinity to a point be equal fixes the surface to be a paraboloid, with no freedom left. The sine condition is a second requirement on the same rays and there is nothing left to satisfy it with.

A converging lens making a real image. An object 2.44 focal lengths from a thin converging lens. The image sits where the construction rays cross, at 1.69 focal lengths, magnified -0.69×.
Fig. 3 What the counting argument is counting: the paraxial construction, in which every ray from a point on the object arrives at one point on the image and the whole system is described by a single number. That behaviour is what a single surface can be shaped to deliver — on the axis, for one conjugate pair, and nowhere else. Everything on this page is the question of what happens to the rays this construction leaves out, and the answer is that one shape function has already been spent buying this picture.

Give the system two surfaces and there are two shape functions, and both conditions can be met. That is the Ritchey–Chrétien design: two hyperboloids, chosen so that the pair is stigmatic on axis and obeys the sine condition, at the price that neither mirror alone is any good. It is what the Hubble telescope is, and what nearly every large professional telescope has been since about 1930.

The classical Cassegrain — a paraboloid primary and a hyperboloid secondary — is stigmatic on axis and does not satisfy the sine condition, so it has the same coma as the paraboloid alone. The change from one to the other is entirely about which conditions the two available freedoms are spent on.

A parabola brings every ray to one point. Parallel rays reflected off a parabolic mirror, each by the exact law of reflection about the local normal. Every ray crosses the axis at the same place, which is the defining property of the shape.
Fig. 4 Parallel rays reflected off a paraboloid, every one crossing the axis at the same point. This is the on-axis perfection the sine condition is indifferent to, and everything in this essay is about a defect this picture is incapable of showing.

Where the condition comes from, in one line

The derivation is short enough to give, and it is the reason the condition is a theorem rather than a design heuristic.

Take a system that images the axial object point OO to the axial image point II perfectly, so every path from OO to II has the same optical length. Now consider a second object point OO' a small distance η\eta from OO, perpendicular to the axis, and ask for its image II' at η\eta' from II.

A ray leaving OO' instead of OO has its path lengthened at the start by nηsinuo-n\eta\sin u_o, where uou_o is the angle it leaves at, because the extra distance is the projection of η\eta on the ray. Arriving at II' instead of II shortens it by nηsinuin'\eta'\sin u_i. For the image at II' to be sharp, the total path length must be the same for every ray, which requires

nηsinuo=nηsinuin\eta\sin u_o = n'\eta'\sin u_i

for all rays, and hence sinuo/sinui\sin u_o/\sin u_i constant — which for an object at infinity is h=fsinuh = f\sin u.

The whole argument is Fermat’s principle applied twice: once to fix the axial image and once, to first order in the displacement, to fix its neighbour. There is no approximation about small angles anywhere in it, which is why the sine appears and the tangent does not.

It also produces a second reading of the same equation. nηsinun\eta\sin u is, up to a factor, the étendue of the bundle, so the sine condition says that an aplanatic system conserves étendue ray by ray rather than merely in total — and the brightness no lens can increase is the same statement made about a whole beam. An instrument that violates it is one whose rays are not carrying their share.

The rule of thumb it produces

The practical form of all this is a formula for how large a field a given telescope has, and it is worth stating because it is unforgiving.

Coma’s angular blur for a paraboloid is approximately

θcomaη16F2\theta_{\text{coma}} \approx \frac{\eta}{16 F^2}

where η\eta is the field angle and FF is the focal ratio. The F2F^2 is what hurts. A slow instrument at f/10f/10 has a comatic blur sixteen times smaller than a fast one at f/2.5f/2.5 at the same field angle, so the usable field of a fast Newtonian is small in a way that surprises people who chose it for its speed.

Setting the coma equal to the diffraction limit gives the field over which the instrument is diffraction-limited, and for a fast paraboloid that is a couple of minutes of arc — smaller than the moon by a factor of fifteen. Everything outside it is comatic, and the fact that the mirror is perfect on axis is no consolation.

The standard the aberrations have to be compared against is diffraction. Two point sources are resolved or not according to a limit set by the aperture alone, and an aberration blur smaller than that limit is invisible — so the design target is not zero aberration but aberration below the diffraction spot. That is what makes the sine condition a practical criterion rather than an ideal: it says when the residual has stopped mattering.

What coma looks like, and why it is the worst-behaved aberration

The name comes from the shape: a comatic image is a small comet, with a bright head and a fan spreading away from it, and the fan always points either toward or away from the axis depending on the sign.

The asymmetry is what makes it so much more damaging than its size suggests. Spherical aberration and defocus produce blurs that are symmetric about the ideal image point, so the centroid of the light is still in the right place and a measurement of position is unbiased. Coma’s blur is not symmetric: the centroid is displaced from the head of the comet by about a third of the flare’s length, and the displacement grows with the field angle.

For astrometry that is fatal. A star measured near the edge of a comatic field is reported at a position that is systematically wrong, in a direction that depends on where in the field it fell, and no amount of averaging removes it because it is not noise. Every wide-field astrometric survey therefore uses an aplanatic design or a corrector, not because the images look better but because the positions are otherwise biased.

For photometry it is a different problem: the flare spills outside whatever aperture the photometry uses, so the measured brightness falls off toward the edge of the field in a way that has to be calibrated out. And for spectroscopy it is worse still, because the flare and the slit are not aligned in any consistent way.

The general point is one worth carrying past optics. A symmetric error degrades a measurement; an asymmetric one biases it, and the second is much harder to live with, because averaging more data reduces the first and not the second.

The correctors, and what they are correcting

A coma corrector is a small lens assembly placed near the focus of a fast paraboloid, and its job is stated exactly by the condition above: it is not correcting the mirror’s figure, which is right, but adjusting the mapping between ray height and convergence angle so that the combination satisfies the sine condition.

That is why such a corrector is a couple of centimetres across and works for a mirror half a metre wide. It does not touch the wavefront’s overall shape; it applies a small height-dependent change of angle. The same is true of the Schmidt plate, invented for the opposite problem — a spherical primary with a stop at its centre of curvature has no coma at all, by symmetry, and enormous spherical aberration, and the plate removes the second without disturbing the first.

The Schmidt case is the cleanest demonstration in the whole subject that these two aberrations are independent. A sphere with the stop at the centre of curvature has, from any direction, exactly the same geometry — every direction is an axis — so there is no way for it to distinguish on-axis from off-axis and no coma is possible. It has spherical aberration to spare, and one thin aspheric plate fixes that. Two problems, two independent cures, and the reason a Schmidt camera photographs a field several degrees wide.

The paraxial construction locates an image and says nothing whatever about its quality. Every statement in this essay lives in the gap between where that construction puts the image and where the light actually goes — and the mirror equation, which is exact in the paraxial limit, is silent about the entire subject. That silence is not a defect of the equation; it is what “paraxial” means.

What the condition buys: one blur for the whole field

The sine condition is usually presented as a way of removing coma, which is true and undersells it. What it actually secures is that the shape of the blur stops depending on where in the field the point was.

An aplanatic system’s image of a point near the axis is, to first order in the field angle, the same function translated — same size, same shape, same orientation. A system that violates the condition has an image whose flare grows and rotates as the point moves outward, so no two points in the field are imaged alike.

That distinction decides whether the instrument can be described at all. If the blur is the same everywhere, the image is the object convolved with a single point-spread function, and everything that follows from a convolution becomes available: a transfer function that says how much contrast survives at each spatial frequency, a single number for resolution, and the possibility of undoing some of the blur computationally. If the blur varies across the field, none of that holds — a deconvolution with one kernel sharpens the middle and smears the edges, and there is no such thing as “the” transfer function of the instrument.

The property has a name, isoplanatism, and the region over which it holds is an isoplanatic patch. It is why an aplanatic design is worth its extra surfaces even when the comatic flare would have been small enough to tolerate: it is the difference between an instrument that has a point-spread function and one that has a different one at every point.

The condition for the other direction, and why both cannot hold

Abbe’s condition asks that a point displaced sideways from the axis be imaged sharply. There is an obvious companion question — what does it take to image sharply a point displaced along the axis? — and it has an answer of the same form, found by Herschel in 1821, half a century earlier.

Running the same path-length argument for a longitudinal displacement gives Herschel’s condition: nsin2(uo/2)/nsin2(ui/2)n\sin^2(u_o/2)\,/\,n'\sin^2(u_i/2) must be constant across the aperture, where Abbe’s demanded that nsinuo/nsinuin\sin u_o / n'\sin u_i be constant.

Both are requirements on the same rays, and they can be divided one by the other. Using sinu=2sin(u/2)cos(u/2)\sin u = 2\sin(u/2)\cos(u/2), the quotient reduces to the demand that cos(uo/2)/cos(ui/2)\cos(u_o/2)/\cos(u_i/2) be constant too — and three constraints of that kind on one function force uo=±uiu_o = \pm u_i for every ray. That fixes the magnification: the two conditions hold together only at m=n/n|m| = n/n', which for an instrument working in air at both ends means unit magnification and nothing else.

The consequence is a genuine impossibility rather than a difficulty. No optical system that magnifies can image a three-dimensional region sharply. It may be aplanatic, in which case it images a surface well and loses sharpness immediately in front of and behind it, or it may satisfy Herschel’s condition and image a line along the axis well while losing the field — and there is no design, however many surfaces it is given, that does both.

Every microscope is built on the first choice, which is why depth is recovered by moving the focus and stacking rather than by capturing it, and why the depth of field of a high-aperture objective is measured in fractions of a micrometre. It is not a limitation of the glass. It is two path-length conditions that cannot both be true.

What it costs

The sine condition constrains a microscope objective more than a telescope. For an object at a finite distance the constant in h/sinuh/\sin u involves both conjugates, and satisfying it at high numerical aperture is the whole difficulty of designing an objective. A modern one has a dozen elements and most of them are there for the field rather than for the axis.

And it produces the reason immersion works twice. The condition in a medium of index nn is nh/sinunh/\sin u constant, so raising the index raises the achievable aperture — which is why an immersion objective resolves better — and it also means the correction has to be redone if the immersion medium changes. An objective designed for oil is not merely degraded in water; it is a different instrument.

Satisfying it costs light and space. Every additional surface reflects, absorbs and scatters, and a two-mirror aplanat has a hole in the middle of its primary and a secondary obstructing the aperture. The obstruction moves energy from the central peak of the diffraction pattern into the rings, which lowers contrast on extended objects — so the aplanatic design wins on field and loses on contrast, and which matters depends on what is being looked at.

The condition is about geometry and says nothing about colour. An aplanatic system built of glass is aplanatic at one wavelength, and the index of every glass depends on the wavelength, so a refracting aplanat has to satisfy the condition and be achromatic at the same time — two constraints on the same small number of freedoms. That difficulty is one of the strongest arguments for mirrors, which have no dispersion at all and therefore satisfy their conditions at every wavelength simultaneously.

And nothing here fixes astigmatism or field curvature. An aplanat is corrected for two aberrations out of the classical five. The Ritchey–Chrétien has astigmatism and a curved focal surface, which is why its detectors are either small or curved and why a third element is added when a really wide field is wanted.

A refracting design pays a cost a reflecting one does not. Two glasses’ focal lengths run differently with wavelength, and a doublet brings two wavelengths together and leaves the rest scattered about — so a refracting aplanat has to satisfy the sine condition and the achromatic one simultaneously, with the same two surfaces. A mirror has no dispersion at all, which is why every large telescope is one.

The sine condition and the diffraction bound are two halves of one statement about apertures. An aperture can collect a certain range of angles; the sine condition says what a system must do to image all of them to the same place, and the diffraction limit says how small a spot that range of angles can produce. Numerical aperture appears in both, with the immersion index inside it, which is why oil buys resolution and why it buys nothing if the sine condition is not met.

What the picture cannot show

The trace is meridional and two-dimensional. Real coma is a three-dimensional figure — a family of overlapping circles growing along the axis of the flare — and the blur plotted here is the spread of the meridional fan, which is the right order and not the right shape. A proper spot diagram needs rays from the whole pupil.

Diffraction is absent from the trace and present in reality. The blurs computed here are geometric, and below about the diffraction limit they are not what is seen. Whether a given amount of coma matters is always a comparison against that limit, and the comparison depends on the aperture, which the geometric trace does not know about.

The sine condition is a first-order statement in the field angle. It guarantees a sharp image of points near the axis, which is what it was derived for. A wide field needs the higher-order conditions too, and satisfying all of them simultaneously is not possible with finitely many surfaces — which is why lens design is an optimisation rather than a construction.

And the two surfaces drawn are ideal. Every real mirror has a figure error, a support-induced deformation and a thermal gradient, and on a large telescope those are comparable with the aberrations discussed here. The design condition sets what is achievable; what is achieved is a manufacturing question.

The ladder from here

Later rungs on this anchor: the five Seidel aberrations as one expansion, and why they are the complete list at that order; the aplanatic points of a sphere, which are the two conjugate positions where a single spherical surface is both stigmatic and aplanatic and which every immersion objective’s front element uses; field curvature and the Petzval sum, a constraint that survives every other correction; the wavefront description, in which all of this becomes a set of polynomial coefficients and the geometric picture is dropped; and adaptive correction, where the aberrations are measured and cancelled in real time rather than designed away.

The neighbouring ladders are what a lens is doing, where the paraxial construction is set up, and the mirror that cannot focus, whose defect the paraboloid cures and whose cure this essay is the price of. The image as a diffraction pattern is the limit all of these aberrations are measured against.

Part 4 of 6

This essay is one argument about Imaging. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

AberrationAplanaticComaField of viewOptical path lengthRay tracingSine conditionSpherical aberration