Optics

The focus that is a slab, not a plane

A lens images one plane and no other, which would make every photograph and every micrograph almost entirely out of focus. What rescues them is a tolerance — and there are two of them, one from rays and one from waves, which give different answers and stop being interchangeable exactly where microscopes work.

Assumes: What a lens is doing, and why three rays are enough · How far apart two things have to be

A thin lens obeys one equation, and the equation has one object distance on one side and one image distance on the other. Everything at that object distance is imaged sharply onto the sensor; nothing else is. Taken literally, every photograph ever made is a picture of one plane surrounded by blur.

Image distance against object distance. Image distance in focal lengths against object distance in focal lengths. At exactly one focal length the image runs off to infinity; inside it the image distance goes negative, which means virtual.
Fig. 1 The lens equation, drawn: image distance against object distance for a lens of a hundred and ten millimetres. One object distance maps to one image distance, and moving the object at all moves the image. Nothing in this relation permits two object planes to be in focus at once.

Photographs are not like that, and the reason is not that the equation is wrong. It is that “in focus” is not a statement the equation can make. Sharpness is a statement about how much blur can be tolerated, which is a statement about who is looking and how closely — and the moment a tolerance is named, the single sharp plane becomes a slab.

The ray account, and what it depends on

How far out of focus a thing can be before it looks it. The diameter of the blur circle a point at a given distance makes on the sensor, for a 50 mm lens focused at 3 metres, at 3 apertures. Only one distance is truly in focus — the curves all touch zero at 3 m and nowhere else — and everything about depth of field is the horizontal line: a tolerance of 20 micrometres, below which nothing looking at the picture can tell. At f/2 that tolerance is met between 2.86 and 3.15 m; at f/16, between 2.18 and 4.82 m. Depth of field is therefore not a property of the lens; it is a property of the lens, the distance, and how closely the result is going to be examined.
Fig. 2 The blur circle a point at a given distance makes on the sensor, for a fifty-millimetre lens focused at three metres, at three apertures. Every curve touches zero at three metres and nowhere else. The horizontal line is the tolerance, and it is the only thing that turns a single sharp plane into a range.

The geometry is elementary. A point out of focus sends a cone of rays through the lens which is intercepted before or after its apex, so it lands as a disc; the disc’s diameter is the aperture’s diameter times the fractional error in image distance. Fix a tolerance on that diameter and the range of object distances that meet it follows.

Two features of the result are worth pulling out because both are usually stated as though they were properties of lenses.

The tolerance is a decision. Twenty micrometres is a common figure for a full-frame camera, and it comes from asking that the blur be invisible in a print of a certain size at a certain viewing distance. Examine the same negative under a microscope and the tolerance shrinks to a wavelength, and the depth of field collapses to nearly nothing. The lens has not changed.

And the aperture enters only through the f-number, which is the focal length divided by the aperture diameter. That is why the same f-number gives similar depth on lenses of different focal length at the same framing — the two effects, a longer lens and a bigger aperture, cancel — and it is the reason photographers speak in f-numbers at all rather than in millimetres.

There is a distance at which the far limit runs off to infinity, called the hyperfocal distance, and it is worth naming because it is where the arithmetic changes character. Focus closer than it and the far limit is finite; focus at it and everything from half that distance to infinity is within tolerance. The transition is not gradual in the useful sense: a small change of focus near that point moves the far limit from twenty metres to unbounded.

Where rays stop being able to answer

The ray account says the blur can be made as small as wanted by stopping down, and therefore that the depth of field can be made as large as wanted. That is false, and the figure that shows why has no rays in it.

How thin a focus is, when rays are not enough. The intensity at the centre of a focused spot as the plane of observation is moved along the axis, at 550 nanometres, for numerical apertures of 0.25, 0.6, 1.3. Rays say a focus is a point and the intensity is infinite at one plane and zero elsewhere. Waves say it is a region: the intensity falls smoothly, reaching four fifths of its peak at λ over twice the square of the aperture, which is Rayleigh's quarter-wave criterion and is 0.16 micrometres at NA 1.3. The square is the important part. Doubling the aperture doubles the resolution and quarters the depth, so a high-power objective's focus is a slice a few hundred nanometres thick — and everything above and below it is a haze the image contains and cannot separate.
Fig. 3 The intensity at the centre of a focused spot as the plane of observation moves along the axis. Rays say the focus is a point and the intensity is a delta function; waves say it is a region, and the region’s thickness goes as one over the square of the numerical aperture. The marked points are Rayleigh’s criterion — a quarter of a wavelength of path error across the pupil, which leaves eighty-one per cent of the peak.

This is the point at which the ray description has to be given up. A lens does not make a point. It makes a spot whose width is set by the aperture and the wavelength — that is the whole of the diffraction limit — and moving along the axis away from focus, the spot does not immediately blur in the ray sense. It stays roughly the same width and loses intensity, and the distance over which it loses an appreciable amount is the depth of focus.

The criterion that has stuck is Rayleigh’s and it is worth stating in its own terms because it is not a criterion about blur at all. Defocusing by a distance zz puts a path error across the pupil, largest at the rim, of about zNA2/2z\,\mathrm{NA}^2/2. Rayleigh’s rule is that a quarter of a wavelength of path error is the most that can be tolerated, which gives a depth of ±λ/2NA2\pm\lambda/2\mathrm{NA}^2, and the figure shows that this leaves 81 per cent of the peak intensity. That last number is a computed consequence rather than part of the definition, and it is what makes the criterion a physical statement instead of a convention.

The square is where the two accounts part company. The ray depth goes as the f-number, so it improves in proportion as the aperture shrinks; the wave depth goes as the inverse square of the aperture, so it improves faster. At small apertures the wave depth is enormous and irrelevant. At large apertures it is the only one that matters, and it is smaller than the ray formula predicts.

The exchange rate

Resolution and depth, bought with the same money. Transverse resolution and axial depth of focus against numerical aperture, at 550 nanometres. One falls as the aperture and the other as its square, so the depth is proportional to the square of the resolution with only the wavelength in the coefficient — a relation that holds at every aperture and is what makes the trade a law rather than a habit. An objective resolving 200 nanometres has a focus about 195 nanometres thick, and one resolving twice as coarsely has four times the depth. Nothing about the design of the lens enters: this is the diffraction limit on both quantities at once, and it is the reason a confocal microscope's optical sectioning and its resolution improve and worsen together.
Fig. 4 Transverse resolution and axial depth against numerical aperture. One falls as the aperture, the other as its square, so the depth is proportional to the square of the resolution with only the wavelength in the coefficient. That relation holds at every aperture and is not a property of any lens design.

Written as a relation between the two things anybody wants, the result is stark: the depth of focus is proportional to the square of the transverse resolution, with a coefficient containing nothing but the wavelength. An objective that resolves twice as finely has a quarter of the depth. There is no design, no aspheric surface, no correction that changes the exchange rate, because both quantities are diffraction’s and neither is the lens’s.

For a microscopist this is the central constraint of the instrument. A ×100 oil-immersion objective at NA 1.4 resolves about 240 nanometres and has a depth of focus of about 280 — so its focus is a slab a quarter of a micrometre thick, thinner than most of the objects being looked at. Everything above and below that slab is still in the image, contributing a haze that carries no information and cannot be separated from what is in focus. Optical sectioning — confocal microscopy, two-photon excitation, light-sheet illumination — exists entirely to remove that haze, and each of those methods works by refusing to collect the out-of-focus light rather than by improving the depth, because the depth cannot be improved.

For a photographer the same relation appears in a different guise. Stopping down improves the geometric depth and worsens the diffraction blur; the two cross at an aperture that depends on the tolerance, and past it the whole image gets softer while the range within tolerance stops growing. The best small aperture on a full-frame camera is around f/8 to f/11 for exactly this reason, and it is a diffraction number rather than a lens number.

The number a designer actually uses

Between the two accounts sits a practical question — at what aperture is a lens best? — and the answer is a crossing rather than a formula.

The geometric blur falls as the aperture shrinks and the diffraction blur grows, so their sum has a minimum. Setting the two equal for a fifty-millimetre lens and a twenty-micrometre tolerance puts the crossing near f/11 on a full-frame sensor and near f/5.6 on a sensor half the size, because the smaller sensor is enlarged more and its tolerance is correspondingly tighter. That is the whole reason small cameras are used at wide apertures and large ones are not.

The same crossing exists in a microscope and is resolved the other way. There the tolerance is a wavelength, diffraction wins at every aperture, and the correct move is always to open up — which is why objectives are specified by numerical aperture rather than by anything else, and why immersion oil, whose only function is to raise the aperture past one, is worth the inconvenience.

And in lithography the crossing has been pushed so far that the trade has become the industry’s central constraint. Printing finer features means a larger aperture, which means a depth of focus of tens of nanometres, which means the wafer must be held flat and in focus to that tolerance across three hundred millimetres. The depth of focus, not the resolution, is what limits how small a feature can be made in practice — and every technique invented to get round it, from immersion to multiple patterning, is a way of buying resolution without paying the depth.

What is actually being traded

The relation looks like a coincidence of two formulas and is not. Both quantities are statements about the same object: the volume in which a converging wave is appreciably concentrated.

A lens collects light over a range of directions of half-angle θ\theta, so the transverse spread of what it can make is λ/sinθ\lambda/\sin\theta and the axial spread is λ/(1cosθ)\lambda/(1-\cos\theta), which for small angles is 2λ/θ22\lambda/\theta^2. Transverse goes as one power of the angle and axial as two, because the axial direction is where the collected directions all agree and the transverse is where they differ most. The asymmetry is geometrical and unavoidable: no lens collects light from behind the object.

That last observation is the one worth carrying. The reason a focus is much longer than it is wide is that a lens sees a cone rather than a sphere, and the axial resolution would equal the transverse only for an instrument collecting over the full solid angle. Some do — a sample between two opposed objectives, imaged by both and combined — and their axial resolution approaches their transverse. The trade drawn here is not a law of optics; it is a law about lenses that face one way.

What a photograph is instead

A converging lens making a real image. An object 2.36 focal lengths from a thin converging lens. The image sits where the construction rays cross, at 1.73 focal lengths, magnified -0.73×.
Fig. 5 The construction that produces the sharp plane in the first place: three rays whose intersection is the image, which exists for one object distance only. Everything in the preceding figures is about what happens to points that are not on this plane, which is almost all of them.

It is worth putting the ordinary construction back after all this, because it makes the situation clear. The three rays meet at a point, and that meeting is what an image is. There is nothing in the construction that goes soft gradually: a point off the plane produces three rays that meet somewhere else and pass the sensor at three different heights.

So a photograph is not a picture with a sharp part and a soft part. It is a picture in which every point except one plane’s worth is a disc, and the discs happen to be small enough not to be noticed over some range. The distinction matters when the tolerance changes — an image acceptable on a screen and unacceptable when enlarged — and it matters even more when something automatic is deciding what is sharp, because a measurement of sharpness is a measurement against a threshold that has to be stated.

Two ways of not caring where the focus is

If the depth cannot be increased, the alternative is to stop needing it, and there are two quite different ways of doing that. Both are worth knowing because they show what the constraint actually forbids.

The first is to record more than an image. A camera that captures the direction of every ray as well as its position holds the entire light field, and from it any focal plane can be computed afterwards — because the information about where a ray came from was never discarded. That does not beat the trade; it pays for it in resolution, since the sensor’s pixels are being spent on angle instead of position, and the recovered image at any one plane is coarser than a conventional one. The etendue argument that fixes what a fibre can accept is the same argument here: position and angle are one budget.

The second is to make a beam that has no focus to be out of. A ring-shaped pupil produces a beam whose central spot stays narrow over a much longer axial range than a filled pupil of the same width — the axial concentration is traded for a much larger fraction of the light going into rings. That is genuinely useful where a long focus matters more than efficiency, in machining and in some kinds of alignment, and it is again not a defeat of the trade but a payment: the light removed from the focus is still there, spread out where it does harm.

Both make the same point about what the constraint is. It is a statement about the volume a converging wave can occupy, and any arrangement that seems to beat it is spending something else — resolution, or light, or a dimension of the sensor.

Where the model stops

The lens is thin, paraxial and perfect. Real lenses have aberrations that vary with aperture, and a lens stopped down is not only diffraction-limited but better corrected — spherical aberration falls as the fourth power of the aperture — so the practical optimum aperture is set by three effects rather than two. The figures here describe a perfect lens, where the only competition is between geometry and diffraction.

The circle of confusion is a single number and a real detector is not. A sensor has pixels of a definite size, an anti-aliasing filter, and a demosaicing step, and what counts as resolved is a property of that chain. Quoting one tolerance is a summary of a system whose response falls off gradually.

Rayleigh’s quarter-wave is a rule of thumb with a number attached. It is the point at which the axial intensity has fallen to 81 per cent, which is a defensible place to draw a line and is not the only one; a criterion at 50 per cent gives a depth about half again as large. What is not a convention is the scaling with the inverse square of the aperture.

And the two accounts are added here by taking whichever is larger, which is not quite right. The true axial response is the convolution of the geometric and diffractive spreads, so near the crossover the depth is smaller than either estimate. The error is worst exactly where photographers work, around f/8, and it is why the optimum aperture is usually quoted from measurement rather than from arithmetic.

The tolerance that is not a choice

One case removes the arbitrariness entirely and is worth ending the argument on, because it shows what the essay’s central claim really amounts to.

If the detector is a single photon counter and the question is whether two point sources can be told apart, the tolerance is no longer a decision about viewing distance: it is set by how many photons arrive and by the statistics of counting them. A measurement with enough photons can separate sources far closer than Rayleigh’s criterion, and one with few cannot separate sources far apart — so the “resolution” of an instrument becomes a function of exposure, which no formula about apertures contains.

The same is true along the axis. The depth over which two planes can be distinguished depends on the signal, and the quarter-wave criterion is a convenient standard rather than a barrier. What does not depend on the signal is the shape of the axial response, which is the diffraction physics, and that is what the figures here compute.

So the honest summary is in two parts. The falloff of intensity with defocus is optics and is fixed. Where along that falloff the line is drawn is a decision — about a print size, about a pixel, or about a photon budget — and every quoted depth of field is that decision made silently. It is worth knowing which of the two a number is, and the way to tell is to ask what would change it: a better lens, or a longer look.

What the pictures cannot show

The axial-intensity figure draws the intensity at the very centre of the spot and says nothing about its shape. Out of focus, a diffraction-limited spot does not simply dim: it develops rings, and at certain defocus distances the centre goes dark while a ring carries the light. A measurement that samples only the centre therefore reports zero intensity at places where there is plenty of light, which is a real trap in autofocus systems that use exactly that measurement.

Nor does any figure here show what the out-of-focus light looks like, which is what a photographer cares about most. The shape of an out-of-focus point is the shape of the aperture — a fact used deliberately in the ringed pupil that images by diffraction alone — and it is a picture of the lens rather than of the subject. A defocused image is the aperture convolved with the scene, which is why the character of the blur is an optical designer’s decision and not the photographer’s.

Where the ladder goes next

The imaging ladder began with what a lens is doing, went through the mirror that cannot focus and the shape that can, the image that is a diffraction pattern twice and the condition a lens must meet to image an area rather than a point. This rung asks how thick the object plane is. The rungs after it: the confocal arrangement, where a pinhole in front of the detector rejects what is out of focus rather than resolving it; the depth of field of a wave, where a beam engineered to have no focus at all trades peak intensity for an axial range; and computational refocusing, where the whole light field is recorded and the plane is chosen afterwards.

The habit worth carrying away is to ask what tolerance a word is hiding. “In focus” is not a property of an optical system, and every number quoted for depth of field, depth of focus or sharpness is a statement about a threshold somebody chose. The physics is in the shape of the curve; the number is in the line drawn across it.

Part 5 of 6

This essay is one argument about Imaging. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

ApertureDiffraction limitFocal lengthImagingMagnificationNumerical apertureThe paraxial approximationReal and virtual imagesResolutionWavefront