How far apart two things have to be
Assumes: Where rays stop being enough, and a shadow acquires a bright centre · When two waves meet, they simply add
A perfect lens does not form a point image of a point source. It forms a small bright disc surrounded by faint rings, and the size of that disc has nothing to do with how well the lens was made. It is set by the fact that the lens has an edge.
That is the whole subject in one picture, and everything below is either where the numbers come from or what they cost.
Where the pattern comes from
Light passing through an aperture arrives at the image plane by many routes of slightly different length, and the amplitudes add with their phases. Straight ahead every route is the same length and everything adds; a little off-axis the routes across the aperture disagree, and at some angle they disagree by exactly enough that the contributions cancel completely.
Where the pattern comes from is a single aperture, and it is worth being clear that the resolution limit is a fact about one hole rather than about two objects. Light passing through an opening spreads, and the narrower the opening compared with the wavelength the wider the spread — not as a smooth blur but as a pattern with zeros in it. Two point sources seen through the same aperture each acquire that pattern, and asking whether they can be told apart is asking whether their two patterns can be told apart.
For a circular aperture the same calculation gives a first zero at rather than , and the extra factor is entirely geometry: the aperture’s width varies across it, so the cancellation is spread out and the first zero pushed further out. The number is the first zero of the Bessel function , which is 3.8317, divided by π. It is computed rather than tabulated in the figures here, by summing the Bessel series and bisecting on it, so the 1.22 is a result and not an input.
The criterion, and what it actually is
Rayleigh’s rule is that two points are just resolved when the maximum of one pattern falls on the first zero of the other. It is a convention. What makes it a good convention is the number that comes out of it.
Because it is a convention, it can be beaten, and modern practice beats it routinely. If the shape of the pattern is known exactly and the data are good, a pair can be fitted rather than looked at, and the separation recovered well below the criterion — the accuracy is then set by the signal-to-noise ratio rather than by the wavelength. Astrometry, single-molecule localisation microscopy and radio interferometry all depend on that, and all of them are measuring the position of a known shape rather than resolving an unknown one, which is a different and much easier question.
What cannot be beaten is the information: two points closer than the criterion produce a summed pattern that differs from a single point’s by an amount that falls very fast as the separation shrinks, and below some separation that difference is smaller than any noise present. The limit is soft and it is real.
The line with a slope of exactly minus one
Written as an angle, the criterion contains two quantities and no others.
That plot is the reason telescopes are large. Light-gathering is the usual explanation and it is the secondary one: collecting area goes as and matters for faint objects, but resolution goes as and matters for every object. The two together are why the useful figure of merit for a telescope goes as a high power of the diameter and why the cost does too.
It is also why the atmosphere is the enemy. A ground-based telescope larger than about 20 cm is not limited by its own aperture at all but by the turbulent cells of air above it, which scramble the wavefront and produce a blur of about an arcsecond — the aperture’s own limit for a lens the size of a jam-jar lid. Everything from Hubble to adaptive optics is an attempt to get back onto the line in that figure.
The other thing that stops an instrument is aberration, and it behaves in the opposite direction. A spherical mirror does not bring parallel rays to a point, and the spread is a geometric error with nothing to do with diffraction — but it gets worse as the aperture grows, where diffraction gets better. So every optical design has a diameter at which the two curves cross: below it the instrument is diffraction-limited and worth enlarging, above it aberration-limited and the extra glass is wasted.
The same formula with a different wave in it
The limit contains a wavelength, and nothing in the argument says the wavelength has to be light’s.
The electron microscope makes the point sharply because it changes only one thing. The optics are worse than a light microscope’s by every conventional measure: magnetic lenses have aberrations a glass lens designer would find embarrassing, and the usable aperture is a fraction of a degree rather than a hemisphere. It resolves a hundred thousand times better anyway, because the wavelength in the numerator is a hundred thousand times smaller.
Only in the last twenty years, with aberration correctors, has the instrument got close to being diffraction-limited rather than lens-limited.
The same reasoning runs in the other direction. Radio astronomy works at wavelengths of centimetres to metres, so a single dish of any buildable size has a resolution measured in arcminutes — worse than the naked eye. The fix is to make the distance between two dishes rather than the size of either, which is interferometry, and it is why radio astronomy went from the worst angular resolution in the subject to the best, with baselines that now include an orbiting antenna.
The microscope’s version, and the oil
An astronomer’s aperture is far away and the natural quantity is an angle. A microscopist’s aperture is right next to the specimen and the natural quantity is a distance, so the same limit is usually written the other way round:
where α is the half-angle of the cone of light the objective collects and is the refractive index of what it is collected through. The numerical aperture is doing exactly the job did before: it is a measure of how much of the diffracted light gets in. Detail in the specimen diffracts light through an angle; if that angle is outside the cone the objective accepts, the information about that detail never enters the instrument at all, and no processing afterwards can recover it.
That formulation explains the oil. Sine is at most one, so a dry objective cannot exceed NA 1 however wide its cone, and in practice tops out near 0.95. Filling the space between specimen and lens with an oil of index 1.5 multiplies the numerical aperture by 1.5, because the same physical cone now corresponds to a larger — and, just as importantly, because rays that would have been totally internally reflected at the cover slip now get through. The gain is about a third in resolution, and it is the difference between seeing a bacterium’s outline and seeing its flagellum.
What the objective has to collect is the other way of stating the same limit, and it is Abbe’s. Detail in a specimen behaves as a set of closely spaced sources, and the interference between them sends light out at angles set by their separation — the finer the detail, the wider the angle. An objective that does not accept those angles never receives the information at all, so the resolution limit is a statement about which diffraction orders get into the lens, and the oil is there to let the wide ones in.
A design that stops exactly at the limit
The best evidence that the limit is real is that things built by evolution sit on it.
A human pupil in daylight is about 2 mm across, which puts its diffraction limit at roughly 70 microradians — a minute of arc, or about 20/20 on the optician’s chart. The spacing of cone cells in the fovea is about 2.5 µm, which at the eye’s focal length of 17 mm subtends very nearly the same angle. That is not a coincidence: a retina with finer sampling would resolve nothing more, since the image arriving is already blurred to the Airy disc, and a retina with coarser sampling would waste the optics. Two entirely separate parts of an organ have been matched to a quantity neither of them controls.
The same equality shows up in a well-designed camera, where the pixel pitch is chosen to match the lens’s blur at its working aperture, and stopping down past about f/11 on a full-frame sensor makes the picture worse rather than better — the depth of field grows and the diffraction blur grows faster. It is one of the few places where a photographer meets a physical limit directly and has to trade against it.
The eye’s match is not perfect and the mismatch is instructive. In dim light the pupil opens to 7 mm, which improves the diffraction limit by a factor of three and a half — and acuity gets worse, because the eye’s own aberrations grow faster than the diffraction blur shrinks and because the retina switches to rods, which are wired together in groups and sample far more coarsely. The optics and the detector are matched at one working point and at no other, which is what an evolved design looks like when a physicist inspects it.
The aperture the air imposes
The remark above that a telescope larger than about twenty centimetres is limited by the atmosphere rather than by its own aperture deserves a number, because the number governs the design of every large telescope built.
Turbulence in the air mixes parcels at slightly different temperatures, and therefore at slightly different refractive indices, so a wavefront arriving flat from a star is crumpled by the time it reaches the ground. The size of the crumples is what matters, and it is quantified by one length: the diameter over which the wavefront is still flat to within about a radian of phase. It is called the Fried parameter, and at a good site in visible light it is ten to twenty centimetres.
Everything follows from comparing an aperture with that length. A telescope smaller than it sees an undistorted wavefront and performs at its own diffraction limit. A telescope larger than it sees a crumpled one, and produces not a blur but a speckle pattern — a swarm of roughly bright spots, each one the size the full aperture’s diffraction limit would give, dancing over a region of size and rearranging themselves every few milliseconds. A long exposure averages them into a disc about an arcsecond across, which is what “seeing” means, and it is the same figure for an eight-metre telescope as for a twenty-centimetre one.
Two scalings decide what can be done about it, and both favour longer wavelengths. The Fried parameter grows as , so at two micrometres in the infrared it is about a metre rather than fifteen centimetres; and the angle over which the distortion is the same — so that one reference star can correct another’s neighbourhood — grows the same way. That is why adaptive optics worked in the infrared for fifteen years before it worked in the visible: the number of actuators needed goes as and the correction has to be applied faster as shrinks, so the visible costs a hundred times more of both.
Adaptive optics measures the crumpled wavefront many hundreds of times a second and flattens it with a deformable mirror. Its usual difficulty is finding something bright enough to measure against, close enough on the sky that its distortion is the same as the target’s — which for the visible means within a few arcseconds, and there is rarely anything there. The fix is to make one: a laser tuned to the sodium transition, fired upward, excites a patch of the sodium layer ninety kilometres up and produces an artificial star wherever it is pointed.
There is also a way of recovering the resolution without correcting anything, and it is older. Expose for a few milliseconds and the speckles are frozen rather than averaged, and each speckle carries information at the full aperture’s resolution. Labeyrie showed in 1970 that averaging the power spectra of thousands of such frames recovers the diffraction-limited information about the object, even though no single frame is a picture of it. Speckle interferometry doubled the number of known binary stars before adaptive optics existed.
The hole too small to let anything through
The formula says the pattern from an aperture widens as the aperture narrows, and it is worth asking what happens in the limit — when the hole is much smaller than a wavelength.
The answer is not that the light spreads over a hemisphere and carries on. It is that almost none of it gets through at all. A hole of radius in a conducting screen, with much less than the wavelength, transmits a power that falls as the fourth power of — an extra factor beyond the geometric area, computed by Bethe in 1944, with the same fourth power that appears in Rayleigh scattering and for a related reason: a small hole radiates as a dipole rather than as an aperture.
The everyday demonstration is on every microwave oven. Its door carries a metal mesh with holes a millimetre or two across. Visible light, at half a micrometre, passes through freely — the holes are thousands of wavelengths wide and the mesh is simply a grid to look through. The microwaves are at 12.2 centimetres, so is about one part in a hundred, the fourth power of that is , and nothing measurable escapes. One screen, transparent to one wave and opaque to another, with the difference entirely a ratio of a hole to a wavelength.
The same argument is why a Faraday cage may be a mesh rather than a sheet, and why the specification for one names a maximum hole size rather than a material.
There is one celebrated exception, and it is worth the sentence. Arrays of subwavelength holes in a thin metal film were found in 1998 to transmit far more than Bethe’s expression allows — in some cases more than the holes’ total area would pass if the metal were not there at all. The extra light travels as a surface wave on the metal, is funnelled into the holes and re-radiated on the far side, so the holes are not acting independently and the single-aperture calculation does not apply. It is a good reminder that a limit derived for one hole is a limit about one hole.
What the limit costs
A microscope cannot see a virus in visible light. With immersion oil and the largest usable cone of angles, the best numerical aperture is about 1.4, and is about 200 nm at 550 nm wavelength. That is a hard floor for anything using an ordinary lens and ordinary illumination, and it sits an order of magnitude above the objects of most interest in cell biology.
Every image is a convolution. The point-spread function does not merely blur two points together, it blurs everything: a real image is the true scene convolved with the Airy pattern, and information at spatial frequencies above the cutoff is not attenuated but absent. Deconvolution can restore what is attenuated and cannot invent what is missing, which is the mathematical statement of why enhancing a photograph beyond its limit produces plausible fiction. The information was not degraded on the way through the instrument; it never entered, which is a stronger and less recoverable condition.
The limit propagates into the diagnosis. The dish of a radio telescope, the pixel pitch of a camera, the depth of focus of a lithography stepper and the size of a transistor are all set by the same expression, and the semiconductor industry’s whole history is a fight against it: shorter wavelengths, larger numerical apertures, immersion, and finally patterning a feature with two exposures because one cannot resolve it. The current generation illuminates at 13.5 nm, in vacuum, off mirrors made of alternating layers a few atoms thick, because no material is transparent at that wavelength and no lens can be used at all.
Two people, six years apart
The limit was arrived at twice, from opposite ends of the subject, by people with different purposes.
Abbe came to it commercially. Zeiss hired him in 1866 because the firm’s microscopes were made by trial and error and could not be improved systematically; Abbe’s answer, published in 1873, was that resolution is set by the diffraction of light at the specimen and the angular cone the objective collects, and that a lens designer’s job is therefore to maximise the numerical aperture rather than to polish more carefully. It turned optics from a craft into a calculation, made Zeiss the dominant instrument-maker in Europe within a decade, and produced the immersion objective as a direct consequence of the formula.
Rayleigh came to it astronomically, in 1879, from the problem of when a double star can be split — and he arrived at the criterion that carries his name and at the quarter-wave tolerance on optical quality in the same series of papers. His treatment is about incoherent point sources at infinity, Abbe’s about coherently illuminated periodic structure at a finite distance, and the two limits differ by a factor of about 1.2 for that reason. They are usually quoted as though they were the same statement, which is close enough for most purposes and wrong in exactly the situation — a modern microscope with controlled illumination — where the difference is worth money.
What the picture cannot show
The summed patterns above are for two incoherent sources — two stars, two fluorescent molecules — whose intensities add. Two coherent sources add amplitudes instead, and then the dip depends on their relative phase: two mutually coherent points at the Rayleigh separation can produce a summed pattern with a deep dip or no dip at all depending on a quantity the drawing does not contain. The standard treatment of resolution silently assumes incoherence, and every measurement made with a laser breaks that assumption.
The figures also assume a clear circular aperture. A telescope with a secondary mirror has an annular one, which narrows the central peak slightly and throws far more light into the rings; a rectangular aperture gives a cross-shaped pattern; and the diffraction spikes on every photograph of a bright star are the pattern of the vanes holding the secondary. Aperture shaping is a whole discipline, and it exists because the pattern is the Fourier transform of the aperture and can therefore be designed.
The domain of validity is: far field, monochromatic, incoherent, a clear aperture, and no aberration. Real instruments violate all five, and the formula survives as the thing they are compared against — which is the usual fate of a limit.
The ladder from here
Later rungs on this anchor: the Airy pattern derived properly, as the Fourier transform of a circular aperture; the optical transfer function, which restates resolution as a cutoff spatial frequency and is the form used in design; near-field microscopy, where the far-field assumption is broken deliberately and the limit does not apply; the super-resolution techniques that beat the criterion by making sources take turns to emit; and interferometry, where the aperture is synthesised from the separation between telescopes.
The neighbouring ladders are diffraction itself, which is where the pattern comes from, interference, which is the addition rule the pattern is built by, and matter waves, which is where a much shorter wavelength was found.
Part 2 of 8
This essay is one argument about Diffraction. The others:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
Airy patternAngular resolutionApertureDe broglie wavelengthDiffraction limitNumerical aperturePoint-spread functionRayleigh's criterion