Optics

Where rays stop being enough, and a shadow acquires a bright centre

Light going through a narrow gap spreads. No amount of ray tracing predicts it, the size of the spreading is set by one ratio, and taking that ratio to zero is exactly what the ray model is.

Every figure in this site’s optics essays draws light as rays: straight lines that bend at surfaces and otherwise go where they are pointed. The model is enormously successful, and it contains a claim that is simply false — that light passing through an opening continues in a beam the width of the opening, throwing a shadow with a sharp edge.

It does not. It spreads, by an angle that can be computed, and the smaller the opening the more it spreads. Narrowing a slit to sharpen a beam works down to a point and then reverses, and the reversal is not a defect of the apparatus.

Single-slit diffraction at three slit widthsIntensity against angle behind a single slit, evaluated from the integral across the aperture, for slits two, six and twenty wavelengths wide. A wide slit throws a nearly sharp shadow; a narrow one spreads light through a wide angle.-1-0.500.5100.20.40.60.81angle from the axis (radians)intensityslit = 2 λfirst zero 30.0°slit = 6 λfirst zero 9.6°slit = 20 λfirst zero 2.9°what rays predict
Fig. 1 Intensity against angle behind a single slit, for slits two, six and twenty wavelengths wide, evaluated from the integral across the aperture. The dashed outline is what the ray model predicts: a top-hat the width of the slit, and nothing outside it.

Adding up the aperture

The calculation needs one idea, which is already available: waves add.

Treat every point across the open slit as a source of a wave — Huygens’ construction, which is a statement that a wavefront’s future is determined by treating each of its points as a fresh emitter. Behind the slit, the disturbance in any given direction is the sum of the contributions from all those points.

Straight ahead, every contribution has travelled the same distance, so they all arrive in step and add to a maximum. Off-axis, the contribution from one edge of the slit has further to travel than the contribution from the other, by asinθa\sin\theta for a slit of width aa. When that extra distance is exactly one wavelength, something specific happens: the contributions from the top half of the slit are each exactly out of step with a partner in the bottom half, they cancel in pairs, and the total is exactly zero.

Two waves 180° out of step, and their sumTwo sine waves differing in phase by 180 degrees, drawn faintly, with their sum drawn solid. The sum is computed point by point.00.511.52-2-1012positionphase difference 180°they cancel completely
Fig. 2 Two contributions half a cycle apart, cancelling completely. The first dark direction behind a slit is where every point in the top half of the aperture has a partner in the bottom half in exactly this relationship — so the cancellation is exact rather than approximate, and it needs no integration to establish.

That gives the first dark direction with no calculus at all:

sinθ1=λa.\sin\theta_1 = \frac{\lambda}{a}.

Carrying the sum through properly gives the whole pattern, whose intensity is sinc2\mathrm{sinc}^2 of the phase difference across the aperture, with zeros wherever the pairing argument can be repeated on quarters, sixths and so on.

The single ratio λ/a\lambda/a controls everything. It is not the wavelength that matters, nor the slit width, but their quotient — which is why the same equation governs light through a pinhole, sound through a doorway, radio past a hill and electrons through a crystal.

The ray model, recovered as a limit

Single-slit diffraction at three slit widthsIntensity against angle behind a single slit, evaluated from the integral across the aperture, for slits two, six and twenty wavelengths wide. A wide slit throws a nearly sharp shadow; a narrow one spreads light through a wide angle.-1-0.500.5100.20.40.60.81angle from the axis (radians)intensityslit = 1 λno zeroslit = 3 λfirst zero 19.5°slit = 10 λfirst zero 5.7°slit = 40 λfirst zero 1.4°what rays predict
Fig. 3 The same pattern at four slit widths, from one wavelength to forty. As the slit widens the central lobe narrows toward the geometric beam, and the first zero moves from 90° to 1.4°.

Reading the figure from left to right is reading the ray model being born.

At a=λa = \lambda the first zero is at 90°: the light spreads through the entire forward hemisphere and the slit acts as a point source, with no memory of having had a width. At a=3λa = 3\lambda the first zero is at 19°. At a=20λa = 20\lambda it is at 2.9°, and at a=40λa = 40\lambda at 1.4°. The pattern is collapsing toward a beam.

So rays are not an approximation to be corrected. They are what the wave theory becomes when every dimension in the problem is large compared with the wavelength, in exactly the way Newtonian mechanics is what relativity becomes at low speed, and the small-angle pendulum is what a real one becomes at small amplitude. Visible light has a wavelength around half a micron, and the openings in ordinary life are millimetres and metres, so λ/a\lambda/a is between 10310^{-3} and 10610^{-6} and the spreading is invisible. That, and nothing else, is why the ray model was believed for two thousand years.

Sound gives the contrast for free, because its wavelengths are comparable to the objects around it. A 100 Hz tone has a wavelength of 3.4 metres, so a doorway is a fraction of a wavelength wide and the sound spreads through the whole room beyond. A 10 kHz tone has a wavelength of 34 millimetres, so the same doorway is twenty-five wavelengths across and the treble arrives in a beam. Standing outside a room and hearing bass but not treble is a diffraction measurement, performed daily and rarely noticed — and it is the same ratio that decides whether a wave reflects or wraps around an obstacle.

The same ratio decides why long-wave radio reaches around hills and behind buildings while a 100 MHz signal needs line of sight, and why a loudspeaker’s tweeter must be small: a 25 mm dome at 10 kHz has a/λ=0.74a/\lambda = 0.74 and radiates broadly, while a 300 mm cone at the same frequency has a/λ=9a/\lambda = 9 and beams the sound at whoever is directly in front.

Narrower in, wider out

The inverse relationship between slit width and spreading angle is worth stating on its own, because it is the whole content of the subject and it recurs far outside optics.

Confining a wave in space forces it to spread in direction. Not as a rule of thumb — as an equality with a constant in it. A slit of width aa produces a beam whose angular half-width is about λ/a\lambda/a, so the product of the spatial confinement and the angular spread is fixed at about λ\lambda, and no arrangement of apertures, lenses or mirrors reduces it.

Two sources 5 wavelengths apartCircular wavefronts from two sources, with the lines along which they arrive in step drawn through the pattern. Those lines are where the path difference is a whole number of wavelengths.sourcesourcesolid: waves arrive in stepbetween them they arrive opposed
Fig. 4 Two sources further apart, giving a finer pattern of reinforcing directions. The same inverse relation governs both figures: a wider arrangement produces finer angular structure, and a narrower one produces coarser.

Stated in terms of the wave’s spatial frequency rather than its angle, the relationship becomes ΔxΔk1\Delta x \, \Delta k \gtrsim 1: a wave restricted to a small region must be built from a broad range of spatial frequencies. That is a theorem about Fourier transforms, it applies to every wave that has ever been studied, and it is the same theorem that stops a short note from having a pitch with position in place of time.

The reason to put it that way is what happens next. Multiply both sides by Planck’s constant, and kk becomes momentum. The statement that a wave confined to a small region must contain a spread of spatial frequencies becomes the statement that a particle confined to a small region must have a spread of momenta — the uncertainty principle, with the same proof and the same content. Electrons fired at a narrow slit produce exactly the pattern at the top of this page, with λ\lambda their de Broglie wavelength, and the spreading is not a disturbance caused by the measurement. It is what waves do at apertures, and it was known to be what waves do at apertures for a century before anybody suspected electrons were involved.

What the spreading costs every instrument

The consequence that matters most is not the pattern behind a slit. It is that an instrument’s aperture is a slit, and its image is therefore a pattern rather than a point.

A converging lens making a real imageAn object 2.44 focal lengths from a thin converging lens. The image sits where the construction rays cross, at 1.69 focal lengths, magnified -0.69×.FFobjectimageu = 2.44 fv = 1.69 fmagnification -0.69
Fig. 5 A lens forming an image by three construction rays. Even with the grinding perfect and the rays crossing exactly, the image of a point is not a point — the lens has a finite aperture, and the light passing it spreads by the angle that aperture permits.

A perfect lens of diameter DD produces, from a distant point source, a small disc surrounded by faint rings — the Airy pattern, which is the circular version of the pattern above. Its angular radius is

θ=1.22λD,\theta = 1.22\,\frac{\lambda}{D},

with the 1.22 coming from the first zero of a Bessel function rather than from anything adjustable — the same function Airy introduced for the rainbow. Two point sources closer together than this cannot be told apart, because their patterns overlap into one blur. That is the Rayleigh criterion, and it is a limit on information rather than on engineering: no improvement in polishing, materials or alignment relaxes it.

The numbers are worth having, because they explain the sizes of things.

The eye, with a 3 mm pupil in daylight, has an Airy radius of about 0.8 arcminutes. Measured visual acuity is about 1 arcminute. The eye is therefore working within a small factor of the diffraction limit, and the density of cones in the fovea is matched to it — evolution has stopped adding detectors at the point where more would resolve nothing.

The optical microscope has a resolution of roughly 0.61λ/NA0.61\lambda/\mathrm{NA}, which with the best oil-immersion objectives is about 240 nanometres. That number is why cell biology stalled where it did, why electron microscopy exists — electrons have a far shorter wavelength — and why the fluorescence techniques that beat the limit had to do so by a trick that avoids resolving two sources at once rather than by resolving them.

The telescope gains resolution in proportion to diameter, which is the main argument for building large ones. It is also why radio astronomy, working at wavelengths a million times longer, must spread its dishes across continents to reach the resolution an optical telescope gets from a two-metre mirror.

And stopping a camera lens down reduces aberration and increases the Airy disc at the same time, so every lens has an aperture at which the two errors balance. The trade is between an error of manufacture and a consequence of the wave nature of light, and only one of them can ever be improved.

What it costs to see something smaller

The resolution limit is not merely a nuisance to be quoted; it is the constraint every technique for seeing small things has been built to evade, and each evasion has a price that can be stated.

Shorten the wavelength. An electron accelerated through 100 kV has a de Broglie wavelength of about 4 picometres, a hundred thousand times shorter than visible light, and an electron microscope’s resolution is correspondingly better. The costs are severe and specific: a vacuum, because electrons do not travel through air; a conductive specimen, or a coating, because charge accumulates; a dead specimen, because the beam deposits enough energy to destroy anything living; and lenses so aberrated that the practical resolution is a hundred times worse than the wavelength would allow. Electron optics trades a diffraction limit for an aberration limit, which is a different limit rather than no limit.

Increase the aperture. Resolution improves in proportion to DD, so a telescope four metres across resolves twice as finely as one two metres across — at four times the mirror area, considerably more than four times the cost, and with the atmosphere spoiling the result unless the instrument is above it or corrected in real time. Ground-based optical telescopes above about 20 cm are seeing-limited rather than diffraction-limited, so for two centuries the aperture bought light-gathering power and not resolution.

Get closer than the far field. Bringing a probe within a fraction of a wavelength of the object samples the pattern before it has had room to spread, which sidesteps the limit entirely. Near-field scanning microscopy reaches tens of nanometres with visible light this way. The cost is that it is a scanned contact measurement rather than an image — slow, restricted to surfaces, and unable to see inside anything.

Or refuse to resolve two things at once. The modern super-resolution methods work by making sure only one emitter is lit at a time, locating its blurred pattern to a precision far better than the pattern’s width, and repeating. The limit on distinguishing two simultaneous sources is untouched; what changed is that the problem was reformulated so that it never arises. The cost is time — thousands of frames per image — and a requirement that the sample be labelled with something that can be switched on and off.

The pattern across all four is worth naming. A diffraction limit is not a limit on precision; it is a limit on separating two things. Locating one isolated point source is limited only by how many photons are collected, which is why astronomers routinely measure positions to a small fraction of an Airy radius while being unable to split a close double at all.

The prediction made to destroy a theory

The history here is the best in optics, because the decisive experiment was proposed by someone trying to prove the opposite.

Two sources 3 wavelengths apartCircular wavefronts from two sources, with the lines along which they arrive in step drawn through the pattern. Those lines are where the path difference is a whole number of wavelengths.sourcesourcesolid: waves arrive in stepbetween them they arrive opposed
Fig. 6 Interference between two sources, with the reinforcing directions computed from the path difference. Fresnel’s account of diffraction is this construction with the two sources replaced by every point of a wavefront, and the sum taken as an integral rather than as a pair.

In 1818 the French Academy set a prize on diffraction, expecting to strengthen the particle theory of light. Fresnel submitted a wave treatment: Huygens’ construction, with the contributions from every point of the wavefront added with their phases — the calculation that produces the curves at the top of this page.

Poisson, on the judging committee and a committed opponent, worked through Fresnel’s mathematics and found what he took to be a fatal absurdity. Applied to a circular obstacle rather than an aperture, the theory predicts that the very centre of the shadow should be bright — because every point on the rim of the disc is exactly the same distance from the axis, so all their contributions arrive in phase and add. A bright spot in the middle of a shadow was, to Poisson, obviously ridiculous, and he presented it as a refutation.

Arago went and looked. The spot is there. It is faint and it is unmistakable, and it can be seen with a ball bearing, a pinhole and some patience.

Two things are worth taking from this beyond the anecdote. The first is that a theory is most convincingly confirmed by a prediction its own supporters did not want to make — Fresnel had not noticed the bright spot, and its discovery by a hostile reader is what made it evidence rather than a fitted parameter. The second is that the spot exists for any obstacle with a smooth rim, and it does not depend on the size of the disc, which is why it survived the objection that it must be a manufacturing artefact.

Where the model stops

The treatment above has its own domain, and three of its assumptions fail in useful places.

The screen is far away. Everything here is the far-field pattern, where the outgoing directions can be treated as parallel. Close behind the aperture the pattern is quite different and much messier, with the light showing structure on the scale of the aperture itself. The boundary between the two regimes is at a distance of about a2/λa^2/\lambda, which for a 1 mm slit in visible light is two metres — so an aperture that seems small can have a near field that fills a room.

The aperture is a hole in an opaque, infinitely thin screen. Real edges have thickness, conductivity and a finite absorption, and light interacts with the material of the edge. For microwaves passing a metal plate this matters a great deal, and the exact treatment is a substantial calculation that Sommerfeld solved for a perfectly conducting half-plane in 1896 — a boundary-value problem in the electromagnetic field rather than a statement about light.

The light is monochromatic. The pattern’s angular scale depends on wavelength, so white light produces overlapping patterns with a common centre and coloured fringes further out. That is why white-light diffraction looks like a coloured smear rather than a set of rings, for the same reason white-light interference fringes are coloured at the edges — and why a diffraction grating, which is the same physics with many apertures, works as a spectrometer.

And the wave is scalar. Treating light as a single oscillating quantity ignores polarisation. For apertures much larger than a wavelength that is harmless; for apertures comparable to one it is not, and the transmission through a subwavelength hole depends strongly on the polarisation relative to the hole’s shape. That dependence is what the whole field of plasmonics is built on.

The ladder from here

Later rungs: the double slit calculated properly, with the single-slit pattern as an envelope over the two-slit fringes — a fact that explains why some expected fringes are missing. The diffraction grating and its resolving power. The Airy pattern derived, and the 1.22 accounted for. Fresnel zones and the near field. The Arago spot computed rather than described. Diffraction by a straight edge, which produces fringes inside the bright region and a gradual fade into the shadow. X-ray diffraction from crystals, where the aperture is the lattice and the pattern reveals the structure. Electron diffraction, which showed that matter does this too. And Fourier optics, in which a lens performs a transform and the far-field pattern of an aperture is the transform of its shape — the statement that makes every result on this page a corollary of one theorem.

The connection worth carrying forward is that the rainbow’s supernumerary bows and the resolution of a telescope are the same calculation, and that both are invisible to a ray picture that gets every other feature of both right.