Optics

The path that does not change

Reflection, refraction and the angle of the rainbow are not three laws. They are one condition — that the optical path length is stationary — and the word stationary rather than shortest is the whole of what makes an elliptical mirror and a rainbow the same statement.

Assumes: The bend at the boundary, and what it is really about · What a lens is doing, and why three rays are enough

Snell’s law is a rule about angles at a surface, and it looks like a fact about surfaces. The law of reflection is a different rule about angles at a different kind of surface, and it looks like a different fact. Neither appearance survives an afternoon with a pocket calculator, because both are the same statement made twice: light travels along a path whose optical length is stationary with respect to small changes in it.

Snell's law, found by searching. Paths from a point in a medium of index 1 to a point in one of index 1.5, and the optical path length of each against where it crosses the boundary. The curve is that length; the marked point is its minimum, located by golden-section search and not by any use of a law of optics. The angles there are 55.80° and 33.46°, which satisfy n₁sin θ₁ = n₂sin θ₂ to 1.0e-8. Every other drawn path is longer, and the flatness of the curve near the bottom is why light is not fussy: a path a tenth of the way off costs almost nothing.
Fig. 1 Paths from a point in air to a point in glass, and the optical path length of each against where it crosses the boundary. The curve is that length — n1n_1 times the distance in the first medium plus n2n_2 times the distance in the second — and the marked point is its minimum, located by golden-section search. The angles there are 55.80° and 33.46°, which satisfy n1sinθ1=n2sinθ2n_1\sin\theta_1 = n_2\sin\theta_2 to one part in a hundred million. No law of optics was used to find that point.

That figure is the argument in its entirety. A quantity is defined for every conceivable path; the quantity is minimised numerically; the answer is Snell’s law to eight figures. Nothing about refraction was assumed and nothing about angles was imposed.

What optical path length is

The quantity being minimised is not distance and not quite time either, though it is proportional to time. It is

OPL=inii,\text{OPL} = \sum_i n_i \ell_i,

the geometric length of each segment weighted by the refractive index of the medium it crosses. Since the speed in a medium is c/nc/n, dividing by cc turns it into the travel time, so minimising one minimises the other. Working with OPL rather than time is a convenience with one real advantage: it makes the quantity a length, comparable directly with a wavelength, which is what the wave explanation needs.

The same crossing drawn as a ray is at the angle the search returns. Snell’s law is normally taken as the starting point and the bending deduced from it; here it is the output of a search over paths, which is a different logical position with the same content — and the difference matters, because the search generalises to cases where there is no boundary and no angle to write down.

The flatness of the curve near its minimum is worth noticing, because it does real work later. A path a tenth of the way off the optimum crossing point costs almost nothing in length. That is what “stationary” means quantitatively — the first derivative is zero, so the penalty for deviating is second order — and it is why light is not fussy and why a lens can be built out of a curved surface rather than requiring a point.

Snell's law, found by searching. Paths from a point in a medium of index 1 to a point in one of index 2.42, and the optical path length of each against where it crosses the boundary. The curve is that length; the marked point is its minimum, located by golden-section search and not by any use of a law of optics. The angles there are 59.47° and 20.85°, which satisfy n₁sin θ₁ = n₂sin θ₂ to 5.5e-8. Every other drawn path is longer, and the flatness of the curve near the bottom is why light is not fussy: a path a tenth of the way off costs almost nothing.
Fig. 2 The same search into diamond, whose index is 2.42. The minimum moves toward the entry point — light spends as little of its path as it can in the expensive medium — and the angles come out at 65.20° and 21.98°, again satisfying Snell to eight figures. The curve is also visibly steeper on one side than the other, which is the asymmetry a high index produces and the reason a diamond’s facets are cut to the angles they are.

The mirror, and the reason the word is not “shortest”

The same construction with both points on the same side of a surface gives reflection.

Equal angles, found by searching. Paths from a source to a detector by way of a flat mirror, and the total length of each against where it meets the mirror. The curve is that length; the marked point is its minimum, located by golden-section search and not by any use of a law of optics. The angles there are 48.01° and 48.01°, which satisfy equality to 5.3e-8 radians. Every other drawn path is longer, and the flatness of the curve near the bottom is why light is not fussy: a path a tenth of the way off costs almost nothing.
Fig. 3 Paths from a source to a detector by way of a flat mirror, with their total lengths plotted against where each meets it. The minimum is at 48.01° on both sides, equal to one part in 10710^7, and the equality is the law of reflection — obtained here, again, by searching. For a flat mirror the stationary path really is the shortest, which is why this is the case everyone remembers and why the general statement gets mislearned from it.

Take a curved mirror and the situation changes. The stationary path can be a maximum, and it can be neither, and there is one case in which the whole family is stationary together.

Every path the same length. An elliptical mirror of eccentricity 0.6, with a source at one focus. Every ray that leaves it arrives at the other focus, and every one of them travels exactly the same distance — 2.000000 in units of the semi-major axis, for all 400 sampled points, spreading by 8.9e-16. So the stationary path is not the shortest, or the longest, or unique: the whole family is stationary together. That is why the principle has to be stated with the word stationary, and it is also the design rule for every focusing surface there is.
Fig. 4 An elliptical mirror with a source at one focus. Every ray that leaves it arrives at the other focus, and every one travels exactly the same distance — 2.000000 in units of the semi-major axis, for all four hundred sampled points, with a spread of 9×10169\times10^{-16}. There is no shortest path and no unique one, and the plot on the right is a horizontal line because that is what “every path is stationary” looks like when it is drawn.

The ellipse is the reason the principle has to be stated in terms of a stationary point. It is also the design rule for every focusing surface in existence: an optical system forms an image at a point exactly when all the paths from object to image are equal in optical length, so that everything arriving there arrives in step.

The relation that follows from imposing equal path lengths is the lens equation, usually derived from refraction at two surfaces. Deriving it from Fermat instead needs no ray tracing at all: a lens images a point when every route between the two points takes the same time, and the thickness profile that achieves that is what a lens is.

A lens does this by making the paths through its thick middle geometrically shorter and optically longer by exactly the same amount.

A lens is an equaliser of paths rather than a bender of rays. The ray through the centre travels the shortest geometric distance and the most glass; the ray through the edge travels further and meets almost none — and the shape is chosen so that the two costs cancel. Every route takes the same time, which is why they all arrive together.

The rainbow is a stationary point with nothing on the other side

The most interesting stationary points are the ones that are not about a surface at all.

Rays through a raindrop. Parallel rays entering a spherical drop at different heights, refracting in, reflecting once from the back, and refracting out. The outgoing rays crowd together near one particular direction, and that crowding is the bow.
Fig. 5 Parallel rays entering a spherical drop at different heights, each refracting in, reflecting once inside, and refracting out, with every angle computed from Snell’s law. The exit directions are not spread evenly: they bunch up sharply near one particular deviation, because the deviation as a function of where the ray struck has a minimum there. Rays that struck at quite different places emerge in almost the same direction, which is what concentrates the light.

The rainbow’s 42° is a stationary point of the deviation, and the reason a rainbow is bright at that angle is the reason the minimum in the first figure was flat: near a stationary point, a whole range of inputs gives nearly the same output. Move away from the rainbow angle and the light thins out immediately; that is why the sky inside the bow is brighter than the sky outside it, and why there is a dark band between the primary and secondary bows where no ray of either kind can go at all.

The angle itself is a consequence of the refractive index of water, and because the index depends on colour the stationary point sits at a slightly different angle for each — which is the whole of why a rainbow has colours in the order it has them.

Computed for two colours, the stationary point moves: water’s index is 1.3435 at 400 nm and 1.3305 at 700 nm, so the minimum deviation is 40.6° for violet and 42.4° for red. The bow’s width is that difference and nothing else — the stationary condition is solved once per wavelength, and the colours separate because the index does, which is why a rainbow is a spectrum rather than a white arc.

Why a ray cannot survey its options

Stated as a principle about paths, Fermat’s rule has an air of the teleological about it that made it controversial for two centuries. Light appears to consider the alternatives and choose. It does no such thing, and the resolution is that the principle is a summary of something local.

Light is a wave. Every path from source to detector contributes an amplitude whose phase is 2π2\pi times the optical path length divided by the wavelength. Near a stationary point, neighbouring paths have path lengths that agree to first order, so their contributions are in phase and add. Everywhere else the phases run through the full circle as the path is varied and the contributions cancel. So the amplitude at the detector is built almost entirely from paths near the stationary one, and the “choice” is a cancellation.

That account makes three predictions the ray version cannot. It says the cancellation is imperfect when the alternatives are restricted — block all but a few paths with a screen and the light goes where the ray picture forbids, which is diffraction. It says the width of the bundle of paths that matter is set by the wavelength, so the ray approximation is good exactly when the apertures are large compared with it. And it says that a stationary maximum works as well as a minimum, since only the vanishing of the first derivative matters — which is the case the elliptical mirror shows and which no argument about light preferring speed can accommodate.

Past the critical angle no refracted path exists at all, and Fermat’s principle handles that gracefully: the search returns no stationary path in the second medium, so no ray goes there. That is a better answer than the ray picture’s, which has to report that a formula has no solution — here the absence of a path is the result rather than a failure of one.

When the index varies smoothly, the path is a curve

Everything so far has had a surface in it, and surfaces are the special case. If the refractive index varies continuously from place to place, the stationary path is no longer a set of straight segments but a smooth curve, bending toward the region of higher index.

The mechanism is the same delay argument arrived at continuously: the side of a wavefront in the slower region falls behind, so the front turns. A layer of hot air near a road surface has an index a few parts in 10510^5 lower than the air above it, which is enough to bend a nearly horizontal ray upward over a few hundred metres and produce the image of the sky on the tarmac that everyone reads as water. The same calculation applied to the whole atmosphere gives astronomical refraction: a star seen on the horizon is about half a degree higher than it really is, which is slightly more than the sun’s own diameter, so the sun is entirely below the horizon at the moment it appears to set.

The curvature is worth putting a number on because it decides a question that otherwise seems to need one. A ray bends with a radius of curvature n/nn/|\nabla n|; for the standard atmosphere that is about five times the radius of the Earth, so light does not follow the curvature of the planet and the horizon is real. Under a strong enough temperature inversion the ratio falls below one, light does follow the curvature, and the horizon effectively disappears — which is a rare but well-documented condition at sea and the origin of a good deal of maritime folklore.

The tolerance that follows from being stationary

The flatness at a stationary point has a consequence that decides how good a telescope has to be, and it is a consequence about errors rather than about paths.

If the optical path lengths from every part of an aperture to the image agreed exactly, the wave would arrive perfectly in step. They never do: the glass has a figure error, the mirror has a polishing error, the air has a temperature gradient. The question is how large a disagreement can be tolerated, and the answer is set by the wavelength rather than by any dimension of the instrument. Rayleigh’s rule, established empirically in 1879 and derivable from the interference calculation, is that a wavefront error of a quarter of a wavelength is barely noticeable and half a wavelength is ruinous.

For visible light that is 140 nanometres of path error across an aperture that may be eight metres wide — a relative accuracy of one part in 6×1076\times10^7. It is why telescope mirrors are specified in fractions of a wave rather than in microns, and why the figure of a large mirror is a harder problem than its size.

A parabolic mirror makes every path from infinity to the focus exactly equal, which is the tolerance that follows from being stationary: near a stationary point the path length changes only in second order, so a small error in the surface produces a much smaller error in the arrival time. That is why a mirror can be figured to a fraction of a wavelength and still work — and why it has to be, since second order of a large number is still large.

The Hubble Space Telescope’s mirror is the standing example: ground to a superb finish with the wrong conic constant, it was out by about two wavelengths, and it was useless until the error was cancelled by a matched one deliberately introduced into a corrector.

The same search, through a planet

The stationary-path condition is not a fact about light, and its largest application is to a wave that travels through rock.

An earthquake sends elastic waves through the Earth’s interior. Their speed varies with depth — it rises with pressure through the mantle, jumps at the core boundary, and changes again at the inner core — so a ray travelling from a source to a distant seismometer is refracted continuously and reflected at discontinuities, and its path is exactly a stationary path of the travel time.

That is the same calculation as the first figure on this page with a continuous index profile and a spherical geometry, and it is done in the same way: propose a speed profile, compute the stationary path from every source to every receiver, predict the arrival times, compare, adjust.

What comes out of that comparison is a map of the interior of a planet nobody can visit. The Earth’s core was found this way in 1913 — a shadow zone appeared between about 103 and 143 degrees from every large earthquake, where the direct wave was absent, and the only structure that produces such a zone is a sphere of markedly lower wave speed occupying the middle of the planet. The inner core was found the same way in 1936, from faint arrivals inside the shadow zone that could only be rays refracted by a further boundary deeper down.

Both are inferences from arrival times alone, made by a search over stationary paths, with no direct observation of anything. The radius of the Earth’s core is known to a few kilometres, and it was measured with clocks and a variational principle.

The modern version inverts the whole problem at once: millions of travel times from thousands of earthquakes, solved simultaneously for a three-dimensional speed structure, giving a tomographic image of the mantle in which cold sinking slabs and hot rising plumes are visible. The physics in the innermost loop is the search this essay’s first figure performs, run some very large number of times.

The path is the same both ways

One property of a stationary path is easy to overlook and has consequences reaching well outside optics: it does not know which end is the source.

The optical path length between two points is a property of the geometry, and it is symmetric — swapping the endpoints changes nothing about the sum. So a path that is stationary from A to B is stationary from B to A, and every ray path is traversable in either direction.

That is the geometrical-optics statement of reciprocity, and it has an immediate practical form. If a source at A can be seen from B, then a source at B can be seen from A, along the same path — which is why a periscope works in both directions, why an observer who can see a mirror is visible in it, and why the one-way mirror in an interrogation room is not one: it is a partial reflector with a very large difference in illumination on the two sides, and turning the lights on in the dark room reverses which side sees which.

The same symmetry, taken up one level to amplitudes rather than paths, is the reciprocity theorem that forbids any arrangement of lenses, mirrors and ordinary media from passing light one way and blocking it the other. Fermat’s principle is where that impossibility begins: the paths are the same, so nothing built out of choosing paths can distinguish the directions.

It also gives a useful check on any ray calculation. Trace a ray from object to image, then trace it back, and it should return to where it started; a computation that does not is wrong, and the failure is often easier to spot in the reverse direction than in the forward one. Optical design software runs rays both ways for exactly that reason.

What it grew into

Fermat stated the principle in 1662, in a letter, to derive Snell’s law and to argue against Descartes’ derivation, which had assumed light travels faster in the denser medium. The two derivations give the same law and opposite predictions about the speed, and settling which was right took Foucault’s measurement of the speed of light in water in 1850. Fermat was right and had waited a hundred and eighty-eight years.

The larger consequence was structural. A law expressed as “the path that makes this quantity stationary” turned out to be a form that other physics could be forced into, and Maupertuis, Euler, Lagrange and Hamilton spent the eighteenth and nineteenth centuries doing exactly that for mechanics. The result — the principle of least action — is the form in which mechanics, field theory and general relativity are all now written, and Hamilton’s route to it was explicitly the optical analogy: he wrote the mechanics of a particle as the optics of a wave and then, a century early, could not say what the wave was.

Schrödinger could. The relationship between the ray and the wave here — a variational principle emerging as the short-wavelength limit of an interference calculation — is exactly the relationship between classical mechanics and quantum mechanics, and the correspondence between them is the same mathematics with different names.

What the picture cannot show

The plots on this page draw the optical path length against one parameter — where the ray crosses a surface. A real path can be varied in infinitely many ways, and being stationary means being stationary under all of them. The one-parameter picture is a slice through a function of infinitely many variables, and a point can be a minimum along one slice and a maximum along another.

The figures also draw a single ray, which is a construction rather than a thing. A ray is the normal to a wavefront, defined where the wavefront is smooth over many wavelengths; where it is not — at a focus, at an edge, in a caustic — the ray picture predicts infinite intensity and the wave picture predicts a finite one with fringes. Every bright caustic on the bottom of a swimming pool is a place the ray drawing fails and the failure is visible.

The domain of validity: media whose properties change slowly compared with the wavelength, apertures large compared with the wavelength, and no interest in amplitudes or phases. Fermat’s principle gives the path and says nothing whatever about how much light goes along it. That is a separate calculation entirely, and for the flat boundary above it gives 4% reflected at normal incidence — a number the whole of this page is silent about.

The silence is worth one more sentence, because it is the same silence in every variational principle. A rule that identifies a path by making a quantity stationary has, by construction, thrown away everything except the path: the amplitude, the phase, the polarisation and the fraction that went the other way are all quantities the stationary condition never mentioned and cannot be interrogated about. Fermat’s principle is therefore complete as a theory of where and empty as a theory of how much, and every practical optical design needs both halves.

The ladder from here

Later rungs on this anchor: the derivation of the eikonal equation, which is Fermat’s principle in differential form and the bridge to the wave equation; light in a medium of continuously varying index, where the path is curved and mirages and gradient-index lenses follow; the caustic as the envelope of a family of stationary paths, and why it is bright; the principle of least action, with the optical-mechanical analogy written out; and the path integral, in which the cancellation argument above is taken literally and stops being an analogy.

The neighbouring ladders are refraction, which is the local rule this principle produces, the rainbow, which is a stationary point of a different quantity, and diffraction, which is what happens when the cancellation the principle relies on is prevented.

Part 1 of 4

This essay is one argument about Fermat. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Fermat's principleOptical path lengthRainbow angleReflectionRefractive indexSnell's lawStationary pointVariational principle