Optics

The fringe and the spectrum are one measurement

An interferometer with no prism and no grating in it measures a spectrum, because what it records as the path difference is scanned is the Fourier transform of the source's spectrum. Coherence length and linewidth are the same fact stated twice, and the resolution is bought in centimetres of travel.

Assumes: How far a wave can remember · Why two lamps never interfere

Split a beam, delay one half, recombine, and measure the power. Do it again at a slightly different delay, and again, over a range of path differences, and what comes out is a curve — the interferogram. It contains no wavelengths, no colours and no dispersion. It is a single power measured against a single distance.

It is also the entire spectrum of the source, in a form that needs one arithmetic operation to read.

The spectrum, and what the interferometer records instead. On the left, a source spectrum: 1 line near 2000 reciprocal centimetres. On the right, what a detector behind a two-beam interferometer reads as the path difference is scanned — the interferogram. It is the cosine transform of the spectrum, so the two panels carry exactly the same information and neither is more fundamental. The fast oscillation is the mean wavenumber; the envelope that decays over about 0.133 centimetres is the reciprocal of the linewidth, which is the coherence length; and where two lines are present, the beat between them is the splitting. Nothing disperses anything anywhere in the instrument.
Fig. 1 On the left, a source spectrum: one line near two thousand reciprocal centimetres. On the right, what a detector behind a two-beam interferometer reads as the path difference is scanned. The fast oscillation is the mean wavenumber; the envelope that decays is the reciprocal of the linewidth. The two panels carry the same information and neither is more fundamental.

The reason is one line of algebra. Each wavelength in the source interferes with itself, giving a cosine in the path difference at its own spatial frequency. Different wavelengths do not interfere with each other — two independent sources never do — so the powers add. The total is a sum of cosines weighted by the spectrum, which is the definition of a cosine transform.

Coherence length is a linewidth

The envelope in that figure is the part with a familiar name. As the path difference grows, the cosines at different wavelengths fall out of step with each other and their sum washes out; the distance over which that happens is the coherence length, and it is the reciprocal of the spread of wavenumbers in the source.

Stated as a transform, that ceases to be two facts. A narrow line has a long envelope and a broad one has a short envelope, for the same reason a narrow function has a broad transform — and the coherence length is not a separate property of the light to be measured alongside its spectrum. It is the width of the transform of the spectrum, which is to say the spectrum measured in different units.

The spectrum, and what the interferometer records instead. On the left, a source spectrum: 2 lines near 2002 reciprocal centimetres. On the right, what a detector behind a two-beam interferometer reads as the path difference is scanned — the interferogram. It is the cosine transform of the spectrum, so the two panels carry exactly the same information and neither is more fundamental. The fast oscillation is the mean wavenumber; the envelope that decays over about 0.455 centimetres is the reciprocal of the linewidth, which is the coherence length; and where two lines are present, the beat between them is the splitting. Nothing disperses anything anywhere in the instrument.
Fig. 2 The same instrument on a source with two close lines. The envelope now beats: the two cosines drift in and out of step, and the beat period is the reciprocal of the splitting. A doublet four reciprocal centimetres apart shows a null every quarter of a centimetre of path difference, and counting the nulls measures the splitting without ever resolving the lines.

The beating case is the one that makes the identification concrete. Michelson used exactly this in the 1890s to measure line splittings far below what his gratings could resolve: he did not transform anything, because the arithmetic was impractical by hand, but he read the visibility against path difference and inferred the line structure from where the fringes disappeared and returned. The measurement of the cadmium red line’s fineness, which underwrote the definition of the metre for half a century, was made this way.

The resolution is a distance

The spectrum, recovered from a scan that had to stop. The spectrum recovered by transforming the interferogram back, for scans of 0.1, 0.4, 2 centimetres, all normalised to the same peak. A scan of infinite length would return the source spectrum exactly. A scan that stops multiplies the interferogram by a rectangle, which convolves the recovered spectrum with a sinc — so every line acquires a width of about 0.6 divided by the scan length, and rings with negative side lobes that are not in the source at all. Resolution is therefore bought in centimetres of travel: to resolve a tenth of a reciprocal centimetre the mirror must move six. Nothing about the light or the detector enters that number.
Fig. 3 The spectrum recovered by transforming the interferogram back, for three scan lengths. An infinite scan would return the source exactly. A scan that stops multiplies the interferogram by a rectangle, which convolves the spectrum with a sinc — so every line acquires a width of about 0.6 divided by the scan length, and rings with negative side lobes that are not in the source.

Here is the instrument’s defining number. Resolution is bought in centimetres of mirror travel, and in nothing else: to resolve Δσ\Delta\sigma reciprocal centimetres, the path difference must reach about 0.6/Δσ0.6/\Delta\sigma. Nothing about the source, the detector, the wavelength or the optical quality substitutes for it.

That is a very different economy from a grating instrument, where resolution is bought in the number of grooves illuminated and is proportional to the size of the grating and the order used. Both are ultimately statements about the largest path difference the instrument creates between two parts of the same wavefront, which is the deep reason a grating’s resolving power is the number of lines times the order: it is the total path difference across the grating measured in wavelengths. The interferometer’s advantage is that a path difference is easy to make large by moving a mirror, where a grating has to be physically that big.

The consequences are the standard advantages of the technique. All wavelengths are measured at once rather than one at a time, so for a fixed measuring time the signal-to-noise is much better when detector noise dominates. There are no slits, so all of the light collected gets to the detector — the etendue is limited only by the beam rather than by an entrance aperture. And the wavenumber scale is set by the mirror’s position, which can be measured with a laser, so the instrument is self-calibrating to a precision no grating can approach.

The visibility, read as a spectrum

What is lost is exactly what is recorded. Fringe visibility against the distinguishability of the record left in the environment, for six couplings between the interferometer and a marker. The points lie on the quarter circle V² + D² = 1, computed here to a part in 10¹² — the visibility from the output probabilities with the marker traced out, and the distinguishability from the overlap of the marker's two states, with nothing shared between the two calculations. The relation is the quantitative form of complementarity, and it is stronger than the usual statement: interference is not lost because something was disturbed, and not lost only when a measurement is made. It is lost exactly to the extent that the environment could in principle say which way the particle went, whether or not anybody looks at the environment. That is why decoherence is a matter of correlation rather than of disturbance: the coherence has not been destroyed but relocated, into a correlation between the particle and something else.
Fig. 4 Fringe visibility as a general measure of how much interference survives. In the spectroscopic case the quantity that erodes it is the spread of wavelengths, and the visibility against path difference is the modulus of the transform of the spectrum — so a visibility curve is a spectrum with its phase discarded.

Before the arithmetic was practical, the whole technique lived in one number: how deep the fringes were at each setting. It is worth seeing what that number is in the transform language, because the answer explains both what Michelson could do and what he could not.

The visibility at a given path difference is the modulus of the transform of the spectrum. So a visibility curve carries the spectrum’s magnitude and throws away its phase — which means a symmetric spectrum can be recovered from it and an asymmetric one cannot. Two lines of equal strength give a visibility that goes to zero and returns; two of unequal strength give one with a non-zero minimum, and the ratio of the minimum to the maximum gives the ratio of the strengths. Both were measurable by eye.

What was not recoverable was any spectrum without a symmetry to lean on, and that is precisely the case an infrared chemist has. The step from visibility to interferogram — from a modulus to a signed record — is the step that turns a technique for measuring line doublets into a technique for measuring anything, and it needed a detector that records a signal rather than an eye that judges a contrast.

The same statement in the other direction is worth having: a measurement of how much interference survives is always a measurement of a spread, whether the spread is in wavelength, in position or in whatever a which-path marker is recording. The visibility is the transform’s modulus in every case, and the thing that erodes it is whatever variable the transform is over.

What the window costs

Ringing, or resolution — the window decides which. The same 0.4 centimetre scan transformed two ways: cut off abruptly, and tapered smoothly to zero at the end. The abrupt cut resolves the finer detail and surrounds every line with negative side lobes reaching a fifth of the peak, which are an artefact of the window and can be mistaken for absorption. The taper removes them almost entirely and widens every line by about half. Neither is more correct: they are two choices about what to do with information the scan does not contain, and the choice is made according to whether a weak line beside a strong one matters more than the last few per cent of resolution.
Fig. 5 The same scan transformed two ways: cut off abruptly, and tapered smoothly to zero at the end. The abrupt cut resolves the finer detail and surrounds every line with negative side lobes reaching a fifth of the peak. The taper removes them and widens every line by about half.

Truncating a scan is multiplying the interferogram by a rectangle, and multiplication in one domain is convolution in the other. The rectangle’s transform is a sinc, whose side lobes reach 22 per cent below zero — so the recovered spectrum has dips beside every line that look exactly like absorption and are not.

This matters because the artefact is the same shape as the signal. In an absorption spectrum, a negative feature beside a strong band is precisely what a weak absorber would produce, and mistaking the window’s ringing for a chemical species is a real and common error. The remedy is to taper the interferogram to zero at the end, which removes the lobes at the cost of widening the lines — and the choice between the two is made by asking which mistake is worse.

Neither choice is more correct. Both are decisions about what to do with information the measurement does not contain: the interferogram beyond the end of the scan was not recorded, and a transform has to assume something about it. A sharp cut assumes it is zero; a taper assumes it dies away smoothly. Neither assumption is knowledge, and the difference between the results is the size of what is being assumed.

That is a general point about any transform of a truncated record, and it is why the same vocabulary — windows, apodisation, side lobes, resolution — appears wherever one is taken.

The trade nobody escapes

Reading the whole thing as a transform makes the instrument’s limits obvious rather than empirical.

The scan length sets the resolution. The sampling interval sets the highest wavenumber that can be measured without ambiguity, because a cosine sampled too coarsely is indistinguishable from a slower one. The number of samples is the product of the two, so a spectrometer covering a wide range at high resolution needs a great many samples and a long scan — and that is why such an instrument’s scan takes seconds rather than milliseconds, and why fast measurements are made at low resolution.

And the interferogram’s peak, at zero path difference, carries the total power while everything about the structure of the spectrum is in the small oscillations far from it. A detector must therefore have the dynamic range to record a large central value and small wings on the same scale, which is the technique’s real difficulty and why it arrived only when detectors and digitisers were good enough.

The three advantages, and what they cost

The technique’s reputation rests on three specific advantages over a dispersive instrument, and each has a corresponding cost worth stating alongside it.

All wavelengths at once. A grating instrument looks at one resolution element at a time; the interferometer measures every one throughout the scan. Where the noise is in the detector rather than in the light, that is a signal-to-noise gain of the square root of the number of elements — a factor of thirty or more in the infrared. Where the noise is in the light, as it is in the visible with a photon-counting detector, the advantage vanishes and reverses, because every element’s noise is now spread across all of them. That is precisely why the technique dominates the infrared and is rare in the visible.

No slit. A dispersive instrument’s resolution is set by an entrance slit, so raising the resolution throws light away. An interferometer’s resolution is set by the scan, and its aperture is limited only by the off-axis path-difference smearing mentioned below — which for a given resolution is a much larger solid angle. The gain is roughly two orders of magnitude in throughput and is the reason the technique reaches sources a grating cannot see at all.

A laser-referenced scale. The mirror position is measured by counting fringes of a stabilised laser in the same interferometer, so every wavenumber in the spectrum is tied to that laser’s frequency. The accuracy is parts in 10810^8, against parts in 10510^5 for a good grating, and it is why line positions in reference databases come from this kind of instrument.

The costs are the mirror’s mechanical precision, the dynamic range problem, and the fact that everything is computed rather than observed — an artefact in the transform is not visible as an artefact, where a scratch on a grating usually is.

Where the same transform is doing the same job

The relation between a scan and a transform is not a fact about interferometers, and it is worth listing the places it recurs because recognising it saves rederiving the same limits.

A diffraction pattern is the transform of an aperture, and the finite aperture is the window: that is why a slit of finite width gives a pattern with side lobes, why a thousand slits resolve better than two, and why apodising an aperture — making its transmission fall off smoothly at the edge — suppresses the rings round a star’s image at the cost of widening it. The vocabulary is identical because the mathematics is.

A pulse’s spectrum is the transform of its shape in time, so a short pulse has a wide spectrum with the same reciprocal relation and the same window effects. Truncating a pulse abruptly puts side lobes in its spectrum in exactly the way truncating an interferogram does.

And a diffraction pattern from a crystal is the transform of its electron density, with the finite crystal as the window and the resolution limited by the highest angle at which reflections can be measured — which is a statement about the largest path difference the experiment creates, again.

In each case the pair of conjugate variables is different and every conclusion transfers: resolution is one over the range scanned, artefacts come from the window, and information outside the record has to be assumed rather than measured. That is why the habit at the end of this essay is worth more than any of the particular numbers in it.

Where the model stops

The two beams are assumed to divide the amplitude equally and recombine perfectly. A real beamsplitter’s ratio varies with wavelength, so the instrument’s response is not flat and every spectrum has to be divided by a background measured with no sample. That division is why an infrared spectrum is quoted as a transmittance rather than as an intensity.

The source is assumed stationary over the scan. A source whose brightness changes while the mirror moves puts a spurious signal into the interferogram that appears in the spectrum at whatever wavenumber corresponds to the rate of change — so a flickering lamp produces a line that is not there. Rapid scanning and modulation exist to move that artefact out of the band of interest.

The transform assumes the interferogram is symmetric about zero path difference and it never is. Dispersion in the beamsplitter and any error in locating the zero introduce a phase that varies with wavenumber, and the recovered spectrum is complex rather than real. Correcting it is a standard step and it is the reason a low-resolution double-sided scan is taken alongside the long single-sided one.

And nothing here is quantum. The whole argument is about adding amplitudes and squaring, which is classical wave optics, and it says nothing about photon statistics. A measurement of the second-order coherence asks a different question of the same light and is not obtainable from the interferogram at all.

Reading a real infrared spectrum

It is worth following what happens to one measurement from mirror to answer, because every step is one of the effects above and the order matters.

The instrument records a voltage against mirror position, sampled at every fringe of a reference laser — typically sixty thousand points over a couple of centimetres. That record is divided by nothing and corrected for nothing: it is the raw interferogram, and it looks like a single enormous spike with small wings, which is what a broadband source gives.

The wings are then multiplied by a taper, the phase is corrected using a short double-sided section around the spike, and the transform is taken. What comes out is a single-beam spectrum: the source’s emission, times the beamsplitter’s response, times the detector’s response, times the absorption of everything in the beam including the air. Nothing about it looks like a chemical spectrum.

A second measurement with no sample gives the same product without the sample’s absorption, and dividing one by the other leaves the transmittance. Every instrumental factor cancels, which is the reason the technique is quantitative at all, and it cancels only to the extent that the instrument has not drifted between the two measurements — which is why a background is taken every few minutes and why the two scans are as close together in time as the experiment allows.

The step that surprises people is the last one. The published spectrum is a ratio of two transforms of two scans, and there is no point in the chain at which anything resembling a spectrum was observed.

What the pictures cannot show

The interferogram is drawn against path difference and the instrument scans in time, so the horizontal axis of the real measurement is a mirror position inferred from a reference laser’s fringes. Everything about the accuracy of the wavenumber scale lives in that inference, and none of it is visible in a figure that plots against the path difference as though it were known.

Nor do the figures show what happens off-axis. Rays through the interferometer at an angle have a slightly different path difference, so a finite aperture smears the interferogram — and the smearing grows with path difference, which limits the resolution by exactly the mechanism the scan length was supposed to control. The trade between the aperture and the resolution is the instrument’s real design constraint and it is a geometrical effect that a one-dimensional picture cannot contain.

Where the ladder goes next

The coherence ladder began with why two lamps never interfere, went through how far a wave can remember, the fringe that measures a star, the grain that is in the light and the correlation that survives what the phase does not. This rung says that the first two of those are the same measurement as a spectrum. The rungs after it: the spatial version, where the transform is over a separation rather than a delay and returns the source’s shape rather than its spectrum; heterodyne detection, where the transform is performed by mixing rather than by scanning; and the coherence of a light field described completely, as a function of two positions and two times, of which everything in this ladder is a slice.

The habit worth carrying away is to ask what a measurement is the transform of. An instrument that scans one variable and records a power is almost always measuring the transform of a distribution in the conjugate variable, and recognising which pair is involved usually says immediately what sets the resolution, what sets the range, and what the artefacts will look like.

Part 6 of 6

This essay is one argument about Coherence. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

BandwidthCoherenceFourier transformInterferenceLinewidthPath differenceResolving powerSpectrumVisibilityWavefront