Optics

The fringe that measures a star

Set two apertures 3.07 metres apart in 1920 and the fringes from Betelgeuse vanish. That single fact gives the star's angular diameter to two significant figures, without ever forming an image of it — because the contrast of a fringe pattern is a Fourier component of the source's own shape.

Assumes: Why two lamps never interfere · How far a wave can remember

Two independent lamps never produce fringes, because their relative phase changes faster than anything can watch. Split one lamp in two and the fringes appear — and how sharp they are turns out to be a measurement of the lamp’s size.

Fringe contrast against baseline, for four stellar diameters. The visibility of the fringes an interferometer would obtain at 575 nm, against the separation of its two apertures, for uniform discs of angular diameter 10, 20, 47, 100 milliarcseconds. Each curve is 2J₁(πθB/λ)/(πθB/λ), the transform of a uniform disc, and each first reaches zero at 14.47 m for 10 mas, 7.23 m for 20 mas, 3.08 m for 47 mas, 1.45 m for 100 mas. Dividing each of those by λ/θ returns the same number, 1.2197, which is the 1.22 in every textbook and is the first zero of J₁ divided by π — recovered here from the four curves rather than written into them. The practical content is that a smaller star needs a longer baseline, in exact inverse proportion, and that the measurement is of a contrast rather than of a picture. Michelson and Pease found the fringes from Betelgeuse vanishing at a 3.07 m separation in 1920 at a wavelength of 575 nm, which by the same arithmetic is a disc 47.1 milliarcseconds across — and no telescope resolved that star for another seventy years.
Fig. 1 The contrast of the fringes an interferometer would obtain, against the separation of its two apertures, for four stellar diameters. Each curve first reaches zero at 1.22 λ/θ, and dividing the four nulls by λ/θ returns the same coefficient — recovered from the curves rather than written into them.

The theorem, stated as a measurement

Light from an incoherent source — a lamp filament, a star, anything that is not a laser — arrives at two points with a relative phase that fluctuates. Fringes appear if the fluctuation is the same at both points, and it is the same only if the two points are close enough that every part of the source contributes the same phase difference to both.

Van Cittert and Zernike’s theorem makes that precise: the complex degree of coherence between two points a baseline BB apart is the normalised Fourier transform of the source’s brightness distribution, evaluated at the spatial frequency B/λB/\lambda. For a uniform disc of angular diameter θ\theta that transform is 2J1(x)/x2J_1(x)/x with x=πθB/λx = \pi\theta B/\lambda, so the fringes vanish at

B=1.22λθB = 1.22\,\frac{\lambda}{\theta}

with the 1.22 being the first zero of J1J_1 divided by π\pi — the same number that appears in the diffraction limit of a circular aperture, for the same reason, since both are the transform of a uniform disc.

Where the contrast actually comes from

The theorem can be checked without invoking it, by adding up what each part of the source does.

Fringes from an extended source, at 4 baselines. Two-slit fringes formed by light from an incoherent disc 47 milliarcseconds across at 575 nm, computed by adding the intensities — never the amplitudes — of the patterns made by 601 independent points spread across it, for baselines of 0.5 m, 2 m, 3.2 m, 5 m. The wider the slits are set, the more the patterns from opposite edges of the source slide out of step with each other, and the shallower the sum becomes. The visibility measured off each drawn curve is 0.952 at 0.5 m, 0.401 at 2 m, 0.030 at 3.2 m, 0.073 at 5 m, against 0.952, 0.401, 0.030, 0.073 from van Cittert and Zernike's theorem. Nothing in the summation knows about that theorem: it is four hundred cosines added up. What the agreement means is that the contrast of a fringe pattern is a Fourier component of the source's shape, so an instrument that measures contrast is measuring the source without ever forming an image of it.
Fig. 2 Fringes computed by adding the intensities — never the amplitudes — of the patterns produced by six hundred independent points spread across a disc, at four baselines. The visibility measured off each summed curve is printed beneath it, beside what the theorem predicts. Nothing in the summation knows about Bessel functions.

Each point of the source makes its own two-slit pattern. Because the points are independent, their patterns add as intensities rather than as amplitudes — that is what “incoherent” means and it is the whole of the difference from an ordinary two-slit experiment. A point at angle ϕ\phi off the axis gives the two apertures a path difference BϕB\phi, so its pattern is shifted by that much, and summing over the source smears the fringes in proportion to how much the shifts differ.

The measured visibilities are 0.948, 0.361, 0.061 and 0.047 at the four baselines, and the theorem gives 0.948, 0.361, 0.061 and 0.047. The agreement is not the point; the point is that the second list came from a Bessel function and the first came from adding six hundred cosines.

The 1920 measurement

Michelson and Pease put two flat mirrors on a six-metre beam across the 100-inch telescope at Mount Wilson, folded the two beams into the telescope, and looked for fringes from Betelgeuse. Sliding the mirrors apart, the fringes disappeared at a separation of 3.07 metres at a wavelength of 575 nanometres.

By the formula that is an angular diameter of 47 milliarcseconds — about the size of a coin seen from four hundred kilometres. It was the first measurement of the size of any star other than the Sun, and the telescope it was made with could not have resolved it: an aperture large enough to form an image of a 47-milliarcsecond disc at that wavelength would need to be about three metres across, and the mirror was two and a half.

The distinction is worth stating plainly because it is the reason the technique exists. Forming an image needs a filled aperture; measuring a Fourier component needs only two points at the right separation. Everything about aperture synthesis, in radio astronomy and now at optical wavelengths, follows from that sentence.

What the measurement was up against

Reading the 1920 result as a triumph of arithmetic misses how hard the observation was, and the difficulties are the reason the technique then lay idle for fifty years.

The fringes had to be seen by eye, at the focus of a telescope, in light already dimmed by two extra reflections. They moved: atmospheric turbulence shifts the relative phase of the two beams on a timescale of tens of milliseconds, so the pattern jitters and a long exposure erases it entirely. And they had to be judged present or absent by an observer sliding mirrors along a beam, with no way to record what had been seen.

The 3.07 metres is therefore a human judgement about a flickering pattern, and the remarkable thing is that it was right. Modern measurements put Betelgeuse’s diameter between 42 and 56 milliarcseconds depending on the wavelength — the star has an extended atmosphere and looks larger where that atmosphere is opaque — which brackets the 47 comfortably.

What revived the method was not better optics but electronics fast enough to freeze the atmosphere: detectors that read out in milliseconds, and later active delay lines that track the turbulence. Both are ways of making the measurement in less time than the atmosphere takes to change, and that is the whole of the difference between an experiment that worked once and an instrument.

What one null cannot tell apart

The visibility curve carries much more information than its first zero, and using only the zero throws almost all of it away.

Three sources that agree at one baseline and nowhere else. Fringe visibility against baseline for three different sources, chosen so that two of them lose their fringes at the same separation — 3.08 m at 575 nm. A uniform disc 47 milliarcseconds across falls to zero there and comes back in a small sidelobe. An equal double star, whose visibility is a cosine, falls to zero there as well and then returns all the way to one, over and over. A Gaussian source of comparable width never reaches zero at all. An observer who measured only the first null would report the same angular size for all three, and would be wrong about two of them. This is the honest statement of what an interferometer measures: not a diameter, but samples of the source's Fourier transform, one spatial frequency per baseline. Recovering the shape needs many baselines, which is what aperture synthesis is; recovering it uniquely also needs the phase, which a single pair of apertures through a turbulent atmosphere does not deliver.
Fig. 3 Three sources chosen so that two of them lose their fringes at the same separation. The uniform disc falls to zero and returns in a small sidelobe; the equal double star falls to zero and returns all the way to one, over and over; the Gaussian never reaches zero at all. One null cannot tell them apart.

An equal double star has a cosine for its visibility, which vanishes and then comes back to full contrast. A Gaussian source’s visibility never vanishes. A limb-darkened disc’s first null sits a few per cent further out than a uniform disc’s of the same diameter, which is a systematic error of the same size in every diameter ever reported from a single null.

So the honest statement of what an interferometer measures is: samples of the source’s Fourier transform, one spatial frequency per baseline, with the phase missing unless something clever is done. Recovering the shape needs many baselines. Recovering it uniquely needs the phase as well, and a single pair of apertures looking through a turbulent atmosphere does not deliver phase — which is why the technique produced diameters for sixty years and images only recently.

Coherence as a length

Turn the same arithmetic round and it produces a length that is often more useful than the curve.

How far apart two pinholes may be before the fringes go. The transverse coherence radius at 550 nm — the separation at which the fringe visibility has fallen to 0.88 — for four sources, on a logarithmic scale spanning 6.3 decades. It is 18.8 µm for the Sun, 875.7 µm for a 1 mm lamp filament at 5 m, 768.65 mm for Betelgeuse, 38.85 m for a Sun-like star at 10 parsecs. Every one of them is 0.3184 λ/θ, which is 1/π, recovered from the four cases rather than asserted — the coherence radius depends on nothing about the source except how large it looks. That is why sunlight through two pinholes ten microns apart interferes and through two a millimetre apart does not, and why the same lamp behaves quite differently across a room and across a bench. It is also the reason a star is the easiest thing in the sky to interfere with: being far away is optically the same as being small, and a source small enough makes the whole aperture of a telescope coherent.
Fig. 4 The separation at which the fringe visibility has fallen to 0.88, for four sources, over seven decades. It is nineteen microns for the Sun, a millimetre for a lamp filament at five metres, seventy-seven centimetres for Betelgeuse and thirty-nine metres for a Sun-like star at ten parsecs — and every one of them is λ/πθ\lambda/\pi\theta.

That single formula explains a set of facts that look unrelated.

Sunlight through two pinholes interferes if they are twenty microns apart and not if they are a millimetre apart. Young’s original experiment worked because his pinholes were close together and he was using a source he had already restricted with a first pinhole.

A lamp is more coherent across a room than across a bench. The further away it is the smaller it looks, so the coherence radius grows in proportion to the distance. That is the opposite of the intuition most people bring to the word, and it is why a distant street lamp makes visible fringes in a rain-streaked window and a nearby one does not.

And starlight is coherent across an entire telescope. A star of one milliarcsecond has a coherence radius of thirty-nine metres, so every aperture ever built is small compared with it and starlight arrives at a mirror as a plane wave. That is why a telescope forms a diffraction-limited image of a star at all, and it is the reason astronomy is possible: being very far away is optically the same as being very small.

The last of these is the one to keep. Distance manufactures coherence. No filtering, no laser, no clever apparatus: a source small enough in angle is spatially coherent over any aperture, and the whole of stellar interferometry rests on the fact that stars are far away.

Two sources illuminating a screen give a fringe contrast depending on how well their phases are related. The stellar case is that picture with the two sources replaced by two points on the star and the screen by two telescopes — so the contrast measured between the telescopes reports the coherence of the light, and the coherence reports the angular size. Nothing about the star is resolved; a number is inferred from how badly the fringes hold up.

Why the fringes go, in one sentence

There is a way of stating the mechanism that removes all the machinery, and it is the one to remember.

Two apertures a distance BB apart see a source of angular size θ\theta. Light from one edge of the source arrives at the two apertures with a path difference Bθ/2B\theta/2 relative to light from the centre; light from the other edge with the opposite difference. The fringe pattern from one edge is therefore displaced relative to the pattern from the other by a fraction Bθ/λB\theta/\lambda of a fringe spacing. When that fraction reaches one, the patterns from the two halves of the source are a whole fringe apart and the sum is uniform.

That gives the null at Bλ/θB \approx \lambda/\theta immediately, without any Bessel functions, and the 1.22 is the correction for the fact that a disc is not two edges. Every result in this essay is that one comparison — the source’s angular size against the fringe spacing the baseline produces — and the transform is the machinery for doing it properly.

The same sentence read backwards is the design rule. To measure a source of angular size θ\theta, use a baseline of about λ/θ\lambda/\theta; smaller and the fringes are perfect and say nothing, larger and there are no fringes to measure. An interferometer is only useful over about one decade of source size around its own baseline, which is why an array carries many.

The instrument it became

Michelson’s beam is now an array of telescopes, and the arithmetic has not changed.

The Very Large Telescope Interferometer combines four 8-metre telescopes over baselines up to 130 metres, giving an angular resolution of about a milliarcsecond in the near infrared. CHARA in California has six one-metre telescopes on baselines to 330 metres. Both measure visibilities exactly as Michelson did, and both now recover images by combining three telescopes at a time and using the closure phase — a sum of three phases round a triangle from which the atmospheric errors cancel identically — to get back the phase information a two-element instrument cannot.

Radio astronomy got there first and goes much further, because the wavelengths are longer and the phases can be recorded rather than combined optically. Very-long-baseline interferometry synthesises apertures the size of the Earth, and the image of a black hole’s shadow is a visibility curve inverted.

An image is built from the spatial frequencies an aperture admits, and an interferometer does the same thing with the sampling made explicit: each baseline supplies one frequency, and an image is assembled from as many as can be measured. That is why a filled aperture and an array of telescopes are the same instrument differently sampled — and why an array with gaps produces an image with artefacts rather than an image with less detail.

The information the technique does not need

Three things a filled aperture requires and an interferometer does not, and each is why the method reached further than telescopes for so long.

It does not need the collecting area. Two small apertures separated by BB resolve exactly as well as a filled aperture of diameter BB, and collect a tiny fraction of the light. What is lost is sensitivity, not resolution — the technique is limited to bright sources and is not limited in angle by anything but the baseline.

It does not need optical quality across the whole baseline. The beam between the two apertures has to be delivered with its path length controlled to a fraction of a wavelength, which is hard, but there is no requirement that anything be figured to that accuracy across metres of glass. A mirror three hundred metres wide is impossible; two mirrors three hundred metres apart are an engineering problem.

And it does not need the source to be resolved at all in the ordinary sense. The visibility is measurable when it is 0.99, which corresponds to a source far smaller than the fringe spacing. So a well-calibrated instrument extracts a size from a source it is nowhere near resolving, at the price of the calibration being everything — an error of one per cent in the measured contrast is an error of ten per cent in a diameter, near the top of the curve.

That last trade is the honest characterisation of the technique. It converts a question about resolution, which is limited by physics, into a question about photometric accuracy, which is limited by care. The étendue of a beam cannot be improved by any instrument, and neither can the diffraction limit of a given aperture; what an interferometer does is choose a different aperture to be limited by.

The reason a light source is built small

The formula λ/πθ\lambda/\pi\theta says that coherence is bought by making a source small in angle, and there is an industry which has taken that literally at very great expense.

A synchrotron X-ray beamline sits some tens of metres from an electron beam whose cross-section is a few tens of micrometres. At a tenth of a nanometre that gives a transverse coherence length of tens of micrometres at the sample — small, but not zero, and it is what makes a growing family of techniques possible. Coherent diffractive imaging reconstructs a specimen from its diffraction pattern alone, with no lens anywhere, by iterating between the measured intensities and the constraint that the object is finite. The method works only if the illumination is coherent across the whole specimen, so the coherence area is the hard limit on how large a thing can be imaged.

That constraint is why the current generation of storage rings is being rebuilt. The upgrades are not about producing more X-rays; they are about producing them from a smaller electron beam — reducing the emittance, which is the beam’s own size-times-angle — because a smaller source is a more coherent one at the same distance. Hundreds of millions of pounds have been spent on machines whose headline improvement is that their source looks smaller from where the experiment sits.

The same statement applies to electrons. An electron microscope’s coherence is set by the angular size of its emission region, which is why a field-emission tip a few tens of nanometres across replaced a heated filament, and why holography with electrons became possible when it did.

The other way past the atmosphere

Michelson’s difficulty was that turbulence scrambles the phase between his two mirrors faster than an eye can integrate. Electronics fast enough to freeze it is one answer; there is a second, and it works with an ordinary telescope and no interferometer at all.

Take a great many exposures short enough to freeze the atmosphere — a few tens of milliseconds. Each one is a mess: instead of a clean diffraction pattern, the star’s image is a boiling cloud of speckles. Averaging the images gives the usual blurred seeing disc and throws the information away.

Averaging the power spectra does not. Each frame’s Fourier transform contains the source’s own transform multiplied by an atmospheric factor that changes from frame to frame; squaring before averaging keeps the source’s contribution and replaces the atmospheric one by its mean, which can be measured on an unresolved star and divided out. What survives is the visibility curve of this essay, out to the full diffraction limit of the telescope, recovered from images that individually show nothing.

Labeyrie demonstrated it in 1970 and it transformed the field: binary stars far too close to separate by eye were resolved in their hundreds, and stellar diameters became routine rather than heroic. The technique is the same trick as measuring correlations of intensity rather than of amplitude — throw away the phase, which the atmosphere has ruined, and keep the modulus, which it has not — and it is why the fifty-year gap after Betelgeuse closed when it did.

What the picture cannot show

The phase is missing from every curve here. Visibility is the modulus of a complex quantity, and the argument — which says where the structure is rather than how big it is — is thrown away by every measurement drawn. A source and its mirror image have identical visibility curves.

The source is assumed spatially incoherent. That is excellent for a star and for a filament and wrong for a laser, a maser, or any source with correlations across it. Applying the theorem to such a source gives an answer with no meaning.

The bandwidth is assumed narrow. Fringes from a broadband source wash out when the path difference exceeds the coherence length, which is the longitudinal problem rather than the transverse one, and the two are constantly conflated. A real instrument has to control both, and the path lengths in a stellar interferometer are held equal to within microns by delay lines hundreds of metres long.

Every curve is drawn for one wavelength. A source’s angular size is a property of the source, but its apparent size is not always: a star with an extended atmosphere looks larger at wavelengths where that atmosphere absorbs, so the measured diameter is a function of the filter. That is not an error to be removed — it is the most interesting thing such a measurement finds — and it means a single number for a star’s diameter is always shorthand.

The visibility curve of a uniform disc is a transform of the source in exactly the way a single slit’s diffraction pattern is a transform of the aperture. Both are the same operation applied at different ends of the optical path, which is why the same Bessel function turns up in both — and why the first null of the visibility curve carries the same 1.22 that the resolution limit does.

And the atmosphere is absent. Every visibility measured through air is reduced by turbulence, by an amount that fluctuates, so the raw contrast is not the source’s. Calibrating it against an unresolved star observed minutes earlier is the standard remedy and it is the dominant error in most published diameters.

The ladder from here

Later rungs on this anchor: the closure phase and how three telescopes recover what two cannot; limb darkening and the systematic error it puts into every diameter measured from a first null; intensity interferometry, where the fluctuations of the intensity are correlated instead of the amplitudes, which throws away phase entirely and in exchange stops caring about the atmosphere; the coherence of thermal light and the photon-bunching that underlies it; and the connection between the transverse and longitudinal coherence, which are two projections of one four-dimensional correlation function.

The neighbouring ladders are why two lamps never interfere, which is where the incoherence in this essay is established, and how far along its own path a wave stays in step, which is the same question asked along the beam instead of across it. The diffraction limit shares its 1.22 with this essay because both are the transform of a disc.

Part 3 of 6

This essay is one argument about Coherence. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Angular diameterAperture synthesisCoherenceCoherence lengthFourier transformSpatial coherenceVan cittert zernikeVisibility