Optics

The rings that belong to the edge

Every account of diffraction so far asks what the size of an aperture does. The rings around a star are not about its size: they are the transform of a discontinuity, they do not shrink relative to the core when the telescope grows, and the only way to remove them is to stop the transmission falling to zero abruptly. Softening the edge buys forty-five decibels of contrast and costs eighty per cent of the resolution.

Assumes: How far apart two things have to be · What a thousand slits buy that two cannot

The limit on telling two things apart establishes that what an instrument can separate is set by the width of the hole light comes through. Every essay since has asked a question about a size — how wide the slit is, how many lines the grating has, how far away the screen is.

There is a second property of an aperture and it decides something different. An aperture is not only a size; it is a transmission at every point across it. Writing that transmission as a function and calling it the pupil function, the far field is its Fourier transform and the image is the squared modulus of that. Everything about the pattern is in the function.

A plain hole is the particular pupil function that equals one inside and zero outside, and it has a discontinuity at the rim.

The rings belong to the edge, not to the size. The far-field intensity of 3 apertures of the same width, against angle in units of the diffraction limit, on a logarithmic intensity axis spanning ten decades. They differ only in how the transmission falls off toward the rim. With a hard edge the first sidelobe is 13.3 decibels down and the core is 0.89 wide. With a Hann taper the first sidelobe is 31.5 decibels down and the core is 1.44 wide. With a Blackman taper the first sidelobe is 58.1 decibels down and the core is 1.64 wide. The hard edge's rings are not a defect of the optics and are not reduced by making it larger — they are the transform of a discontinuity, and the only way to remove them is to remove the discontinuity. What it costs is the width of the core, which is the resolution.
Fig. 1 The far field of three apertures of the same width, on a logarithmic intensity axis spanning ten decades. They differ only in how the transmission falls off toward the rim. A hard edge puts its first sidelobe 13.3 decibels below the peak; a Hann taper, 31.5; a Blackman taper, 58.1 — and the cores are 0.89, 1.44 and 1.64 diffraction limits wide.

The rings are the transform of a cliff

A function that drops abruptly has a transform with a long tail, and the tail’s envelope falls like a power of the distance rather than exponentially — which is the same statement a truncated interferogram makes about a spectral line. That is not an optical fact; it is a general property of transforms, and it is why the same shapes appear in signal processing, in radio astronomy and in any measurement made over a finite window.

For a hard edge in one dimension the transform is a sinc, its sidelobes fall as the inverse square of the angle, and the first one sits 13.26 decibels below the peak. For a hard circular edge the transform is an Airy pattern, its rings fall as the inverse cube, and the first ring is 1.75 per cent of the peak — which is the pattern the resolution criterion is laid over.

Both numbers are properties of the shape of the pupil, and not of its size. Scale the aperture and the pattern scales with it: everything moves inward in proportion and every ratio is untouched. A hundred-metre telescope has its first ring exactly 13 decibels below its peak, exactly as a hundred-millimetre one does.

That is the first thing worth saying clearly, because the instinct is that a bigger instrument has a cleaner image and in this respect it does not. What a bigger instrument buys is that the rings are closer in, so the region beyond them where the light has genuinely fallen away begins at a smaller angle. If the object of interest sits at a fixed angle from the star, that helps; if it sits at a fixed number of resolution elements, it does not help at all.

What softening the edge costs

Make the transmission fall smoothly to zero at the rim and the discontinuity is gone, so the tail of the transform is gone with it.

The rings belong to the edge, not to the size. The far-field intensity of 3 apertures of the same width, against angle in units of the diffraction limit, on a logarithmic intensity axis spanning ten decades. They differ only in how the transmission falls off toward the rim. With a hard edge the first sidelobe is 13.3 decibels down and the core is 0.89 wide. With a Gaussian taper the first sidelobe is 35.6 decibels down and the core is 1.27 wide. With a Blackman taper the first sidelobe is 58.1 decibels down and the core is 1.64 wide. The hard edge's rings are not a defect of the optics and are not reduced by making it larger — they are the transform of a discontinuity, and the only way to remove them is to remove the discontinuity. What it costs is the width of the core, which is the resolution.
Fig. 2 A Gaussian taper and a Blackman taper against the hard edge. Neither has any structure resembling rings at all past the first few diffraction limits — the light simply falls away — and both pay for it in the width of the core.

Nothing comes free and the exchange rate is worth having in front of one.

What each decibel of contrast costs in resolution. Peak sidelobe level against the width of the core, for 6 standard pupil tapers, each computed from the same transform. The envelope of the points is the trade: the narrowest core has the worst sidelobes, the quietest pupil has the widest core, and no taper sits below and to the left of all the others. Going from a hard edge to a Hann taper buys 18 decibels for a core 1.63 times wider. The points lie near a curve rather than exactly on one, and the exception is worth seeing: a Hamming taper has both a narrower core and a lower first sidelobe than a Hann taper, which looks like something for nothing and is not — the two differ in how fast the FAR sidelobes fall away, and Hamming buys its first sidelobe by giving that up. Which point is wanted is decided by what is being looked for: two stars of similar brightness want the left-hand end, a planet beside a star wants the right, and something at a known separation wants whichever taper happens to be quiet there.
Fig. 3 Peak sidelobe level against the width of the core for six standard tapers, each computed from the same transform. The envelope of the points is the trade: the narrowest core has the worst sidelobes and the quietest pupil has the widest core.

The figure also contains an exception worth reading carefully, because it looks like a free lunch. A Hamming taper has a narrower core than a Hann taper and a lower first sidelobe. Both of those are true, and it is not free: the two windows differ in how fast the far sidelobes fall, and Hamming buys its first sidelobe by giving that up. Hann’s sidelobes decay steeply with angle and Hamming’s do not, so for something far out Hann is much the better choice and for something just outside the core Hamming is.

That is the general shape of the subject. There is no best pupil; there is a best pupil for a stated requirement, and the requirement has to name a separation and a contrast before the question is well posed — which is the same discipline a resolution criterion needs and usually does not get.

Why a taper is the same thing as a longer measurement

There is a second way of reading the trade that makes it feel less like a compromise and more like an identity, and it is worth having because it explains why the same curve turns up everywhere.

A pupil is a window. The far field is the transform of what is inside it, and what a window does to a transform has nothing to do with light: it is that a measurement made over a finite range cannot distinguish frequencies closer than the reciprocal of that range, and that the way it fails to distinguish them is decided by the window’s shape.

So the core’s width is the reciprocal of the pupil’s effective width, and tapering the pupil reduces its effective width — the outer parts contribute less, so the aperture behaves as though it were smaller. That is the whole of the cost. A Hann-tapered aperture of diameter DD resolves like a hard-edged one of diameter D/1.63D/1.63, and its sidelobes are those of a smooth window rather than a sharp one.

Read that way the trade stops being mysterious. Contrast is bought by making the aperture effectively smaller, and the reason it cannot be bought otherwise is that a full-width hard edge is what a full-width measurement is. Anything that uses all of the aperture with equal weight has the discontinuity at the rim; anything that removes the discontinuity is not using all of the aperture.

It also says where the limit of the whole approach is. The best possible window for a stated main-lobe width — the one with the least energy outside it — is a prolate spheroidal function, and it is the solution to exactly that optimisation. Every window in the figure above is an approximation to a member of that family, chosen because it is easy to compute or easy to manufacture, and none of them can beat it.

The requirement that drives it

The reason anybody cares is that some things are very much fainter than the things beside them.

How faint a companion each taper allows. The light a star still puts at a given angular separation, for three pupil tapers at three separations, on a logarithmic axis — which is directly how faint a companion has to be before it is lost. The envelope is taken rather than the value at a point, because a companion does not arrange to sit in a null. The best combination here is a Blackman taper at 16 resolution elements, which leaves 1.3e-9 of the peak. An Earth beside a Sun is one part in ten thousand million at about one tenth of an arcsecond, so nothing on this chart is remotely enough, and the instruments that attempt it use a taper as one stage of several rather than as the answer.
Fig. 4 What each taper leaves at a given separation, taken as the envelope rather than at a point, because a companion does not arrange to sit in a null. A Blackman taper at sixteen resolution elements leaves a thousandth of a millionth of the peak — and an Earth beside a Sun is one part in ten thousand million.

Put the numbers side by side. A Jupiter beside a Sun-like star reflects about one part in a thousand million of the starlight, and sits at about half an arcsecond at ten parsecs. An Earth is one part in ten thousand million and sits at a tenth of an arcsecond. For an eight-metre telescope in the visible, a tenth of an arcsecond is about three resolution elements.

Three resolution elements and ten orders of magnitude. The hard-edged aperture leaves something like one part in ten thousand there; the best taper in the figure leaves perhaps one part in a hundred thousand at three elements. That is five or six orders of magnitude short.

So apodisation is not the answer to the problem it was invented for, and the instruments that attempt it use several stages: an occulting mask to block the star’s core, a shaped pupil or a phase mask to suppress what leaks past — the zone plate’s trick of painting out what would arrive out of step, applied to a star — a second pupil stop to catch the light diffracted by the first mask, and active correction of the wavefront to hold it all still. Apodisation appears twice in that list and is one stage of four.

How faint a companion each taper allows. The light a star still puts at a given angular separation, for three pupil tapers at three separations, on a logarithmic axis — which is directly how faint a companion has to be before it is lost. The envelope is taken rather than the value at a point, because a companion does not arrange to sit in a null. The best combination here is a Hann taper at 12 resolution elements, which leaves 4.4e-8 of the peak. An Earth beside a Sun is one part in ten thousand million at about one tenth of an arcsecond, so nothing on this chart is remotely enough, and the instruments that attempt it use a taper as one stage of several rather than as the answer.
Fig. 5 The same comparison at separations a real instrument works at — three, six and twelve resolution elements. The contrast available falls very steeply as the object moves inward, which is why the inner working angle is the number an instrument is judged by and why every scheme in the subject is a fight over the first few diffraction limits.

Where else the same curve is used

The trade in the second figure is the same one that appears wherever a measurement is made over a finite window, and the vocabulary is shared because the mathematics is identical.

A spectrometer measures an interferogram over a finite path difference and transforms it, so its instrumental line shape is the transform of the window. A hard truncation produces sidelobes of 22 per cent, which put spurious features beside every strong line, and an interferometric spectrometer is always apodised before transforming — trading resolution for a line shape with no feet on it.

A radio array has a pupil that is a set of points rather than a continuous aperture — a grating’s pupil, with the count deciding the width of each maximum — and its point-spread function — the dirty beam — has enormous sidelobes because the pupil is mostly holes. Weighting the measurements differently is apodisation by another name, and the choice between natural and uniform weighting is a choice on exactly the curve above.

An antenna radiating from an aperture has the same transform, and its sidelobe level is a specification with regulatory force: a satellite dish must not radiate above a stated level off-axis, and the way it meets that is by tapering the illumination of its own reflector, at the cost of gain.

And a laser beam is the case where nature has already chosen. A beam leaving a laser has a Gaussian amplitude profile rather than a hard edge, so its far field is Gaussian and has no rings at all — which is why a laser spot is a clean blob and a star through a telescope is not. The price is the same one: a Gaussian beam of a given width diverges more than a hard-edged one of the same width would, and the extra divergence is the widened core in a different vocabulary. The resonator that produced it selected that profile because the mode with the least loss round the cavity is the one with no sharp edges to spill light past the mirrors.

In every case the same two quantities trade and the same curve governs it. That is one of the better arguments for meeting the Fourier relation once rather than four times: the answer to “how are the sidelobes lowered” is the same in optics, spectroscopy, radio and antenna design, because the question is about a transform and not about any of those subjects.

The other way of shaping an aperture, which is to make holes in it

Tapering the transmission continuously is one solution and it is not the one that is built, because a piece of glass whose transmission varies smoothly from one to zero over a metre is difficult to make, difficult to keep clean and impossible to make achromatic.

What is built instead is a binary mask: a piece of metal with holes in it, fully transparent or fully opaque, whose shape is designed so that the transform has the required contrast over the required range of angles. A shaped pupil of that kind is a purely geometric object, it has no coating to degrade, and it works at every wavelength the same way because nothing about it is dispersive.

The design is an optimisation rather than a formula. Maximise the light passed, subject to the constraint that the transform’s intensity stays below a stated level throughout a stated region of the image plane — which is a linear program in the pupil’s own opening function, and is solved as one. The solutions look nothing like a taper: they are sets of slots of varying width, or a mask with a pattern of curved openings, chosen so that the light diffracted by one edge cancels the light diffracted by another in the region that matters.

Two things about them are worth carrying away, because both are counterintuitive.

The contrast is achieved only in part of the image. A shaped pupil produces a dark region — a “discovery space” — over a range of angles and a range of directions, and the light it removed from there is elsewhere in the image at higher intensity than a hard edge would have put it. That is not a defect of the design; it is a consequence of the total being fixed. The designs trade area of dark region against depth of darkness and against throughput, on a three-way version of the curve above.

And the binary mask beats the smooth taper. A continuous apodisation of the kind drawn in the figures cannot reach ten to the minus ten at three resolution elements at any throughput; a shaped binary pupil can, in the calculation. The reason is that the optimisation is over a much larger space — any opening shape at all, rather than a radially symmetric transmission profile — and symmetric solutions are not the best ones.

One dimension, a third of the light, and a pupil with struts in it

The figures are one-dimensional. A slit and a circular aperture have different transforms — a sinc against an Airy pattern — with different sidelobe levels and different falloff. The tapers behave the same way in both and the numbers are not the same numbers, so the values here are right for a slit and indicative for a telescope.

Apodisation by transmission throws light away. A Blackman-tapered pupil passes about a third of the light a hard-edged one of the same size does, so the improvement in contrast is bought with exposure time as well as with resolution. That does not matter for a bright star, which is the case apodisation is usually wanted for, and it matters a great deal for anything else.

And a real pupil is not just a shape. A telescope has a secondary mirror in the middle and struts holding it, both of which are part of the pupil function, and both of which contribute their own diffraction — the struts produce the spikes on bright stars in every astronomical image, by the same boundary wave any obstacle’s rim radiates. An apodisation designed for a clear circular aperture does nothing about either, and the schemes that work on real telescopes are designed around the obscuration rather than in spite of it.

Amplitude across the pupil, intensity on the screen

They cannot show that the transform is of the amplitude and the picture is of the intensity. Apodisation works on the amplitude across the pupil, and the sidelobe levels quoted are intensity ratios, so a taper that halves the amplitude at the rim reduces the intensity there by four. That squaring is why the dynamic range in the figures runs over ten decades of intensity for a pupil whose transmission varies over one.

Nor can they show the phase. Everything here is a pupil whose transmission varies and whose phase is uniform, and there is a whole family of devices — phase masks, vortex coronagraphs — that leave the transmission alone and vary the phase instead, achieving suppression a transmission taper cannot. Those are the subject of a different calculation with the same transform in it.

And they cannot show that the sidelobes are still there. A taper does not remove light; it redistributes it. The light that is no longer in the first ring has gone into the core and into a smooth halo further out, and at some large angle every taper’s pattern is again dominated by whatever the real optic scatters rather than by anything in this calculation.

Still open: how far a shaped pupil can be pushed

The design problem — find the pupil transmission maximising throughput subject to a stated contrast within a stated range of angles — is an optimisation with a known structure, and solutions have been computed since the early 2000s for shaped pupils with contrast of ten to the minus ten over a band of angles. Whether those designs survive the things the calculation leaves out is the open part: a real pupil has a central obscuration, the wavefront is not perfectly flat, the star is not a point, and the optics move.

The quantity that decides it is how stable the wavefront is over the hours an exposure takes, because a contrast of ten to the minus ten requires the wavefront to be held to picometres over that time. No instrument has yet demonstrated it on the sky, several are being built, and the honest position is that the pupil design is the part of the problem that is understood.

The habit worth carrying away is about what a discontinuity costs. A function that stops abruptly has a transform with a long tail, and the tail is usually the thing that limits the measurement. The tail is not an imperfection to be polished away — it is the exact consequence of a sharp edge — and the only repair is to stop having one, which always costs resolution because resolution is what the sharp edge was buying.

Part 7 of 8

This essay is one argument about Diffraction. The others:

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

ApertureApodisationContrastCoronagraphDiffractionDynamic rangeFourier transformPoint-spread functionResolutionSidelobesTrade offWindow function