The wiggle faster than any wave in it
Assumes: The fan of plane waves inside every beam · How far apart two things have to be
Every point of a wavefront can be treated as a source of wavelets, and every beam can be written as a fan of plane waves travelling in different directions. The second description carries the diffraction limit in one line. A plane wave travelling at an angle to the axis varies across the beam at a spatial frequency set by the sine of that angle, and the sine cannot exceed one. So no travelling component of the light varies across the beam faster than one cycle per wavelength, and any detail finer than that — carried by components that would have to vary faster — is evanescent and never reaches the far side of the lens. That is why two things have to be a certain distance apart to be seen as two.
It is natural to read this as a statement about the light itself: that a field made of components none of which varies faster than one cycle per wavelength cannot vary faster than that anywhere. The reading is wrong. A sum of slow waves can oscillate faster than its fastest component — over a stretch as long as desired, as fast as desired — and the phenomenon has a name, superoscillation. What prevents it from dissolving the diffraction limit is not that it is impossible but what it costs.
A sum that outruns its terms
The example that made the idea concrete is a single formula, written down by Michael Berry following work by Yakir Aharonov on quantum measurements in the late 1980s. Take , with a number greater than one, and raise it to the power . Multiplying out gives a sum of terms with running from to , in steps of two. That is all it contains: plane waves whose wavenumbers are integers no larger than . Projecting the function onto or finds nothing there, to the precision of the arithmetic.
Now look near . For small , is approximately , which is approximately , and raising it to the power gives . Near the origin the function behaves like a wave of wavenumber — larger than by the factor , and so faster than any component the function contains. The first figure shows it for and : in the shaded window the solid curve completes three full oscillations while the fastest ingredient, , completes two. Nothing forbids from being 10, or 100. The window can hold as many fast oscillations as desired, at as high a rate as desired.
This looks like a contradiction of Fourier analysis and is not one. The statement that a function contains no wavenumber above is a statement about the whole function over a whole period. It constrains averages, not local behaviour: over a full period the phase of a function built from wavenumbers up to can turn by at most turns, but nothing says how that turning is distributed. Berry’s function spends it lavishly near the origin and has to economise elsewhere.
Eleven beams, most of them cancelling
The formula can be built out of light directly, and doing so shows where the trick is hidden. Each term is a plane wave crossing a screen at an angle whose sine is — for , eleven beams, one straight on and the others tilted by steps of a fifth in the sine, the most tilted grazing the screen. The coefficient of each term is the amplitude its beam must carry. For those amplitudes are 57.7 for the beam with , then , 288, , 150, , 16.6, , 0.4 and for the beams at down to , and zero at : large, alternating in sign, and summing at the origin to exactly 1.
That is the whole mechanism. At the origin the eleven beams, whose sizes run into the hundreds, cancel almost perfectly, leaving a field of size 1 whose phase happens to turn fast. Move a little way from the origin and the beams’ relative phases change, and the fine balance of the cancellation shifts; what remains changes rapidly because it is a small difference of large things, and a small difference of large things can change far faster than either of them. Move to and every beam arrives with the same phase: there the amplitudes add without cancellation, and their sizes sum to 1,024, which is .
So the superoscillation is not something the waves do in addition to interfering. It is interference, of an extreme kind: destructive almost everywhere near the origin, to a precision of one part in a thousand, with the residue shaped to oscillate fast. The same description will turn up again when the spot is built from a mask, and it explains both prices in advance. The energy is in the large beams, and nearly all of it goes where they do not cancel. And a residue of one part in a thousand of the beams is only as accurate as the beams are.
It also shows why the phenomenon was missed for so long. Nobody designing a lens or an antenna would think to feed it with components hundreds of times larger than the field wanted, alternating in sign; the ordinary design principle is to make the components add. Superoscillation was found by people working on quantum measurement, where Aharonov and his colleagues had shown that a suitably weak measurement of a spin can return a value a hundred times larger than any the spin can have — a result that turns out to rest on exactly this arithmetic.
Where the phase turns slowly
The economy is visible in the rate at which the phase turns.
The phase of the function is , and its rate of change, the local wavenumber, is . At the origin that is . At it is , slower than the band limit by the same factor that the origin was faster. The curve crosses the band limit where , which for is at 0.615 radians. Over the half period from to the phase turns through exactly , the same as for : the superoscillating function has the same total phase as its fastest component and has merely redistributed it.
That redistribution is what the rest of this essay is about, because the function cannot redistribute its phase without also redistributing its size. A function whose phase turns fast in one place and slowly in another, and which is a sum of a limited set of waves, has to be small where the phase turns fast. The second figure shows the where; the third shows the price.
The mountains beside the valley
The size of Berry’s function over a whole period tells the rest of the story.
The magnitude is , which is 1 at the origin and at . The superoscillation lives in a valley whose floor is a thousandth of the height of the mountains on either side, for and . Doubling to twenty doubles the number of fast oscillations in the window and squares the ratio: the mountains become a million times higher than the valley. Raising to 3 makes the oscillations half again as fast and raises the mountains to , about sixty thousand. And the whole pattern repeats every : the function returns to magnitude 1 at , with another window of fast oscillation there, and the two windows between them hold every superoscillation the function has.
That is the general rule, and not a peculiarity of one formula. The number of superoscillations a band-limited function can hold in a window grows at best linearly with the effort spent, while the ratio between its size outside the window and inside grows exponentially. Measured in energy, which goes as the square of the amplitude, a superoscillating field puts an exponentially small fraction of itself where it is doing the interesting thing. The mathematics behind this was worked out in the early 1960s by David Slepian, Henry Landau and Henry Pollak, who asked how well a function limited to a band of frequencies could be concentrated in a window of time. The answer is that roughly the product of the bandwidth and the window’s length — a count of degrees of freedom — can be concentrated well, and beyond that count the concentration falls off exponentially. A superoscillation is a function that insists on more oscillations in the window than that count, and the exponential is its bill.
A spot narrower than the limit
For light the practical question is a focus: can a lens-like mask put a spot of light narrower than the diffraction limit into its focal plane, using only travelling waves?
The figure answers yes. Each solid curve is a mixture of twenty-one waves, each with a spatial frequency between zero and one cycle per wavelength — nothing the ordinary spot does not have — with the amounts of each chosen so that the field is 1 at the centre and exactly zero at three pairs of points spaced apart. Of all the mixtures that meet those conditions, the one drawn is the one with least total energy, found by solving the conditions directly. With at half a wavelength it is almost indistinguishable from the ordinary spot. With the central spot is half as wide as the diffraction limit allows, and it is still made of nothing but travelling light.
The idea is old. Giuliano Toraldo di Francia pointed out in 1952 that a pupil divided into concentric rings with chosen amplitudes and phases could make a central spot as narrow as desired, and drew the analogy with antenna arrays: an array of radiators fed with large currents of alternating sign can be made far more directional than its size would suggest. Such superdirective arrays had been analysed a decade earlier and were known to be impractical, for the reason the next figure shows. Only in the last two decades, with masks structured on the scale of a wavelength and a clearer understanding of the cost, has the optical version been built and used to image objects below the diffraction limit.
The side lobes that carry the light
The narrow spot is flanked by something the first figure does not show, because it is off the scale.
Past the designed zeros, the narrowed fields grow into side lobes, and the narrower the spot the larger they are. For the ordinary spacing the distant side lobes are a few thousandths of the spot, much like the outer rings of an ordinary focus. At a spacing of 0.35 wavelength the brightest side lobe is nearly four times the spot; at 0.25 it is five hundred times. The side lobes are Berry’s mountains in a different guise: the same band-limited field, having crowded its oscillations into the centre, has to put its size somewhere.
For imaging this sets a condition. A superoscillating spot is useful only if the object being imaged is illuminated by the spot and not by the side lobes — only if the field of view is confined to the dark region between the designed zeros, or the light from the side lobes is otherwise blocked. Microscopes built this way scan the spot across the object and collect light through a small aperture that excludes the side lobes, point by point. The field of view at any instant is tiny, and the spot’s own brightness is a small fraction of what illuminates the sample as a whole.
The share of light in the spot
The energy accounting makes the cost exact.
Below about 0.4 of a wavelength the curve is close to a straight line on a logarithmic scale, which is to say the cost is exponential. Narrowing the spot from 0.35 to 0.25 of a wavelength costs a factor of two hundred in the share of light it receives. Narrowing it further to 0.2 costs another factor of nearly twenty. The least-energy design is the best that can be done for these conditions — any other mixture that makes the same zeros puts more energy into the side lobes, not less — so this is a bound, not a failure of design.
It is also why superoscillation is not a loophole in the diffraction limit so much as a clarification of what the limit says. The limit is about how much of the light a lens can put into a region of a given size. For regions larger than about half a wavelength, almost all of it. For smaller regions, an exponentially small part, with the rest thrown elsewhere. The ordinary statement — that nothing finer than half a wavelength can be focused — is the practical form of the second clause, and it is the right one whenever the light is limited.
How exactly the mixture must be made
The last figure shows the second price, which is often the one that decides.
The spot at the centre is what remains when components 127 times larger than it very nearly cancel. An error in any one of them that is small relative to the component is large relative to the remainder, and it goes straight into the spot. An error of one part in a thousand of the largest component is already an eighth of the spot’s own height, and twenty-one such errors add. A fabricated mask whose transmission is controlled to one per cent — which is good fabrication — gives a pattern in which the neighbourhood of the spot is dominated by the errors, and at a few per cent nothing of the design survives.
This is exactly the failure that sank superdirective antennas. An array fed with large alternating currents that nearly cancel in every direction but one achieves narrow beams on paper, and in practice the currents cannot be held to the required precision, the losses in the conductors carrying them are enormous, and the arrangement works only over a very narrow band of frequencies. The optical superoscillation inherits all three problems, and the balance between how narrow a spot is wanted and how precisely a mask can be made sets what is achievable. Demonstrations have made spots a few tens of per cent narrower than an ordinary focus of the same aperture, with usable brightness; much narrower spots are made on paper far more easily than in glass.
Still open: what a superoscillation can really tell
The spot can be made. Whether it delivers information the diffraction limit forbids is a more subtle question, and it is disputed. A band-limited field is an analytic function: if it is known exactly on any stretch, it is determined everywhere, and in principle the whole object can be recovered from the light that passes the lens. That principle is practically useless, because the recovery amplifies noise exponentially — the same exponential as the superoscillation’s cost, seen from the other side. The number of independent details that can be recovered in the presence of noise is set by the count of degrees of freedom in the Slepian sense, and a superoscillating spot does not change that count; it trades brightness for a local sharpening of the probe.
What superoscillation clearly provides is a narrow probe in a dark field, useful when the light is plentiful and the region to be examined is small. Whether it provides more than that — whether, combined with prior knowledge of the object, it can resolve detail that could not have been resolved with the same total light and an ordinary focus — depends on how the signal-to-noise ratio and the prior information are counted, and different groups count them differently. Related questions arise for pulses that rise faster than their bandwidth should allow and for quantum measurements whose outcomes lie outside the range of the operator measured, which is where the idea began.
The habit worth carrying from here is to ask whether a limit is local or global. A limit on the frequencies in a field constrains it on average, over the whole of it, not at every place, and a local violation is always paid for somewhere else. The useful question about any claimed escape from such a limit is where the payment has been put.
Part 6 of 6
This essay is one argument about Huygens. The others:
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
AntennaBandwidthDiffraction limitEvanescent waveFourier transformInterferencePhaseResolving powerSignal-to-noiseSuperposition
- The phases that turn a glow into pulses bandwidth, fourier transform, interference, phase, superposition
- How far a wave can remember bandwidth, fourier transform, interference, superposition
- The fringe and the spectrum are one measurement bandwidth, fourier transform, interference, resolving power
- Everything a scatterer removes, from one direction interference, phase, superposition
- The backward wave Huygens had to remove interference, phase, superposition
- The grating that photographs itself fourier transform, interference, phase