Waves

The pulse two failures keep alive

Dispersion spreads a pulse until it is nothing. Nonlinearity steepens it until it breaks. Each on its own destroys a disturbance, and there is exactly one height for each width at which the two cancel completely — leaving a shape that travels for ever and survives being run into by another one.

Assumes: The packet that will not keep its shape · The front that steepens until it cannot

Two quite different things destroy a pulse, and both are ordinary. Dispersion spreads it: the components of the pulse travel at different speeds, drift out of step, and the pulse flattens into a long train of ripples. Nonlinearity steepens it: the tall part travels faster than the shallow part, the front leans forward, and the profile heads for a vertical face.

Two ways of destroying a pulse, and their cancellation. The same starting pulse, carried forward three times by three equations that differ only in which terms are present, all shown at t = 0.097. Dotted: where it began. With the dispersive term alone the pulse spreads and sheds an oscillating tail, because its Fourier components travel at different speeds and drift out of step — the profile departs from what it was by 27 per cent of its own height. With the nonlinear term alone the tall part overtakes the shallow part and the front leans forward: the steepest gradient is 9.3 times what it started as and is on its way to vertical. With both, neither happens — the profile has moved to the right and is otherwise identical to what it was, to 2.7e-12 of its own height. The balanced run is a pseudo-spectral integration and the other two are exact, and the integration conserves the two quantities the equation conserves — mass to 2.2e-16 and the squared integral to 4.7e-15 — which is the check that the answer belongs to the equation rather than to the integrator. What the picture cannot show is why the cancellation is stable: a pulse of the wrong height for its width does not persist in the wrong shape, it sheds the excess as a dispersive tail and settles on the shape that works, which is why these objects turn up in canals and optical fibres rather than only in equations.
Fig. 1 The same starting pulse under three equations differing only in which terms are present, all at the same time. With the dispersive term alone it spreads and sheds a tail. With the nonlinear term alone it steepens nine-fold. With both, the profile has moved to the right and is otherwise identical to what it was, to a part in ten thousand billion of its own height.

Put both in the same equation and they usually make matters worse. But their effects on a profile have opposite signs, and for one shape at each height they cancel — not approximately, but completely. What is left travels without changing at all.

What the equation is

Add a dispersive term to the steepening equation and the result is Korteweg and de Vries’s equation of 1895, written here in its standard scaled form:

ut+6uux+3ux3=0\frac{\partial u}{\partial t} + 6u\frac{\partial u}{\partial x} + \frac{\partial^3 u}{\partial x^3} = 0

The middle term is the amplitude-dependent speed that leans a front forward. The third-derivative term is the lowest-order correction that makes long waves travel slightly faster than short ones — it is what a shallow-water wave gets from the finite depth, and it is a dispersion in the same sense as the spreading of a quantum packet though with a different power of the wavenumber.

The two terms are of comparable size for a pulse of the right proportions and of wildly different size for anything else. A very wide, very low pulse is dominated by the nonlinearity, because the third derivative of a gentle curve is negligible; a very narrow, very low one is dominated by the dispersion. Somewhere between there is one combination for each speed at which the two are equal at every point, and that is the solitary wave.

The narrow packet is the one that spreads. Three packets on deep water, ω = √(gk), all built on the same 4 m carrier and differing only in bandwidth — 10%, 18%, 28% of the carrier wavenumber. Each curve is the width of the emitted envelope, measured as the second moment of its intensity about its own centroid in a frame moving at the group velocity, divided by that width at the start. The starting widths are 4.50 m, 2.50 m, 1.61 m, and the order of the curves is the reverse of the order of the widths: the shortest packet, which is the one with the widest spectrum, is the one that comes apart first. The dashed curves are √(1 + (t/τ)²) with τ = σ₀²/|d²ω/dk²| — computed from the dispersion relation, not fitted. They agree with the measured widths to 0.3% at 10% bandwidth, 3.3% at 18% bandwidth, 8.5% at 28% bandwidth, and that ordering is the second thing the figure says: the closed form keeps only the curvature of ω(k), so it is exact for a narrow spectrum and starts to fail for a wide one, by about as much as the cubic term is worth. A packet with no bandwidth would never spread at all, and would also never begin or end.
Fig. 2 Dispersive spreading on its own, for three bandwidths. A wider spread of components spreads faster, and nothing about a linear medium stops it: the width grows without limit, and the height falls to match.

The shape it has to be

The balance does not permit a family of shapes. It permits exactly one, up to a scale, and the scale is not free either.

Taller is faster, and narrower with it. Three solitary waves of the Korteweg–de Vries equation, drawn at the same instant. Each is a sech² profile and each is a solution — substituted back into u_t + 6uu_x + u_xxx = 0, they leave 4.7e-7. They are not free to be any shape or any size: the amplitude is half the speed and the width goes as one over the square root of it, so a taller solitary wave is a faster and a narrower one, and the three cannot be scaled independently. Amplitude times width squared is the same for all three to 0.24 per cent, measured off the drawn curves rather than derived. That single-parameter family is what the balance produces: the nonlinear term steepens the front at a rate that grows with height, the dispersive term spreads it at a rate that grows with curvature, and only one height goes with each width for the two to cancel. The picture cannot show the consequence that makes these objects remarkable, which is what happens when two of them meet.
Fig. 3 Three solitary waves drawn at the same instant. Each is a sech-squared profile and each solves the equation to a part in ten million. The amplitude is half the speed and the width goes as one over the square root of it, so amplitude times width squared is the same number for all three.

The profile is sech2\operatorname{sech}^2, and it carries one parameter. Fix the speed and everything else follows: the height is half the speed, and the width goes as one over the square root of the height. A taller solitary wave is a faster and a narrower one, and there is no such thing as a tall wide one or a short narrow one.

The tie is worth putting in words that survive the scaling. Doubling a solitary wave’s height doubles its speed and shrinks its width by a factor of the square root of two; it also multiplies the energy it carries by about 2.8, and the momentum by 2. So a train of solitary waves emerging from an arbitrary initial lump is sorted: the tall ones are also the fast ones, and they arrive first. That sorting is visible in a bore running up an estuary, where the leading undulations are the largest, and it is not a coincidence of that geometry — it is the height–speed tie applied to whatever the tide happened to deposit.

That single-parameter family is what the balance produces, and the reason is easy to state. The nonlinear term steepens at a rate that grows with the height; the dispersive term spreads at a rate that grows with the curvature, which for a fixed height grows as one over the width squared. Setting the two equal ties the height to the width, and the tie is the relation the three curves in that figure obey to within a quarter of a per cent, measured off the drawn profiles rather than derived.

The consequence for anybody trying to make one is that the shape cannot be dictated. Launch a pulse of the wrong proportions and it does not simply persist in the wrong shape: it sheds the excess as a dispersive tail and settles onto the shape that works. That is why these objects are found in nature rather than only in equations, and it is why Russell could make one behind a barge.

Two of them meeting

The property that makes a solitary wave a soliton, and that took seventy years after Russell to be noticed, is what happens when a fast one catches a slow one.

Two solitary waves, before and after. The exact two-solitary-wave solution of the Korteweg–de Vries equation at five times, the earliest at the bottom. The taller wave is four times the height of the shorter and four times the speed, so it starts behind and catches up; in the middle frame the two have merged into a single hump that is not the sum of them; and afterwards they separate with their heights unchanged to 0.001 per cent. That is what makes them solitons rather than merely solitary waves: a nonlinear equation has no reason whatever to let two disturbances pass through one another intact, and almost none does. The collision is not free, though, and the price is the one thing that shows: the tall wave comes out 0.548 ahead of where it would have been and the short one 1.100 behind, measured from the peak positions long before and long after. A phase shift is the entire record of the encounter — no radiation is left, no energy is exchanged, and the shapes are the same. The expression drawn satisfies the equation to 1.2e-6.
Fig. 4 Two solitary waves at five times, the earliest at the bottom. The taller is four times the height of the shorter and four times the speed, so it catches up; in the middle frame the two have merged into a single hump that is not the sum of them; and afterwards they separate with their heights unchanged to a thousandth of a per cent.

A nonlinear equation has no reason whatever to let two disturbances pass through one another intact. Almost none does — that is exactly what nonlinearity means, and the previous rung on this ladder is an equation in which two waves interact so violently that they generate frequencies neither of them had. Yet here the two come apart with their shapes and speeds exactly as they went in.

During the overlap the profile is not the sum of the two. It is lower than the taller constituent and has a shape neither of them has; there is a range of parameters for which the two do not even resolve into separate humps at closest approach. So this is not superposition sneaking back in. It is something stronger and stranger: the equation possesses enough conserved quantities to forbid any other outcome.

The collision is not free. The one thing it leaves behind is a phase shift: the fast wave comes out ahead of where a free one would have been, by 0.55 in these units, and the slow one behind by 1.10, both measured from the peak positions long before and long after. No energy is exchanged, no radiation is emitted, and nothing else records that the encounter took place.

Why it is allowed to happen

The equation has an infinite number of conserved quantities, and that is not a figure of speech. The first two are familiar — the integral of uu, which is a mass, and the integral of u2u^2, which is an energy — and the integration in the hero figure holds both to within a part in a hundred thousand billion, which is the check that the answer belongs to the equation rather than to the numerical scheme.

What forbids it is a count. The Korteweg–de Vries equation has infinitely many conserved quantities — not the two or three an ordinary mechanical system has, but a countable infinity, each an integral over the whole profile, each unchanged by the evolution. That is a ferocious constraint. Two solitons coming out of a collision must reproduce every one of those integrals, and the only configurations that do are the two they went in as. The equation is integrable, and integrability is exactly this: so many conserved quantities that the motion has nowhere left to go.

The third and higher ones have no everyday names, and they are the ones that matter. A collision that changed the two waves’ heights would have to change one of those quantities, and the quantities are constant, so the collision cannot. Systems with this property are called integrable, they are rare, and the Korteweg–de Vries equation was the first partial differential equation shown to be one.

Rarity is worth emphasising. Add a small term the equation does not have — a little dissipation, a slightly different dispersion, a second dimension — and the infinite tower of conserved quantities collapses to a handful. Solitons in such a system still nearly survive collisions, and shed a little radiation each time; over enough encounters they change. The exact result is a property of an exact equation, and every physical realisation is an approximation to it.

The experiment that found them

The history is worth a section because it is unusually clean: the discovery was made on a computer, by accident, while looking for the opposite result.

A linear chain of masses and springs has normal modes, and each of them is independent: energy put into one mode stays in that mode for ever, because nothing in a linear equation couples them. That independence is why a nonlinear chain was the interesting experiment. Add a small nonlinearity and the modes do couple, so the energy should wander out of the first mode and spread across the others until it is shared equally — which is the equipartition the whole of statistical mechanics rests on.

In 1953 Fermi, Pasta, Ulam and Tsingou set a chain of masses and springs oscillating on the Los Alamos computer, with a slightly nonlinear spring law. The expectation was thermalisation: energy put into one mode would leak into the others through the nonlinear coupling and end up shared equally, which is what statistical mechanics says should happen and what a linear chain cannot do.

It did not happen. The energy sloshed among a handful of low modes and then came back, nearly all of it, to the mode it started in. The recurrence was clean enough that the run was repeated in case of a coding error. Nothing in the theory of the day accounted for it.

The nonlinearity in question was small. Fit a parabola to the springs’ potential and the departure from it is slight at every amplitude the experiment used — a quadratic term with a cubic correction of a few per cent. That was the point: a chain of springs that are nearly harmonic ought to share its energy out slowly rather than not at all, and a small departure from linearity was expected to produce a slow approach to equipartition rather than a different phenomenon.

Ten years later Zabusky and Kruskal took the continuum limit of that chain and got the Korteweg–de Vries equation. Solving it numerically, they saw the initial disturbance break into a train of solitary waves that overtook one another, passed through, and eventually re-assembled — which is the recurrence, seen in the right variables. They coined the word soliton in that paper, for a solitary wave that behaves like a particle.

The lesson generalises badly and interestingly. A nonlinear system with enough conserved quantities does not explore its phase space, so the statistical assumption that underlies the equipartition of energy fails, and it fails not because the system is too small or too cold but because it is too ordered. How large a perturbation is needed to destroy that order, and how long it takes, is a question still being worked on seventy years later.

Where the balance turns up

Russell’s canal is the least useful instance and by far the most famous. The general condition is a nonlinearity and a dispersion of opposite effect, and that combination is common.

Water is the case the balance was first found in, and it is worth being precise about which half of it is on any ordinary chart. The speed of a water wave depends on its wavelength — that variation is the dispersion, and it is what spreads a pulse. The amplitude dependence that steepens one is not on that chart at all, because a linear dispersion relation has no amplitude in it. The object this essay is about lives in the cancellation between the two, so neither curve alone can show it.

Optical fibres are the instance that pays for itself. A glass fibre is dispersive, so a data pulse spreads and eventually overlaps its neighbours, which is what sets the bit rate of a long link. Glass is also very slightly nonlinear, with an index that rises with the intensity. Choose the wavelength so that the dispersion has the right sign, and the two cancel: the pulse propagates without spreading, over thousands of kilometres.

Bose–Einstein condensates obey a close relative — the nonlinear Schrödinger equation — in which the interaction between atoms supplies the nonlinearity and the kinetic term the dispersion. Both bright solitons and dark ones have been made and photographed.

The atmosphere produces them at a scale nobody could miss. The Morning Glory over the Gulf of Carpentaria is a roll cloud a thousand kilometres long, propagating as a solitary wave on an inversion layer, and it recurs because the balance is stable.

Josephson junctions carry them as fluxons — a single quantum of magnetic flux propagating along a long junction as a solitary wave in the phase difference across it — and the same sine-Gordon equation that describes them describes a chain of coupled pendulums, which is the mechanical demonstration usually built to show the effect. Two fluxons collide and pass through one another with a phase shift, exactly as here.

And a nerve may be one, though this is contested: the mechanical soliton model of the action potential proposes that the pulse travelling along an axon is a solitary wave in the membrane rather than only an electrical event.

The man on the horse

The founding observation is usually given as an anecdote and it deserves to be given as an experiment, because Russell did the work and was disbelieved by better mathematicians than himself for a decade.

In August 1834 he was watching a boat being drawn along the Union Canal near Edinburgh when the tow-rope broke and the boat stopped. The mass of water it had been pushing did not stop. It rolled forward as, in his words, a large solitary elevation — a rounded, smooth and well-defined heap of water — about thirty feet long and a foot or so high, travelling at eight or nine miles an hour without change of form or diminution of speed. He followed it on horseback for a mile or two before losing it in the windings of the channel.

What made it more than an anecdote is what he did next. He built a wave tank at home, generated such waves deliberately by dropping weights into one end, and measured them systematically for ten years. He established that they can be produced reliably, that a disturbance too large breaks into two of them, that they pass through one another, and — the key quantitative result — that the speed depends on the height: c=g(h+a)c = \sqrt{g(h+a)}, with hh the depth and aa the wave’s own amplitude. That is the amplitude–speed tie this essay derives, measured half a century before the equation.

He was not believed. Airy, whose shallow-water theory was the standard treatment, argued that no wave of permanent form could exist: his equations contain the nonlinearity and no dispersion, so every disturbance must steepen and break, and a wave that travels unchanged was therefore impossible. Stokes was similarly sceptical for some years. The objection was not carelessness — it was a correct deduction from an approximation that had thrown away exactly the term that makes the phenomenon possible.

The theory arrived from Boussinesq in 1871 and Rayleigh in 1876, both of whom derived the sech-squared profile by keeping the next order in the depth-to-wavelength expansion, and the equation was written in the form used here by Korteweg and de Vries in 1895. Russell had been dead for a decade.

The general shape of the episode is one worth recognising. An approximation that has discarded a term does not merely lose accuracy; it can make a real phenomenon logically impossible, and the resulting argument against an observation looks exactly like a proof. What settled it was not a better experiment, because Russell’s was good; it was somebody retaining one more term.

The pulses that were supposed to carry the internet

The optical instance was named above as the one that pays for itself, and the honest version of the story is that it did not, in the way it was expected to.

Hasegawa and Tappert pointed out in 1973 that a glass fibre supplies both ingredients. Above about 1,310 nanometres its dispersion has the sign that makes longer wavelengths travel slower, and its refractive index rises very slightly with intensity. The two effects on a pulse are opposite, and at one particular pulse energy for each width they cancel — the same balance, with the nonlinear Schrödinger equation in place of Korteweg and de Vries’s. Mollenauer and colleagues observed it in 1980, and by 1988 had sent solitons four thousand kilometres.

The argument for using them was clean: a link’s bit rate is limited by how far pulses spread before they overlap, and a pulse that does not spread removes the limit. Through the 1990s soliton transmission was widely expected to be how long-haul optical communication would work.

It is not how it works. Two things defeated it.

The first is noise. Every optical amplifier along the line adds spontaneous emission, and a soliton that absorbs some of it has its carrier frequency shifted slightly at random. Because the fibre is dispersive, a frequency shift is a speed shift, so the pulse arrives early or late — and the accumulated timing jitter grows as the cube of the distance. That is the Gordon–Haus effect, it is intrinsic to using a nonlinear pulse in an amplified line, and it sets a hard limit on the product of rate and distance.

The second is that the competition improved faster. Wavelength-division multiplexing puts dozens of channels down one fibre, and solitons in different channels collide as they pass and shift each other’s timing — the same phase shift this essay measures, now a source of error. Meanwhile the linear approach acquired erbium amplifiers, dispersion compensation that undoes the spreading with a length of oppositely-dispersive fibre, and error-correcting codes strong enough to work with badly smeared pulses. Managing dispersion turned out to be easier than cancelling it.

What survived is worth naming, because it is not nothing. Dispersion-managed transmission, in which the sign of the dispersion alternates along the line, grew directly out of the soliton work and is how every long link is built. And mode-locked fibre lasers, which produce the shortest pulses available, are soliton devices outright — the pulse circulating in the cavity is held together by exactly this balance. The physics was correct; the system it was expected to become went another way.

What it is not

Two confusions are common enough to be worth naming.

A soliton is not a wave packet. A packet whose envelope moves at the group velocity is a purely linear object, and it spreads: the components disagree about speed, the phases drift apart, and the envelope widens without limit while its height falls to match. Nothing in a linear medium stops that. What makes the soliton different is the nonlinear term, which is absent from every calculation about group velocity — so the two look alike in a snapshot and behave oppositely over time.

And a soliton is not a shock. Both are what a steepening front turns into, and which one appears depends on which term the medium supplies to oppose the steepening. Dissipation gives a shock: a discontinuity of finite thickness, losing energy. Dispersion gives a train of solitons: an undular bore, losing nothing. The same wave in the same water can produce either depending on how much friction the bed provides.

What the pictures cannot show

Every figure here is one-dimensional. In two dimensions the equation is different, and a solitary wave is generally unstable to bending: it develops a transverse wobble and breaks into lumps. The exceptions — the Kadomtsev–Petviashvili equation with the right sign — are what allow the X-shaped soliton interactions occasionally photographed on a flat beach.

The collision is drawn for a case with a clean separation. When the two speeds are close the merged profile does not split into two humps at all during the overlap, and the “two waves passing through one another” description has to be replaced by an exchange of identity between them.

Taller is faster, and narrower with it. Three solitary waves of the Korteweg–de Vries equation, drawn at the same instant. Each is a sech² profile and each is a solution — substituted back into u_t + 6uu_x + u_xxx = 0, they leave 5.0e-5. They are not free to be any shape or any size: the amplitude is half the speed and the width goes as one over the square root of it, so a taller solitary wave is a faster and a narrower one, and the three cannot be scaled independently. Amplitude times width squared is the same for all three to 0.10 per cent, measured off the drawn curves rather than derived. That single-parameter family is what the balance produces: the nonlinear term steepens the front at a rate that grows with height, the dispersive term spreads it at a rate that grows with curvature, and only one height goes with each width for the two to cancel. The picture cannot show the consequence that makes these objects remarkable, which is what happens when two of them meet.
Fig. 5 The same family at three other speeds. The tie between height and width is the whole of the constraint, and it is what makes a soliton a one-parameter object rather than a shape that can be chosen.

The integration is exact only because the initial condition was. Starting from a general pulse, the shedding of the dispersive tail takes a long time and never quite finishes; the figures showing a perfectly preserved shape start from an exact solitary wave, which is the sharpest way to state the claim and not the way a real one is made.

And nothing here is quantum. The solitons of a field theory — where the conserved quantity is a topological charge and the object is a particle — share the name and a good deal of the mathematics, and are a different subject.

The ladder from here

Later rungs on this anchor: the inverse scattering transform, which solves the equation exactly by turning it into a linear scattering problem and reading the solitons off as bound states; the infinite hierarchy of conserved quantities and where it comes from; the nonlinear Schrödinger equation and the difference between bright and dark solitons; and the Fermi–Pasta–Ulam recurrence, which is the numerical experiment that led to all of this and whose failure to thermalise is still not fully explained.

The neighbouring ladders are the front that steepens until it cannot, which is this equation with the dispersion removed, the packet that will not keep its shape, which is it with the nonlinearity removed, and the speed that depends on the length, where the dispersion relation this balance depends on is measured.

Part 5 of 6

This essay is one argument about Wave packets. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Conserved quantityDispersionGroup velocityIntegrable systemKorteweg de vriesNonlinear wavePhase shiftSolitary waveSolitonWave packets