Mechanics

The fold that has to be there

Two trajectories that separate exponentially, in a region they can never leave, are being asked to do two incompatible things. The resolution is that the motion is folded back on itself over and over, and the object that survives infinitely many foldings is neither a curve nor a patch of surface — it has a dimension between the two, and the number can be measured two entirely different ways.

Assumes: The error that doubles on a schedule · The last curve to go

A chaotic system does two things that cannot both be true of a smooth flow, and the reason it is worth a figure is that neither of them is mysterious on its own.

The set an orbit that never repeats settles onto. 24,000 successive positions of one orbit of the map x' = 1 − 1.4x² + y, y' = 0.3x, after five hundred steps of transient have been discarded. Nearby points separate at e^0.4188 per step, so the orbit is unpredictable in the way the rung below measures; and every one of the 24,000 points lies inside a box 2.558 by 0.767, a diagonal of 2.670, so it is going nowhere. Those two statements are not compatible with a smooth stretching: something has to bring the separated points back, and the bringing back is the visible fold at the left-hand end. The curve is not a curve. Every strand of it is a bundle of strands at any magnification, which is what an area contraction of 0.3 per step leaves behind when the stretching along the other direction is e^0.419. The 24,000 points paint 7,352 distinct marks at the resolution this is drawn at, which is itself a measurement of how little of the plane the set occupies.
Fig. 1 Twenty-four thousand successive positions of a single orbit, after the transient has been discarded. Nearby points separate at a fixed exponential rate, and every point lies inside a box less than three units across. The picture is what those two facts together are forced to produce.

The first is that two trajectories starting a hair’s breadth apart separate exponentially. That is the measurement the first rung on this ladder makes, and it is what makes prediction expensive: every extra decade of accuracy in the starting conditions buys a fixed extra amount of time and no more.

The second is that the motion is bounded. The orbit never leaves a region a few units across, and it never will, because the region is where the system’s energy budget and its dissipation put it.

Neither statement is strange. Put together they are a contradiction unless something specific happens.

Why the two facts cannot both hold for a smooth stretching

Take a small blob of starting conditions and follow it. The exponential separation says its longest dimension grows by a fixed factor each step, so after enough steps that dimension is arbitrarily large. Boundedness says the blob is still inside a box a few units across.

A long thin object inside a small box has to be bent.

Separation, until there is nowhere further to separate. The distance between two orbits started 1e-9 apart, on a logarithmic scale, for 90 steps. The first stretch is a straight line of slope 0.1683 decades per step — 0.3874 in natural logarithms, against the 0.4188 the Jacobians give — and then it stops, flat, at the width of the attractor itself: 2.670, measured off eight thousand points of the orbit rather than assumed. That ceiling is the whole difficulty. Exponential growth with no ceiling is an instability and the system flies apart; exponential growth with a ceiling means the trajectories that ran apart have to be brought back together again, and brought back into a region they have already been through. Every return is a fold, and the folds are what the attractor is made of.
Fig. 2 The distance between two orbits started a billionth apart, on a logarithmic scale. Straight for the first forty steps, at the rate the Jacobians predict, and then flat at the width of the attractor — which is measured off the orbit rather than assumed.

The plot above is the whole argument in one curve. The straight portion is the stretching; the flat portion is the box. What has to happen at the join is that trajectories which ran apart are brought back together, and brought back into a region they have already visited.

That bringing back is the fold. It is not an extra ingredient. It is what boundedness costs once exponential separation has been granted, and any system with both must have it.

The rate on the straight part is 0.38740.3874 per step, fitted over the stretch of the run that is still exponential; the same rate computed from the accumulated Jacobians of the map, without following any pair of orbits, is 0.41880.4188. The two agree to within the noise of a single realisation, which is what fitting one pair of trajectories buys. The second method averages over the whole orbit and is the one the rest of this essay uses.

One step, taken apart

The system drawn here is the smallest thing that does all of this: a map of the plane onto itself,

x=1ax2+yy=bxx' = 1 - a x^2 + y \qquad y' = b x

with a=1.4a = 1.4 and b=0.3b = 0.3. It has no time in it, no forces, and one nonlinearity. Hénon wrote it down in 1976 as a stripped-down stand-in for a slice through a continuous flow, and it does everything the flow does.

A rectangle, and where one step of the map sends it. A rectangle 2.7 wide and 0.8 tall, drawn as its boundary, beside the image of that boundary under one application of the map. The straight top edge comes back as a parabola, because the map's only nonlinear term is the square, and the two ends of the rectangle are brought round to the same side — which is the fold. The area falls from 2.1600 to 0.6480, a ratio of 0.3000 against the exact 0.3 the determinant of the Jacobian gives at every point. Repeat this and the same thing happens to the image: bent again, thinned again, laid alongside what is already there. The attractor is what is left after infinitely many of these, and the reason it has a fractional dimension is that every one of the folds survives in it.
Fig. 3 A rectangle covering the attractor, and the image of that rectangle after a single application of the map. The straight top edge comes back as a parabola, the two ends are brought round to the same side, and the area falls by a factor the Jacobian’s determinant fixes at every point.

The image of the rectangle is the mechanism, drawn once. Three things happen at the same time.

It is stretched. Along one direction the image is longer than the rectangle was, which is the positive Lyapunov exponent expressed as a shape rather than as a number.

It is squeezed. Along the other direction it is thinner. The determinant of the map’s Jacobian is b-b at every point without exception, so the area of any region is multiplied by exactly 0.30.3 per step — measured on the drawn boundary at 0.30000.3000, by the shoelace formula, which is the check that the boundary really was mapped and not merely sketched.

It is bent. The only nonlinear term is x2x^2, and a parabola is what a straight line becomes under it. The bend brings the two ends of the rectangle round to the same side, which is what puts a stretched object back inside a box.

Repeat and the image is bent again, thinned again, and laid down alongside material that is already there. What the orbit settles onto is what remains after infinitely many of those, and the last figure of this essay is what that looks like at three magnifications.

The dimension is not an integer, and two methods agree on it

An adjective — intricate, complicated, infinitely detailed — is not a measurement. The measurement is a dimension, and it can be taken in a way that has nothing at all to do with the pictures.

Counting boxes, and getting a number between one and two. How many boxes of side one part in k contain at least one of 60,001 points of the orbit, against k, both on logarithmic axes. A curve would give a slope of one and a filled region a slope of two; the fitted slope is 1.292, over the 6 scales where the count is neither trivially small nor limited by how many points there are to find. The independent check is that the two Lyapunov exponents already determine a dimension without any counting: 1 + λ₁/|λ₂| = 1.258, which is the fraction of the contracting direction the stretching manages to keep. The two agree to 0.034. A number between one and two is the arithmetic saying what the picture says: the fold has been applied infinitely often, so the set is thicker than a line everywhere and thinner than a patch everywhere.
Fig. 4 How many boxes of a given size contain at least one point of the orbit, against how many boxes there are across the attractor, on logarithmic axes. The slope is the box-counting dimension. Hollow points are scales at which the boxes outnumber the sample, where the count measures the sample rather than the set.

Cover the plane with boxes of side ε\varepsilon and count how many contain at least one point of the orbit. For a curve the count grows as 1/ε1/\varepsilon; for a filled region as 1/ε21/\varepsilon^2. The slope on logarithmic axes is therefore a dimension, and for a curve or a patch it comes out an integer.

Here it comes out 1.2921.292.

The second route uses no boxes and no counting. The two Lyapunov exponents of this map are λ1=+0.4188\lambda_1 = +0.4188 and λ2=1.6228\lambda_2 = -1.6228 per step, computed by following how a pair of vectors is deformed by the product of Jacobians along the orbit. Kaplan and Yorke’s argument says that a set which is stretched at λ1\lambda_1 and squeezed at λ2\lambda_2 retains a fraction λ1/λ2\lambda_1/|\lambda_2| of the contracting direction, giving a dimension of

D=1+λ1λ2=1.258.D = 1 + \frac{\lambda_1}{|\lambda_2|} = 1.258.

The two numbers agree to 0.030.03. They are computed from different data by different means, and neither could return a number between one and two for an object a finite list of curves could describe.

There is a third check available and it is free. The exponents must sum to the logarithm of the area contraction, because the sum of the exponents is the logarithm of the determinant. Here λ1+λ2=1.204\lambda_1 + \lambda_2 = -1.204 and ln0.3=1.204\ln 0.3 = -1.204. Anything wrong with the way the Jacobians were accumulated would show up there first, and the generator refuses to draw if it does not hold.

Where the contraction comes in, and why it is not the enemy

The area contraction is doing two jobs and it is easy to notice only one of them.

The obvious job is that it makes an attractor exist at all. A region of starting conditions shrinks to zero area, so the long-term behaviour lives on a set of no area, and that set is what every orbit ends up on regardless of where it began.

A one-dimensional well makes the contraction familiar: dissipation takes every starting condition to the same point, and volume in phase space shrinks to nothing. The contraction here does the same work and arrives somewhere very different — a set of zero volume that is not a point, not a curve, and not anything with an integer dimension. Same shrinking, incompatible destination, and the difference is the stretching that accompanies it.

The less obvious job is that the contraction is what keeps the layers distinguishable. Each fold lays a new sheet alongside the old ones; if nothing thinned them they would merge and there would be nothing to see at small scales. It is the factor of 0.30.3 per step that keeps every generation of sheets thinner than the last, which is exactly the condition for structure to survive at every magnification.

So the intuition that dissipation destroys structure is the wrong way round here. A system that preserves area — a Hamiltonian one, for instance — has no attractor at all, and the second rung on this ladder is about what that looks like: invariant curves that survive or break, in a phase space where nothing shrinks. Strange attractors belong to systems with friction in them.

An area-preserving map has no contraction at all, and it behaves quite differently. Chaotic regions and surviving invariant curves coexist, neither attracting the other, because there is nothing to attract to — without dissipation there is no attractor, and the chaos is distributed through the space rather than collected onto a set. That is worth knowing because it separates two things routinely run together: chaos needs stretching and folding, and a strange attractor needs contraction as well.

The layering, at three magnifications

The set has no smallest feature. That claim is checkable in the crudest possible way: magnify it and see whether the strands become single strokes.

The same layering, three magnifications apart. The same region of the attractor at ×1, ×8, ×64, each panel centred on the same point and drawn from the same orbit. What looks like a single stroke in the first panel is several strokes in the second and several again in the third, and the counts — 54,725 points at ×1, 993 points at ×8, 141 points at ×64 — are of the same orbit, not of a new computation. There is no last magnification at which the strands become single: the fold that separated them happened at every scale on the way, because the map has been applied an unbounded number of times. That is what a fractional dimension means in a picture, and it is why the set cannot be described by any finite list of curves.
Fig. 5 The same region of the attractor at three magnifications, each panel centred on the same point of the orbit and drawn from the same run. What is one stroke in the first panel is several in the second and several again in the third.

They do not. What reads as a single line at the scale of the whole picture is a bundle at eight times, and each strand of that bundle is a bundle again at sixty-four. Nothing about the computation changes between panels; the same orbit is being drawn, and the points inside each window are simply fewer.

This is where the non-integer dimension stops being an abstraction. A dimension of 1.291.29 says the object is denser than a line everywhere and sparser than a patch everywhere, and the panels are what that sentence looks like.

The limit of the process is not something the pictures can show, and it is worth stating what is being taken on trust. The orbit is finite, the computer’s numbers have sixteen digits, and the panels stop at sixty-four times. What the argument supplies, rather than the pictures, is that the folding has been applied without limit, so there is a generation of sheets at every scale down to zero.

What it does to prediction, and what it does not

The practical consequence is the one the first rung measures: the time over which a prediction is worth anything grows only as the logarithm of the precision of the starting data. Ten times better instruments buy ln10/λ1\ln 10 / \lambda_1 more steps of usefulness, which here is five and a half — a little over two Lyapunov times, whatever the instrument cost.

But the same structure hands back something that is not a prediction and is often more useful.

A random walk’s individual trajectories are unpredictable and the statistic over many of them is smooth and exactly linear. That is the shape of what survives here too: the trajectory is unforecastable beyond a horizon set by the exponent, and the distribution on the attractor is stable, reproducible and computable. Prediction fails and statistics do not, which is the honest statement of what chaos costs.

The attractor itself is stable. Change the starting point anywhere in the basin and the orbit ends up on the same set, with the same dimension and the same exponents. Long-run averages over the orbit — how often it visits each region, what the average of any smooth quantity is — converge, and converge to numbers that do not depend on where the run started. The individual trajectory is unforecastable and the statistics of the trajectory are reproducible to as many digits as the run is long.

That is a different kind of knowledge from a forecast, and it is what makes statistical mechanics possible for systems nobody could integrate. Entropy as a count and the speeds in a still room both rest on trajectories nobody follows, and the reason the averages are trustworthy is the mechanism drawn here.

The same construction in a flow

A map is not a physical system. The reason to write one down is that a continuous flow in three dimensions, cut by a plane, is a map: record where the trajectory crosses the plane, and the next crossing is a function of the last.

A double pendulum is a continuous system with the same behaviour, and it shows why these pictures are drawn as maps. Its trajectory lives in four dimensions and cannot be drawn; take a slice through it — record the state each time the system passes through some surface — and what appears is precisely a stretch-and-fold map. The map is not an analogy for the flow. It is the flow, sampled.

Three dimensions is the minimum for a continuous flow to do this, and the reason is the same folding argument run backwards. Trajectories in a continuous flow cannot cross, because the equations give one velocity at each point. In two dimensions a bounded non-crossing trajectory has nowhere to go but onto a closed loop or a fixed point, which is the Poincaré–Bendixson theorem. In three there is room to pass over and under, and the folded sheet becomes possible without any trajectory meeting another.

A two-dimensional phase portrait cannot be chaotic at all, and the reason is worth stating because it bounds the whole subject. Trajectories in a plane cannot cross, so they have nowhere to go but onto closed curves or off to infinity — there is no room to stretch and fold. Three dimensions is the minimum for a flow, two for a map, and the separatrix in such a portrait is the only orbit that never repeats.

So the dimension count is not an accident of which system was picked. It is the reason chaos appears in a driven pendulum and not in a free one, in three coupled equations and not in two, and it is why the second rung’s area-preserving map needs two dimensions rather than one.

The horseshoe, and what it counts

The stretch-and-fold picture is a description, and it has an exact form which turns it into a theorem. The theorem is Smale’s, from the early 1960s, and it converts geometry into combinatorics.

Take a square. Stretch it in one direction, squeeze it in the other, bend the result into a horseshoe and lay it back across the original square, so that the horseshoe crosses it in two vertical strips. That is the map, and it is the folding of this essay stripped of everything inessential.

Now ask which points stay in the square for ever, under the map and its inverse. Following the construction through, the answer is a set that is a Cantor set in each direction — a set of no area, infinitely subdivided, exactly the kind of object the dimension measurement above reports.

The step that makes it powerful is a relabelling. At each step a surviving point is in one of the two strips; call them 0 and 1. Its whole history, forward and backward, is therefore a doubly infinite sequence of noughts and ones, and applying the map shifts that sequence one place along. The correspondence is exact: every sequence corresponds to exactly one point, and the dynamics is nothing but the shift.

Everything then follows by counting sequences, with no geometry at all.

A periodic orbit of period nn is a sequence that repeats every nn places, and there are 2n2^n of them. So the invariant set contains periodic orbits of every period, infinitely many of them, all unstable and all woven through one another.

Sequences that never repeat outnumber those that do, uncountably, so almost every orbit in the set is aperiodic.

Some sequence contains every finite block of symbols somewhere along it, so there is an orbit that comes arbitrarily close to every point of the set.

And two sequences agreeing on the first thousand places may differ at the thousand and first, so two points arbitrarily close together separate — which is sensitive dependence, obtained by looking at symbols rather than by measuring anything.

Two further facts make the horseshoe more than a worked example. It is structurally stable: perturb the map a little and the horseshoe survives, so none of this is fragile. And Smale proved that a horseshoe is present whenever a system has a transverse homoclinic point — a trajectory that leaves an unstable fixed point and returns to it, crossing itself.

Poincaré had found exactly such a crossing in the three-body problem in 1890, drawn the resulting tangle of intersections far enough to see what it implied, and written that he would not attempt to draw the figure, so struck was he by its complexity. Smale’s theorem says that what Poincaré declined to draw and what this essay’s figures do draw are the same object.

Stirring, which is the same operation

Stretch and fold is not only how a chaotic system behaves. It is how anything gets mixed, and the connection is exact rather than metaphorical.

Two fluids side by side mix by diffusion, and diffusion over any macroscopic distance is hopeless: the time goes as the square of the distance, and stirring a cup of tea by molecular diffusion alone would take months. What stirring does is not to move the molecules further but to stretch the interface between the two fluids, so that diffusion has only microns to cross rather than centimetres.

Stretching an interface inside a bounded container is exactly the situation this essay opened with. The interface’s length grows exponentially while it is confined, so it must fold, and it folds again at every turn of the spoon. The layers get thinner exponentially, and mixing is complete when they are thin enough for diffusion to finish in the time available.

Turbulence supplies the stretching at high Reynolds number. The interesting case is when it cannot. In a microfluidic channel a few tens of micrometres across, or in a polymer melt, or in the Earth’s mantle, the Reynolds number is so small that the flow is smooth and laminar and turbulence is impossible.

Such a flow can still mix chaotically, and that is the result that founded a subject. The velocity field may be perfectly smooth and perfectly periodic in time, and the trajectories of the fluid particles in it can still be chaotic — because the trajectories obey a map, and a map does not need a complicated velocity field to fold. Aref named it chaotic advection in 1984.

The engineering consequence is a device. A microfluidic channel with a herringbone pattern of grooves cut into its floor drives a transverse circulation whose direction alternates along the channel, so a fluid particle sees a two-stage map repeated. That is a stretch and a fold, the striation thickness falls exponentially rather than as a square root, and the length of channel needed to mix two streams drops from metres to millimetres.

The oldest version of the same operation is in a bakery. Roll the dough out and fold it in half; roll and fold again. The mathematician’s name for the idealised map — the baker’s transformation — is not a joke, and its Lyapunov exponent is the logarithm of two, one bit of stretching per fold. A taffy puller does the same with three rods rather than two hands, and its exponent is the logarithm of the golden ratio.

Where the model runs out

The map is not a physical system and its numbers are not physical numbers. a=1.4a = 1.4 and b=0.3b = 0.3 are chosen values in a parameter plane, and the attractor exists only for part of that plane; elsewhere the orbit runs to infinity, which the generator refuses to draw rather than showing the first few hundred points before it escapes.

Counting boxes, and getting a number between one and two. How many boxes of side one part in k contain at least one of 90,001 points of the orbit, against k, both on logarithmic axes. A curve would give a slope of one and a filled region a slope of two; the fitted slope is 1.280, over the 7 scales where the count is neither trivially small nor limited by how many points there are to find. The independent check is that the two Lyapunov exponents already determine a dimension without any counting: 1 + λ₁/|λ₂| = 1.258, which is the fraction of the contracting direction the stretching manages to keep. The two agree to 0.022. A number between one and two is the arithmetic saying what the picture says: the fold has been applied infinitely often, so the set is thicker than a line everywhere and thinner than a patch everywhere.
Fig. 6 The same count with a longer orbit and one more scale. The extra scale is usable only because there are more points, which is the general shape of the difficulty: a box count is a measurement of the set only where the sample is dense enough to find it.

The box count measures the sample below a certain scale. Once boxes outnumber points every occupied box holds one point and the count saturates at the number of points, giving a slope that drifts towards zero. The fit here uses only the scales where that is not happening, and the hollow markers show which were excluded. A dimension quoted without that exclusion is a number about how long the run was.

Both dimensions are the crudest of a family. Box counting weights an occupied box the same whether it holds one point or a thousand, so it measures the support and not the distribution on it. The information dimension and the correlation dimension weight by occupancy and come out slightly smaller. For this map they differ in the second decimal place; for a measure with strongly varying density they differ much more, and quoting one of them as “the” dimension is then a mistake.

Kaplan–Yorke is a conjecture, not a theorem. It is proved for special cases and holds numerically in a great many others, and there are constructed counterexamples. Its agreement here is evidence that both computations are right rather than a derivation of either.

And nothing here is about a system with noise in it. Add a small random kick at each step and the fine layers below the size of the kick are washed out, the box count flattens at that scale, and the measured dimension falls. Every experimental determination of a fractal dimension is a determination over the range of scales between the instrument’s resolution and the size of the system, and the honest ones say so.

Separation, until there is nowhere further to separate. The distance between two orbits started 1e-12 apart, on a logarithmic scale, for 90 steps. The first stretch is a straight line of slope 0.1739 decades per step — 0.4003 in natural logarithms, against the 0.4188 the Jacobians give — and then it stops, flat, at the width of the attractor itself: 2.670, measured off eight thousand points of the orbit rather than assumed. That ceiling is the whole difficulty. Exponential growth with no ceiling is an instability and the system flies apart; exponential growth with a ceiling means the trajectories that ran apart have to be brought back together again, and brought back into a region they have already been through. Every return is a fold, and the folds are what the attractor is made of.
Fig. 7 The same separation experiment started a thousand times closer. The straight portion is longer and its slope is the same, which is what makes the exponent a property of the system rather than of the experiment.

The ladder from here

Later rungs on this anchor: the route by which a system becomes chaotic as a parameter is turned, and the universal numbers that route obeys; the reconstruction of an attractor from a single measured time series, which is how the dimension of a real system is obtained without knowing its equations; control of chaos, where the unstable periodic orbits woven through the attractor are used rather than avoided; and transient chaos, where the folding is present but the set it builds is not an attractor and every orbit eventually leaves.

The neighbouring ladders are the error that doubles on a schedule, which measures the stretching, the last curve to go, which is what happens when the contraction is absent, and the equation that only runs forwards, where a reversible microscopic law produces an irreversible average by a mechanism this essay is a picture of.

Part 3 of 5

This essay is one argument about Chaos. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

AttractorBox countingChaosDeterminismDissipationFractal dimensionKaplan yorkeLyapunov exponentPhase space volumePredictabilityStrange attractorStretching and folding