Thermodynamics

The diffusion that slows the longer it is watched

A random walker whose steps come at irregular times still diffuses normally, provided the waits have a finite average. Let the waits have no average — mostly short, occasionally enormous, as they are for a molecule that keeps getting stuck — and three things change at once. The spread grows more slowly than time. A diffusion coefficient measured by following one particle depends on how long it was followed. And following one particle for a long time no longer tells what many particles do: averages along a path and averages over a crowd give different laws. Tracking single molecules inside living cells runs straight into all three.

Assumes: The jiggle that proved atoms · The walk that comes home

The equation that only runs forwards found a coin being tossed beneath the diffusion equation. The jiggle that proved atoms turned the jiggling of a pollen grain into Avogadro’s number, and the walk that comes home found that a walker in three dimensions usually never returns. The later arguments — the summer that reaches the cellar in December, the coupled flows, the gas that flows towards more of itself and the speed limit that does not care how big the molecules are — all used the same foundation: a walk whose mean-square displacement grows in proportion to time, with a diffusion coefficient as the constant.

The last of those ended by noting that inside living cells the foundation cracks. Proteins in cytoplasm and membranes are measured to spread more slowly than their size predicts, and often more slowly than in proportion to time. The walk that comes home had listed anomalous diffusion among the things it did not treat. This essay treats one mechanism for it, the simplest, and finds that it does more than slow diffusion down. It breaks the assumption that a long observation of one particle tells the same story as a short observation of many.

A walk that waits

Take the ordinary random walk — a step left or right with equal chance — and change only when the steps happen. Between each step and the next the walker waits for a time drawn at random. If the waits have a finite average, ⟨τ⟩\langle\tau\rangle, the walk at long times is ordinary diffusion with a diffusion coefficient of one step squared per 2⟨τ⟩2\langle\tau\rangle: the irregularity of the timing averages away, as the central limit theorem says it must.

Suppose instead that the walker keeps getting stuck: in a pocket of a crowded cell, a trap in a disordered solid, a mesh of a gel. If the traps come in many depths and escape from each is thermally activated, the waiting times can have a distribution with a long tail, falling as a power, t−1−αt^{-1-\alpha}. Nothing keeps a magnetisation for ever found that an escape time depends exponentially on a barrier height; a spread of barrier heights turns into a spread of escape times covering many decades, and an exponential spread of barriers becomes a power-law spread of times. For α\alpha less than one the average wait is infinite: long waits are rare, but not rare enough for their average to converge.

Two random walks, one of which keeps stopping. Two random walks of unit steps over 2000 units of time, seeded. In each, the time between steps is drawn from a distribution falling as t^(−1−α). With α = 1.5 (thin line) the mean wait is finite and the walk jitters steadily. With α = 0.5 (thick) the mean wait is infinite: most waits are short but some are enormous, and the walk sits still for long stretches — its longest pause here lasts 558 time units, against 100 for the other. Each individual step is the same, left or right with equal chance; only the waiting differs. A molecule wandering through a cell, a charge carrier in a disordered solid and a particle in a gel all show pauses of this kind, trapped for a while in a pocket of the medium and then released.
Fig. 1 Two seeded random walks of unit steps over 2,000 time units, with waits between steps falling as t−1−αt^{-1-\alpha}: α = 1.5 (thin) and α = 0.5 (thick). The first jitters steadily. The second sits still for long stretches; its longest pause here lasts 558 time units, against 100 for the other. Each step is the same; only the waiting differs.

The figure draws one walk of each kind, seeded so the drawing is reproducible. The walk with α=1.5\alpha = 1.5 has a finite mean wait and wanders steadily. The one with α=0.5\alpha = 0.5 spends most of its time stuck: in two thousand time units its longest pause lasts over five hundred, a quarter of the record. Every step it does take is identical to the other walk’s. The whole difference is in the distribution of waits, and the drawing already suggests what follows: a path dominated by its longest pause is a path whose behaviour depends on when it is watched.

A sum dominated by its largest term

What an infinite mean does to a sum is the key to everything that follows, and it is worth seeing directly. Add up NN waiting times drawn from a distribution with a finite mean and the total is close to NN times the mean, with fluctuations of order N\sqrt{N} — the central limit theorem, which the delay that is a random variable used to turn many small contributions into a normal distribution. Draw them from a distribution falling as t−1−αt^{-1-\alpha} with α<1\alpha < 1 and the theorem fails. The largest of NN such waits is typically of order N1/αN^{1/\alpha}, which for α=0.5\alpha = 0.5 is N2N^2, and the sum of all the others is of the same order: the total time is dominated by a handful of the longest waits, and grows faster than the number of steps.

Turned round, the number of steps taken in a time tt grows only as tαt^\alpha — the square root of time, for α=0.5\alpha = 0.5 — and since each step adds one unit of mean-square displacement, so does the spread. The fluctuations never become small relative to the mean, because the next long wait is always comparable with everything that came before. Sums of this kind converge not to a normal distribution but to one of Lévy’s stable distributions, which have the same heavy tail as the terms, and the whole of the anomalous behaviour below is that replacement of the central limit theorem.

Where heavy tails come from

A power-law distribution of waiting times needs a reason, and the commonest is a spread of trap depths. A walker held in a trap of depth EE escapes at a rate set by the Boltzmann factor, so its mean waiting time is τ0eE/kT\tau_0 e^{E/kT} — the exponential the exponential that decides everything found governing every thermally activated process. If the traps have depths spread exponentially, with a probability falling as e−E/E0e^{-E/E_0}, then the waiting times, averaged over traps, fall as a power: t−1−kT/E0t^{-1-kT/E_0}. The exponent is α=kT/E0\alpha = kT/E_0, the temperature measured in units of the typical trap depth.

That gives the model a physical dial. Hot enough, with kT>E0kT > E_0, the mean wait is finite and diffusion is normal. Cold enough, below a temperature set by the spread of traps, the mean diverges and every anomaly above appears, with an exponent that falls in proportion to the temperature. This is the model Harvey Scher and Elliott Montroll introduced in 1975 to explain why photocurrents in amorphous semiconductors, used in the photocopiers of the time, arrived at their electrodes with a long tail that no ordinary diffusion could produce, and whose shape changed with temperature in exactly this way.

The spread that falls behind time

Average over many walkers, and the mean-square displacement no longer grows in proportion to time.

The spread that grows more slowly than time. The mean-square displacement of 1500 seeded walkers against time, on logarithmic axes, for waiting-time exponents α = 1.50, 0.75, 0.50. With a finite mean wait (α = 1.5) the spread grows in proportion to time, the signature of ordinary diffusion. With an infinite mean wait it grows as a power of time less than one: the measured slopes over the last decade are 0.98, 0.69, 0.49, against the predicted 1.00, 0.75, 0.50. Sub-diffusion is not slow diffusion. No choice of diffusion coefficient makes a straight line of slope one pass through these points: the walker's effective diffusion coefficient keeps falling for as long as it is watched, because the longer it walks, the longer the longest trap it has fallen into.
Fig. 2 The mean-square displacement of 1,500 seeded walkers against time, on logarithmic axes, for α = 1.5, 0.75 and 0.5. The measured slopes over the last decade are 0.98, 0.69 and 0.49, against the predicted 1, 0.75 and 0.5.

For α=1.5\alpha = 1.5 the curve has slope one on logarithmic axes: ordinary diffusion. For α=0.75\alpha = 0.75 and 0.50.5 the slopes are α\alpha itself — the number of steps a walker has taken by time tt grows only as tαt^\alpha, because the longer it walks, the longer the longest wait it has met, and each step contributes a unit of mean-square displacement. The figure measures the slopes from the simulation and finds them within a few hundredths of the prediction.

That is sub-diffusion, and it is not slow diffusion. Slow diffusion would be a line of slope one lower down; these lines have a different slope, and no diffusion coefficient describes them. The best one can do is quote an effective coefficient at a particular time, and it keeps falling: the walkers are not moving through a thick medium at a steady low rate, they are moving at the ordinary rate between ever longer pauses. The curve that is really a staircase found a magnet’s response made of avalanches with no typical size; here the walk is made of pauses with no typical length, and the absence of a typical scale is again the whole story.

One path against the crowd

The deeper consequence appears when the walk is watched the way single-molecule experiments watch it.

A single-particle tracking experiment follows one molecule for a long time and computes a time average: for each lag Δ\Delta, the squared displacement over every window of length Δ\Delta along the record, averaged. For ordinary diffusion this time average and the ensemble average — many particles, each watched once — are the same thing: the system is ergodic, and a long enough single record reproduces the crowd. The jiggle that proved atoms relied on this; Perrin’s measurements of a few grains over long times gave the ensemble’s diffusion coefficient because the two averages agree.

What one trajectory says, and what the crowd says. For the walk with α = 0.5: the time-averaged mean-square displacement of each of 12 single trajectories, each watched for 10000 time units — the average of the squared displacement over every window of a given lag along that one path — against the lag, on logarithmic axes (thin lines), and the mean-square displacement of an ensemble of 800 walkers at the same times (thick). The ensemble spreads as t^0.5. Every single trajectory's time average grows nearly in proportion to the lag — mean slope 0.82, where the ensemble's is 0.5 — as if it were ordinary diffusion, and the trajectories disagree with one another about the diffusion coefficient by more than a factor of ten. A time average along one path and an average over many paths, which for ordinary diffusion are the same thing, here give different laws. That is weak ergodicity breaking, and it is what a single-molecule tracking experiment has to be read against.
Fig. 3 For α = 0.5: the time-averaged mean-square displacement of twelve single trajectories, each watched for 10,000 time units, against the lag, on logarithmic axes (thin), and the ensemble’s mean-square displacement (thick). The ensemble grows as t0.5t^{0.5}; each single trajectory’s time average grows nearly in proportion to the lag (mean slope 0.85), and they disagree about the coefficient by more than a factor of ten.

For α=0.5\alpha = 0.5 they do not agree. The figure computes the time-averaged displacement for twelve single trajectories and lays the ensemble’s mean-square displacement over them. Each single trajectory’s time average grows almost in proportion to the lag, as if the particle were diffusing normally, while the ensemble grows as the square root of time. And the single trajectories scatter: one says the particle diffuses fast, another that it barely moves. An experimenter who tracked one of them, found a straight line of slope one, and concluded that the molecule diffuses normally with a certain coefficient would be wrong twice: about the law, and about the coefficient.

This failure is called weak ergodicity breaking, a name given by Jean-Philippe Bouchaud in 1992 and analysed for this walk by Yong He, Stanislav Burov, Ralf Metzler and Eli Barkai in 2008. It is weak because nothing in the system is split into regions a particle cannot reach, as in a glass or a broken symmetry. Every trajectory can go everywhere. It simply takes too long for the time average to settle, because the waiting-time distribution has no mean to settle towards.

A scatter that does not shrink

The scatter between trajectories can be measured directly.

How much single trajectories disagree. The distribution of the time-averaged mean-square displacement at a lag of 10, over 400 seeded trajectories each watched for 10000 time units, divided by its average: for α = 1.5 and α = 0.5. For the ordinary walk the values cluster near one, with a relative variance of 0.009: one long trajectory gives the diffusion coefficient. For α = 0.5 they spread from near zero to three or four times the mean, with a relative variance of 0.66 — the theory for an infinitely long record gives 0.57 — and it does not shrink as the record lengthens. Repeating the measurement on the same kind of particle in the same medium does not converge; the spread is a property of the process, not of the noise.
Fig. 4 The distribution of the time-averaged displacement at a lag of 10, over 400 trajectories each watched for 10,000 time units, divided by its average, for α = 1.5 (outline) and α = 0.5 (bars). The first clusters near one with a relative variance of 0.009; the second spreads from near zero to three or four times the mean, relative variance 0.66, against 0.57 predicted for an infinitely long record.

For the ordinary walk the time averages cluster tightly about their mean: a relative variance of less than a hundredth, which would shrink further with longer records. One long trajectory tells the truth. For α=0.5\alpha = 0.5 the time averages spread from nearly zero — trajectories that spent most of their time in one long trap — to three or four times the mean. The relative variance is 0.66 here and the theory for an infinitely long record gives 0.57: it does not go to zero as the record lengthens. Repeating the measurement on another molecule of the same kind in the same medium gives a different answer, and no amount of care makes them agree, because the disagreement is the physics.

In single-molecule experiments on living cells exactly this kind of scatter has been reported — potassium channels in the membranes of human kidney cells, tracked by Aubrey Weigel and colleagues in 2011, and insulin granules and lipid particles in other experiments — and the spread of time averages, rather than any single one, was the evidence that the underlying motion involves long, heavy-tailed trapping. Experiments that report a distribution of diffusion coefficients from many tracked molecules are often measuring this, not a genuine variety of molecules.

A clock inside the measurement

The last consequence is the one most likely to catch an experiment out.

A diffusion coefficient that depends on how long it was measured. The time-averaged mean-square displacement at a fixed lag of 10, averaged over 150 trajectories, against the length of the record, on logarithmic axes, for α = 1.50, 0.75, 0.50. For the ordinary walk it does not depend on the record's length (slope −0.02). For α < 1 it falls as the record lengthens, as T^(α − 1): measured slopes −0.30 and −0.51, against −0.25 and −0.50. The same particle watched for an hour appears to diffuse more slowly than when watched for a minute, because a longer record is more likely to contain a very long trap. A diffusion coefficient quoted for such a system without the duration of the measurement is not a number.
Fig. 5 The time-averaged displacement at a lag of 10, averaged over 150 trajectories, against the length of the record, on logarithmic axes, for α = 1.5, 0.75 and 0.5. For the ordinary walk it does not depend on the record’s length (slope −0.02); for α < 1 it falls as the record lengthens, slopes −0.30 and −0.51 against the predicted −0.25 and −0.50.

For the ordinary walk, the time average at a fixed lag does not depend on how long the particle was followed. For the heavy-tailed walk it falls as the record lengthens, as Tα−1T^{\alpha - 1}. Follow a particle for ten times longer and, for α=0.5\alpha = 0.5, its apparent diffusion coefficient is three times smaller. The reason is the same as before: a longer record is more likely to include a very long trap, and a trajectory that spends most of its time trapped has a small time-averaged displacement at every lag.

The system ages. Its properties depend on how long ago the observation began, relative to when the process started, in a way that never settles. The paste that forgets it was stirred found a paste whose stiffness recorded how long it had rested; here a diffusion coefficient records how long it was measured. Both are cases of a system that never reaches a steady state on any timescale an experiment can reach, and in both, a number quoted without its time is incomplete.

The dependence runs the other way in time as well. If the walk has been going for a while before anyone starts watching, the walker is more likely than not to be caught in a long trap at the moment observation begins, because long traps occupy most of the elapsed time. An experiment that starts late therefore sees a population in which a sizeable fraction of walkers do not move at all during the window, and the rest move normally — an apparent split into mobile and immobile particles that is really one population caught at different points of its waits. The fraction that seems immobile grows with how late the watching starts, and for a heavy enough tail it approaches everyone. Experiments on membrane proteins that report a “bound” and a “free” population, with the balance shifting over the course of an experiment, are sometimes seeing this: not two kinds of molecule, but one kind watched at different ages of its walk.

Other ways to be anomalous

Heavy-tailed trapping is one mechanism for anomalous diffusion and not the only one, and the distinction matters because the others do not break ergodicity. A particle moving through a viscoelastic medium — the cytoplasm’s network of filaments, a polymer solution — can sub-diffuse because the medium pushes back with a memory, so that each displacement is followed by a slight tendency to reverse. That motion, called fractional Brownian motion, gives a mean-square displacement growing as tαt^\alpha too, but it is ergodic: the time average of one long trajectory reproduces the ensemble. A particle in a crowded, fractal-like maze of obstacles gives yet another variety. Telling these apart in an experiment requires exactly the analysis the figures above perform: comparing time and ensemble averages, measuring their scatter, and checking whether the answer depends on the length of the record. The same growth law can come from three different kinds of physics, and only the ergodic properties separate them.

There is also the opposite anomaly. If the lengths of the steps, rather than the waits between them, have a heavy-tailed distribution, the walk makes occasional enormous jumps and spreads faster than diffusion: super-diffusion, of the kind called a Lévy flight. The spread of light through certain disordered glasses, and some patterns of animal foraging, have been described that way. The heavy tail is on the other side, and it hastens the walk rather than stalling it.

What the simulation leaves out

The walks here are one-dimensional, on a lattice, with waiting times drawn independently from one fixed distribution. Real trapping has a spatial structure: a molecule leaving a trap enters a neighbouring region with its own traps, and a trap that caught it once may catch it again, correlations the model ignores. The power-law tail is also an idealisation; real traps have a deepest level, so the waiting times have a cutoff, and beyond the corresponding time the diffusion becomes normal again. Whether that crossover time is inside or outside an experiment’s window decides whether the anomalies above are seen, which is why the same molecule can appear anomalous in one experiment and normal in another.

Still open: what the cytoplasm is

For the interiors of living cells the question that the speed limit that does not care how big the molecules are left open is still not settled. Experiments find sub-diffusion in some cells and not others, for some molecules and not others, and they disagree about which mechanism produces it — heavy-tailed binding to cellular structures, the viscoelasticity of the cytoskeleton, or crowding — and about how much the active stirring of the cytoplasm by molecular motors, which consumes energy and has no counterpart in a passive fluid, changes the picture. Reading single-molecule data correctly requires deciding between mechanisms that give the same mean-square displacement and differ only in their ergodic properties, and the statistical tools for doing so on short, noisy trajectories are an active field.

The habit worth carrying away is to ask whether a time average and a crowd average can be trusted to agree. A walk whose waits have no mean spreads more slowly than time, and its single long trajectories each look like ordinary diffusion with a different coefficient that keeps falling as the record lengthens — so one molecule followed for an hour and many molecules followed for a minute give different laws, and both are right. When the longest pause in a record is comparable with the record itself, the record is measuring its own length.

Part 9 of 9

This essay is one argument about Diffusion. The others:

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

AgeingAnomalous diffusionDiffusion coefficientErgodicityHeavy tailed distributionMean square displacementRandom walkSingle particle tracking