Thermodynamics

The second law, with a probability attached

Entropy increases, on average. For a small system pulled quickly, individual runs go the other way — and how often is not a matter of taste but an exact number, fixed by a relation with no adjustable constant in it and no requirement that anything be near equilibrium.

Assumes: Entropy is a count, and the arrow of time is arithmetic · The bit that has to be paid for

The second law is usually stated as a prohibition: entropy does not decrease. Stated that way it is a statement about macroscopic systems, and for those it is as close to certain as anything in physics — the count of arrangements is so lopsided that the probability of a visible decrease is smaller than any number worth writing.

Take a system small enough for its own thermal motion to matter and the statement changes character. Individual runs do go the wrong way, sometimes, and the question of how often has an exact answer.

Runs that break the second law, and how often. The work done in a process repeated many times, and the same for the process run in reverse with its work reflected, for a free-energy change of 4 kT and a dissipation of 3 kT. The average work exceeds the free-energy change, which is the second law, and individual runs do not have to: the shaded tail is the fraction of runs that do less work than the free energy — trajectories in which the entropy of the universe went down — and it is 11.03% here. The two curves cross exactly at the free-energy change, whatever the dissipation, which is what makes an irreversible measurement able to report an equilibrium quantity.
Fig. 1 The work done in a process repeated many times, and the same for the reverse process with its work reflected. The average exceeds the free-energy change, which is the second law. The shaded tail is the fraction of runs that do less — trajectories in which the entropy of the universe went down — and it is a fifth of a per cent here.

The shaded region is not an error bar or a measurement uncertainty. Those runs really happened, really did less work than the reversible minimum, and really left the universe with less entropy than it started with. The law is not violated on average and it is violated in individual cases, and the two statements are consistent because the law was always about the average.

The relation that fixes how often

One straight line, with no adjustable constant in it. The logarithm of the ratio of the two distributions, against work. Crooks' theorem says it is a straight line of slope exactly one per kT, crossing zero at the free-energy change — and neither the slope nor the crossing depends on how far from equilibrium the process was taken, on how it was done, or on what the system is. That is what makes it useful: pull a molecule apart quickly, many times, and both the forward and reverse work distributions can be measured; where they cross is the equilibrium free energy of the transition, which no equilibrium measurement of that molecule may be able to reach.
Fig. 2 The logarithm of the ratio of the two distributions, against work: a straight line of slope exactly one per kT, crossing zero at the free-energy change. Neither the slope nor the crossing depends on how far from equilibrium the process was taken, on how it was done, or on what the system is.

Crooks’ theorem, from 1999, says that the probability of doing work WW going forwards and the probability of doing W-W going backwards are in the ratio e(WΔF)/kTe^{(W - \Delta F)/kT}. Nothing else appears — not the speed of the process, not the protocol, not the size of the system, not what it is made of.

Read at W=ΔFW = \Delta F, the ratio is one: the two distributions cross exactly at the free-energy change. Read in the tail, it gives the odds of a run in which entropy decreased. An entropy production of 10k-10k is e10e^{-10} as likely as the corresponding increase, which is one in twenty-two thousand; 50k-50k is one in 5×10215\times10^{21}; and by the time the system is macroscopic the numbers are of the size that makes the second law feel absolute.

So the transition from “sometimes” to “never” is not a change in the physics. It is an exponential in the size of the entropy change, and the entropy change scales with the system, and the exponential does the rest.

The theorem’s most useful feature is what it does not require. There is no assumption that the process is slow, or reversible, or close to equilibrium; the system need only start in equilibrium and be driven by a protocol that can be run backwards. That is what makes it applicable to processes that no equilibrium argument covers.

Reading a free energy out of an irreversible measurement

The crossing point is the practical part. A free-energy change is an equilibrium quantity, and the classical way to measure one is to do the process slowly enough that the work equals it — which for many systems is impossible, because slowly enough means slower than anything drifts.

Crooks’ theorem removes the requirement. Pull the system apart quickly, many times, and build the work distribution; push it back together quickly, many times, and build the reverse distribution; and where the two cross is the free-energy change. Neither measurement was reversible and the answer is the reversible one.

That is now standard in single-molecule work. A stretch of RNA is unfolded and refolded with an optical trap at speeds that dissipate several kT a time, the two distributions are accumulated over a few hundred pulls, and the crossing gives the folding free energy of a structure that could not be measured any other way. The first such measurement, on an RNA hairpin in 2005, agreed with the equilibrium value where an equilibrium value was available and gave one where it was not.

The technique’s requirement is that both directions be measurable, which is a real constraint: many processes cannot be run backwards at all, and for those the crossing method is unavailable.

The exact equality that does not help as much as it should

An exact equality that converges badly. The running estimate of the free-energy change from Jarzynski's equality — the exponential average of the work, which is exactly e to the minus the free energy — against how many runs have been averaged, for a process dissipating 3 kT. The equality is exact and holds however violently the process is done. The estimator is not: because the exponential weights small works most heavily, the average is dominated by rare trajectories in the low tail, and until enough of those have been seen the estimate sits near the mean work rather than near the free energy. After 200,000 runs it has come down to 3.95 kT against a true 4. The number of runs needed grows exponentially with the dissipation, which is why the method works on molecules and not on anything larger.
Fig. 3 The running estimate of the free energy from Jarzynski’s equality — the exponential average of the work — against how many runs have been averaged. The equality is exact. The estimator is dominated by rare low-work trajectories, so the estimate sits near the mean work until enough of those have been seen.

Jarzynski’s equality, from 1997, is the more famous statement and it follows from Crooks’ by an integration: the average of eW/kTe^{-W/kT} over the forward process is exactly eΔF/kTe^{-\Delta F/kT}. It needs only the forward direction, which removes the constraint above.

It also has a difficulty that the figure exists to show. The exponential weights small works most heavily, so the average is dominated by the rare trajectories in the low-work tail — precisely the ones in which entropy decreased. Until enough of those have been sampled, the estimate sits near the mean work rather than near the free energy, and the number of runs needed to sample them grows exponentially with the dissipation.

The figure’s run dissipates three kT and takes thousands of samples to come down. At ten kT it would take millions; at fifty, more than any experiment. So the equality is exact and the estimator is bad, and the two facts are entirely compatible.

That combination is worth recognising because it recurs. An identity that is exact for all sample sizes says nothing about how quickly an estimate built from it converges, and an average dominated by a rare tail converges slowly however exact the identity behind it. The remedies are the ordinary ones — dissipate less, or sample the tail deliberately — and neither is free.

Runs that break the second law, and how often. The work done in a process repeated many times, and the same for the process run in reverse with its work reflected, for a free-energy change of 4 kT and a dissipation of 0.4 kT. The average work exceeds the free-energy change, which is the second law, and individual runs do not have to: the shaded tail is the fraction of runs that do less work than the free energy — trajectories in which the entropy of the universe went down — and it is 32.74% here. The two curves cross exactly at the free-energy change, whatever the dissipation, which is what makes an irreversible measurement able to report an equilibrium quantity.
Fig. 4 The same process done gently: a dissipation of four tenths of a kT rather than three. The two distributions overlap heavily, the crossing is easy to locate from a handful of runs, and almost a third of the runs do less work than the free-energy change. Slower is better for the measurement and worse for the demonstration.

Comparing the two dissipations shows the trade the technique lives on, and it runs in an awkward direction.

A gentle process gives distributions that overlap, so the crossing point is well determined and the free energy comes out of a few dozen runs. It also gives entropy decreases that are common — a third of the runs, here — which makes the “violation” unremarkable and the demonstration weak.

A violent process gives entropy decreases that are rare and impressive, and distributions so far apart that no practical number of runs locates the crossing. The overlap region, which is where every bit of information about the free energy lives, is exponentially unlikely to be sampled at all.

So the experiment that shows the second law being broken most convincingly is the one that measures the free energy worst, and vice versa. Anyone designing such a measurement is choosing a point on that curve, and the sensible choice is a dissipation of a few kT — enough that the process is genuinely irreversible, little enough that the distributions still meet.

What this does and does not overturn

It is worth being careful about the claim, because the fluctuation theorems are sometimes reported as a refutation of the second law and are nothing of the kind.

What is new is not that small systems fluctuate; that has been known since Brownian motion. What is new is the exact relation between the probability of a fluctuation and the entropy it corresponds to, holding arbitrarily far from equilibrium, with no adjustable constant. Before these results, non-equilibrium statements were mostly inequalities and mostly restricted to near equilibrium. These are equalities and they are not restricted.

What is unchanged is everything about the macroscopic second law. The relations say the probability of a violation falls exponentially with its size in units of kk, and a macroscopic entropy change is of order 1023k10^{23}k, so nothing about a steam engine has moved. Maxwell’s demon does not become possible, because the demon’s cost is an accounting question rather than a fluctuation question.

And what is clarified is the status of the law itself. It is a statement about a probability distribution rather than a prohibition, its strength depends on the size of the system, and the size dependence is now a formula rather than a hand-wave. That is a change in what kind of statement it is, which is worth more than a change in what it says.

Where the entropy actually is

One point deserves care because it is where the statements are most often garbled. The quantity that appears in these theorems is the total entropy production — the system’s plus the bath’s — and for a system driven between equilibrium states it is the dissipated work over the temperature.

That means the “violations” are not runs in which a system spontaneously ordered itself while everything else stayed put. They are runs in which the system took less heat from the bath than the average, so the bath’s entropy increased by less than the system’s decreased. The total went down, and it went down because the bath happened to deliver an unusually convenient sequence of kicks.

The distinction matters for what such an experiment is evidence about. It is not evidence that a system can organise itself; it is evidence that the exchange with a bath is a fluctuating quantity whose fluctuations are quantitatively constrained. That is a statement about the bath as much as about the system, and it is the same statement as the one about a small system’s energy wandering with the accounting done over both.

The measurement of a bit, and what it costs

The work one molecule and one bit are worth. The pressure of a gas of one molecule against its volume, at 300 K, in units of the volume it starts in. The shaded area is the work the molecule does pushing a partition out isothermally, and it is measured here by integrating the drawn curve rather than written down: expanding by 1.5× yields 0.4055 kT against ln 1.5 = 0.4055; expanding by 2× yields 0.6931 kT against ln 2 = 0.6931; expanding by 4× yields 1.3863 kT against ln 4 = 1.3863; expanding by 8× yields 2.0794 kT against ln 8 = 2.0794, agreeing to 2.0e-10. The doubling is the one that matters, because a partition inserted in the middle leaves the molecule on one side or the other, and knowing which is what lets the load be attached to the right face. That single expansion delivers kT·ln2 = 2.87 zeptojoules at 300 K. It looks like work extracted from one temperature, and it is — until the engine is asked to run again, which requires forgetting which side the molecule was on.
Fig. 5 The smallest engine there is: one molecule in a box, a partition, and the knowledge of which side it went. Letting the molecule push the partition out isothermally extracts kT ln 2 of work, which is the same quantity a fluctuation of that size costs in probability — and the connection is not a coincidence.

The Szilard engine belongs beside the fluctuation theorems because it is where the two halves of the subject meet, and because it makes the numbers concrete.

One molecule in a box carries a fluctuation of order its own energy — it is on one side or the other, and which is a fair coin. Knowing which is one bit, and that bit is worth kT ln 2 of work and costs kT ln 2 to erase. The same kTln2kT\ln 2 appears in Crooks’ relation as the work difference corresponding to a factor of two in probability, and that is not two coincidences: entropy is a logarithm of a probability in both statements.

What the fluctuation theorems add is the general version. A trajectory whose entropy production is σ\sigma is eσ/ke^{\sigma/k} times more likely than its reverse; a bit of information is kln2k\ln 2 of entropy; so acquiring one bit shifts the odds of a process against its reverse by exactly a factor of two. Information and irreversibility are the same quantity in different units, and the exchange rate is fixed.

That has been measured. Experiments in the 2010s built Szilard engines from a single colloidal particle in a light trap, extracted work from measurements of its position, and confirmed both the kTln2kT\ln 2 per bit and the modified fluctuation theorem that includes an information term. The engine Szilard invented in 1929 as a thought experiment now runs on a microscope stage.

Where the model stops

The system must start in equilibrium. Both theorems assume the initial state is the equilibrium distribution for the starting value of the control parameter. A process begun from a non-equilibrium state satisfies neither in the form given, and generalising them to such cases is an active and messier subject.

The protocol must be time-reversible in a specific sense. The reverse process is the forward one with the control parameter run backwards, and for a protocol involving a magnetic field or a rotation the reversal has to include reversing that too — the same subtlety that qualifies Onsager’s relations.

The work distributions drawn here are Gaussian, which is a near-equilibrium approximation. The theorems hold for any shape; the Gaussian is what a process dissipating a few kT with many independent contributions gives, and it makes the mean work exceed the free energy by exactly half the variance. Far from equilibrium the distributions are skewed and the relation between the mean and the variance breaks, while the theorems do not.

And the temperature is assumed uniform and constant. A process that heats the system locally, or that couples it to two baths at different temperatures, needs a more careful statement in which the entropy production is a sum over baths rather than a single dissipated work.

Where the theorem comes from

The derivation is short enough to sketch and it is worth having, because it explains why the result contains no adjustable constant.

Consider a particular trajectory of the system under the forward protocol, and the time-reverse of that trajectory under the reversed protocol. The underlying dynamics is reversible, so the two paths are equally likely given their starting points. What differs is the probability of the starting points, and those are equilibrium distributions — proportional to eE/kTe^{-E/kT} at the respective control settings.

Taking the ratio, everything about the path cancels and what is left is the ratio of two Boltzmann factors and the two partition functions. Written out, that is e(WΔF)/kTe^{(W-\Delta F)/kT}, and the theorem follows by summing over all trajectories with a given work.

Three features of that argument explain the theorem’s oddly unconditional character. It never assumed the process was slow, because the path probabilities cancelled whatever the path. It never assumed a particular system, because nothing about the dynamics survived the cancellation. And it produced an exponential with kTkT in it because the only place a temperature entered was the two equilibrium starting distributions.

The same three features explain the theorem’s limits, which are exactly the two assumptions that did not cancel: reversible microscopic dynamics, and equilibrium starting states.

What the pictures cannot show

The distributions are drawn as smooth curves and a real measurement is a histogram of a few hundred pulls. The tails, which are where all the useful information is, are exactly where a few hundred samples are worst — the crossing point is determined by the overlap of the two distributions, and if they barely overlap it is determined by a handful of events. Every reported measurement of this kind lives or dies on that overlap, and designing the experiment is largely a matter of dissipating little enough that the two distributions meet.

Nor does anything here show a trajectory. The whole subject is about the statistics of individual paths, and the figures show only the distribution of one number extracted from each. What a low-work trajectory actually does differently — where in the pull it happened to be helped by the bath — is invisible in a histogram and is what a trajectory-level analysis is for.

What counts as small

The size at which all of this becomes visible is worth a paragraph, because it is the number that decides whether the subject is a curiosity.

The scale is set by comparing the entropy produced with Boltzmann’s constant. A process producing more than about twenty kk has reversals at the level of one in 10910^{9}, which no experiment sees; one producing a few kk has them at the per cent level. So the question is which processes dissipate only a few kk, and the answer is: processes involving a few molecules, over distances of nanometres, at ordinary temperatures.

That is exactly the scale of the machinery inside a cell. A motor protein taking a step hydrolyses one molecule of ATP, worth about twenty kTkT, and dissipates a fraction of it; a ribosome adding an amino acid, an ion channel opening, a molecule folding — all of them operate within an order of magnitude of the thermal scale, and all of them therefore run backwards sometimes.

That is not a defect of biological machines but a condition of their existence. A motor that could not be pushed backwards by thermal motion would be one whose forward step cost far more than kTkT, which would make it wasteful; the ones that exist are close enough to reversible to be efficient and therefore close enough to reversible to slip. Measurements on single motor proteins show exactly that — occasional backward steps, at rates the fluctuation theorems predict from the dissipation.

Where the ladder goes next

The entropy ladder began with entropy being a count, went through the entropy that lives on a surface, the exponential that decides everything, what a system actually minimises, mixing what is already mixed and the bit that has to be paid for. This rung asks how often the law is broken and by how much. The rungs after it: the steady-state fluctuation theorem, which applies to a system held out of equilibrium indefinitely rather than driven between two equilibria; thermodynamic uncertainty relations, which bound the precision of any process by its entropy production; and stochastic thermodynamics generally, in which heat, work and entropy are defined along a single trajectory rather than over an ensemble.

The habit worth carrying away is to ask what kind of statement a law is. “Entropy increases” is a statement about a distribution, and knowing the distribution is worth more than knowing its mean — because the mean says what usually happens and the distribution says what a small system does, how often, and what it can be made to reveal.

Part 7 of 7

This essay is one argument about Entropy. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Arrow of timeDetailed balanceEntropyFluctuationsFree energyIrreversibilityReversibilityThe second lawStatistical mechanicsWork