The second law, with a probability attached
Assumes: Entropy is a count, and the arrow of time is arithmetic · The bit that has to be paid for
The second law is usually stated as a prohibition: entropy does not decrease. Stated that way it is a statement about macroscopic systems, and for those it is as close to certain as anything in physics — the count of arrangements is so lopsided that the probability of a visible decrease is smaller than any number worth writing.
Take a system small enough for its own thermal motion to matter and the statement changes character. Individual runs do go the wrong way, sometimes, and the question of how often has an exact answer.
The shaded region is not an error bar or a measurement uncertainty. Those runs really happened, really did less work than the reversible minimum, and really left the universe with less entropy than it started with. The law is not violated on average and it is violated in individual cases, and the two statements are consistent because the law was always about the average.
The relation that fixes how often
Crooks’ theorem, from 1999, says that the probability of doing work going forwards and the probability of doing going backwards are in the ratio . Nothing else appears — not the speed of the process, not the protocol, not the size of the system, not what it is made of.
Read at , the ratio is one: the two distributions cross exactly at the free-energy change. Read in the tail, it gives the odds of a run in which entropy decreased. An entropy production of is as likely as the corresponding increase, which is one in twenty-two thousand; is one in ; and by the time the system is macroscopic the numbers are of the size that makes the second law feel absolute.
So the transition from “sometimes” to “never” is not a change in the physics. It is an exponential in the size of the entropy change, and the entropy change scales with the system, and the exponential does the rest.
The theorem’s most useful feature is what it does not require. There is no assumption that the process is slow, or reversible, or close to equilibrium; the system need only start in equilibrium and be driven by a protocol that can be run backwards. That is what makes it applicable to processes that no equilibrium argument covers.
Reading a free energy out of an irreversible measurement
The crossing point is the practical part. A free-energy change is an equilibrium quantity, and the classical way to measure one is to do the process slowly enough that the work equals it — which for many systems is impossible, because slowly enough means slower than anything drifts.
Crooks’ theorem removes the requirement. Pull the system apart quickly, many times, and build the work distribution; push it back together quickly, many times, and build the reverse distribution; and where the two cross is the free-energy change. Neither measurement was reversible and the answer is the reversible one.
That is now standard in single-molecule work. A stretch of RNA is unfolded and refolded with an optical trap at speeds that dissipate several kT a time, the two distributions are accumulated over a few hundred pulls, and the crossing gives the folding free energy of a structure that could not be measured any other way. The first such measurement, on an RNA hairpin in 2005, agreed with the equilibrium value where an equilibrium value was available and gave one where it was not.
The technique’s requirement is that both directions be measurable, which is a real constraint: many processes cannot be run backwards at all, and for those the crossing method is unavailable.
The exact equality that does not help as much as it should
Jarzynski’s equality, from 1997, is the more famous statement and it follows from Crooks’ by an integration: the average of over the forward process is exactly . It needs only the forward direction, which removes the constraint above.
It also has a difficulty that the figure exists to show. The exponential weights small works most heavily, so the average is dominated by the rare trajectories in the low-work tail — precisely the ones in which entropy decreased. Until enough of those have been sampled, the estimate sits near the mean work rather than near the free energy, and the number of runs needed to sample them grows exponentially with the dissipation.
The figure’s run dissipates three kT and takes thousands of samples to come down. At ten kT it would take millions; at fifty, more than any experiment. So the equality is exact and the estimator is bad, and the two facts are entirely compatible.
That combination is worth recognising because it recurs. An identity that is exact for all sample sizes says nothing about how quickly an estimate built from it converges, and an average dominated by a rare tail converges slowly however exact the identity behind it. The remedies are the ordinary ones — dissipate less, or sample the tail deliberately — and neither is free.
Comparing the two dissipations shows the trade the technique lives on, and it runs in an awkward direction.
A gentle process gives distributions that overlap, so the crossing point is well determined and the free energy comes out of a few dozen runs. It also gives entropy decreases that are common — a third of the runs, here — which makes the “violation” unremarkable and the demonstration weak.
A violent process gives entropy decreases that are rare and impressive, and distributions so far apart that no practical number of runs locates the crossing. The overlap region, which is where every bit of information about the free energy lives, is exponentially unlikely to be sampled at all.
So the experiment that shows the second law being broken most convincingly is the one that measures the free energy worst, and vice versa. Anyone designing such a measurement is choosing a point on that curve, and the sensible choice is a dissipation of a few kT — enough that the process is genuinely irreversible, little enough that the distributions still meet.
What this does and does not overturn
It is worth being careful about the claim, because the fluctuation theorems are sometimes reported as a refutation of the second law and are nothing of the kind.
What is new is not that small systems fluctuate; that has been known since Brownian motion. What is new is the exact relation between the probability of a fluctuation and the entropy it corresponds to, holding arbitrarily far from equilibrium, with no adjustable constant. Before these results, non-equilibrium statements were mostly inequalities and mostly restricted to near equilibrium. These are equalities and they are not restricted.
What is unchanged is everything about the macroscopic second law. The relations say the probability of a violation falls exponentially with its size in units of , and a macroscopic entropy change is of order , so nothing about a steam engine has moved. Maxwell’s demon does not become possible, because the demon’s cost is an accounting question rather than a fluctuation question.
And what is clarified is the status of the law itself. It is a statement about a probability distribution rather than a prohibition, its strength depends on the size of the system, and the size dependence is now a formula rather than a hand-wave. That is a change in what kind of statement it is, which is worth more than a change in what it says.
Where the entropy actually is
One point deserves care because it is where the statements are most often garbled. The quantity that appears in these theorems is the total entropy production — the system’s plus the bath’s — and for a system driven between equilibrium states it is the dissipated work over the temperature.
That means the “violations” are not runs in which a system spontaneously ordered itself while everything else stayed put. They are runs in which the system took less heat from the bath than the average, so the bath’s entropy increased by less than the system’s decreased. The total went down, and it went down because the bath happened to deliver an unusually convenient sequence of kicks.
The distinction matters for what such an experiment is evidence about. It is not evidence that a system can organise itself; it is evidence that the exchange with a bath is a fluctuating quantity whose fluctuations are quantitatively constrained. That is a statement about the bath as much as about the system, and it is the same statement as the one about a small system’s energy wandering with the accounting done over both.
The measurement of a bit, and what it costs
The Szilard engine belongs beside the fluctuation theorems because it is where the two halves of the subject meet, and because it makes the numbers concrete.
One molecule in a box carries a fluctuation of order its own energy — it is on one side or the other, and which is a fair coin. Knowing which is one bit, and that bit is worth kT ln 2 of work and costs kT ln 2 to erase. The same appears in Crooks’ relation as the work difference corresponding to a factor of two in probability, and that is not two coincidences: entropy is a logarithm of a probability in both statements.
What the fluctuation theorems add is the general version. A trajectory whose entropy production is is times more likely than its reverse; a bit of information is of entropy; so acquiring one bit shifts the odds of a process against its reverse by exactly a factor of two. Information and irreversibility are the same quantity in different units, and the exchange rate is fixed.
That has been measured. Experiments in the 2010s built Szilard engines from a single colloidal particle in a light trap, extracted work from measurements of its position, and confirmed both the per bit and the modified fluctuation theorem that includes an information term. The engine Szilard invented in 1929 as a thought experiment now runs on a microscope stage.
Where the model stops
The system must start in equilibrium. Both theorems assume the initial state is the equilibrium distribution for the starting value of the control parameter. A process begun from a non-equilibrium state satisfies neither in the form given, and generalising them to such cases is an active and messier subject.
The protocol must be time-reversible in a specific sense. The reverse process is the forward one with the control parameter run backwards, and for a protocol involving a magnetic field or a rotation the reversal has to include reversing that too — the same subtlety that qualifies Onsager’s relations.
The work distributions drawn here are Gaussian, which is a near-equilibrium approximation. The theorems hold for any shape; the Gaussian is what a process dissipating a few kT with many independent contributions gives, and it makes the mean work exceed the free energy by exactly half the variance. Far from equilibrium the distributions are skewed and the relation between the mean and the variance breaks, while the theorems do not.
And the temperature is assumed uniform and constant. A process that heats the system locally, or that couples it to two baths at different temperatures, needs a more careful statement in which the entropy production is a sum over baths rather than a single dissipated work.
Where the theorem comes from
The derivation is short enough to sketch and it is worth having, because it explains why the result contains no adjustable constant.
Consider a particular trajectory of the system under the forward protocol, and the time-reverse of that trajectory under the reversed protocol. The underlying dynamics is reversible, so the two paths are equally likely given their starting points. What differs is the probability of the starting points, and those are equilibrium distributions — proportional to at the respective control settings.
Taking the ratio, everything about the path cancels and what is left is the ratio of two Boltzmann factors and the two partition functions. Written out, that is , and the theorem follows by summing over all trajectories with a given work.
Three features of that argument explain the theorem’s oddly unconditional character. It never assumed the process was slow, because the path probabilities cancelled whatever the path. It never assumed a particular system, because nothing about the dynamics survived the cancellation. And it produced an exponential with in it because the only place a temperature entered was the two equilibrium starting distributions.
The same three features explain the theorem’s limits, which are exactly the two assumptions that did not cancel: reversible microscopic dynamics, and equilibrium starting states.
What the pictures cannot show
The distributions are drawn as smooth curves and a real measurement is a histogram of a few hundred pulls. The tails, which are where all the useful information is, are exactly where a few hundred samples are worst — the crossing point is determined by the overlap of the two distributions, and if they barely overlap it is determined by a handful of events. Every reported measurement of this kind lives or dies on that overlap, and designing the experiment is largely a matter of dissipating little enough that the two distributions meet.
Nor does anything here show a trajectory. The whole subject is about the statistics of individual paths, and the figures show only the distribution of one number extracted from each. What a low-work trajectory actually does differently — where in the pull it happened to be helped by the bath — is invisible in a histogram and is what a trajectory-level analysis is for.
What counts as small
The size at which all of this becomes visible is worth a paragraph, because it is the number that decides whether the subject is a curiosity.
The scale is set by comparing the entropy produced with Boltzmann’s constant. A process producing more than about twenty has reversals at the level of one in , which no experiment sees; one producing a few has them at the per cent level. So the question is which processes dissipate only a few , and the answer is: processes involving a few molecules, over distances of nanometres, at ordinary temperatures.
That is exactly the scale of the machinery inside a cell. A motor protein taking a step hydrolyses one molecule of ATP, worth about twenty , and dissipates a fraction of it; a ribosome adding an amino acid, an ion channel opening, a molecule folding — all of them operate within an order of magnitude of the thermal scale, and all of them therefore run backwards sometimes.
That is not a defect of biological machines but a condition of their existence. A motor that could not be pushed backwards by thermal motion would be one whose forward step cost far more than , which would make it wasteful; the ones that exist are close enough to reversible to be efficient and therefore close enough to reversible to slip. Measurements on single motor proteins show exactly that — occasional backward steps, at rates the fluctuation theorems predict from the dissipation.
Where the ladder goes next
The entropy ladder began with entropy being a count, went through the entropy that lives on a surface, the exponential that decides everything, what a system actually minimises, mixing what is already mixed and the bit that has to be paid for. This rung asks how often the law is broken and by how much. The rungs after it: the steady-state fluctuation theorem, which applies to a system held out of equilibrium indefinitely rather than driven between two equilibria; thermodynamic uncertainty relations, which bound the precision of any process by its entropy production; and stochastic thermodynamics generally, in which heat, work and entropy are defined along a single trajectory rather than over an ensemble.
The habit worth carrying away is to ask what kind of statement a law is. “Entropy increases” is a statement about a distribution, and knowing the distribution is worth more than knowing its mean — because the mean says what usually happens and the distribution says what a small system does, how often, and what it can be made to reveal.
Part 7 of 7
This essay is one argument about Entropy. The others:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
Arrow of timeDetailed balanceEntropyFluctuationsFree energyIrreversibilityReversibilityThe second lawStatistical mechanicsWork
- The engine that pays back more than it takes entropy, free energy, irreversibility, reversibility, the second law
- A fridge with no work going into it entropy, reversibility, the second law
- The area that is not allowed to shrink entropy, irreversibility, the second law
- The ceiling on every engine, set before it was designed reversibility, the second law, work
- The engine that has to finish entropy, irreversibility, the second law
- The glow that says nothing about the surface detailed balance, reversibility, the second law