Thermodynamics

The share that is not half a kT

Equipartition is quoted as half a kT for every degree of freedom, and it is nothing of the kind. It is half a kT for every *quadratic* term. A coordinate whose energy is linear in it carries a whole kT, and a gas hot enough that its particles' energy is pc rather than p²/2m therefore holds twice what the counting says — which drops its ratio of specific heats to four thirds and puts a star on the edge of being able to hold itself up.

Assumes: Half a kT for every way of moving · The speeds in a still room

The rule is taught as a count. Work out how many ways a molecule can move, give each half a kT, and the heat capacity follows: three halves for a monatomic gas, five halves for a diatomic one at room temperature, three for a solid.

The rule is not about ways of moving. It is about the shape of the energy in each coordinate, and the half comes from one particular shape.

What a coordinate is worth, by the shape of its energy. The mean energy stored in one coordinate, in units of kT, against the power with which that coordinate enters the energy. The curve is 1/n and the points are the same average obtained by integrating the Boltzmann weight numerically, agreeing to 3.3e-7 per cent at worst. an energy going as |x|^1 holds 1.0000 kT; an energy going as |x|^1.5 holds 0.6667 kT; an energy going as |x|^2 holds 0.5000 kT; an energy going as |x|^3 holds 0.3333 kT; an energy going as |x|^4 holds 0.2500 kT; an energy going as |x|^6 holds 0.1667 kT. The familiar half a kT is the case n = 2 and nothing more general than that: a coordinate whose energy is linear in it — the momentum of an ultrarelativistic particle, or a field with no restoring force but a constant tension — carries a whole kT, and one confined by a very steep wall carries almost nothing. Equipartition is a theorem about quadratic terms, and calling it a theorem about degrees of freedom is the substitution that makes it fail.
Fig. 1 The mean energy held in one coordinate against the power with which that coordinate enters the energy. The curve is kT/n; the points are the same average obtained by integrating the Boltzmann weight numerically, and the two agree to a part in ten million. A quadratic term holds half a kT; a linear one holds a whole kT; a quartic one holds a quarter.

Where the half comes from

Take a coordinate xx whose contribution to the energy is axna|x|^n, and average that energy over the Boltzmann weight eβaxne^{-\beta a |x|^n}. Substituting u=βaxnu = \beta a |x|^n turns both the numerator and the denominator into gamma functions, and their ratio is kT/nkT/n.

axn=kTn.\langle a|x|^n \rangle = \frac{kT}{n}.

Nothing about the coordinate survives into that answer — not aa, not whether xx is a position or a momentum, not what the coordinate is for. Only the exponent.

Put n=2n = 2 and the familiar half appears. A momentum contributes p2/2mp^2/2m; a spring contributes 12kx2\tfrac12 kx^2; a rotation contributes L2/2IL^2/2I. Those are quadratic, so each is worth half a kT, and a count of them is the same as a count of quadratic terms — which is why the shorthand works for every system a nineteenth-century laboratory contained.

Counting terms against measuring them. Measured heat capacities at constant volume at room temperature, in units of R, against the equipartition prediction of half a unit for every quadratic term in the energy. Monatomic gases have three translations and nothing else, and the prediction is exact. Diatomic gases have two rotations as well and the prediction is right if — and only if — the vibration is left out of the count, which nothing in classical physics licenses. The largest disagreement in the table is 0.90R, and it belongs to the molecules with the softest vibrations, which are exactly the ones whose vibrational steps are small enough for room temperature to reach.
Fig. 2 The count set against the measurement, at room temperature, in units of R. A monatomic gas has three translations and nothing else and the prediction is exact; a diatomic gas has three translations and two rotations and the prediction is exact; the vibration each diatomic molecule certainly has contributes nothing at all. Counting quadratic terms works until it is asked to count a term the system is not using.

The distribution of speeds in a still room is what that average is taken over, and its shape comes from three quadratic momentum components and nothing else. The mean kinetic energy it gives is three halves of a kTkT — the count and the shape agreeing, because the shape is the count: a Gaussian in each of three momenta is exactly what three quadratic terms in a Boltzmann weight produce.

The exponent as a measurement

The theorem in its general form is Clausius’s, and it has a version that is easier to use than the one above: for any coordinate xix_i,

xiHxi=kT.\left\langle x_i \frac{\partial H}{\partial x_i} \right\rangle = kT.

That is where the name virial comes from, and it holds whatever the Hamiltonian is. For HaxnH \supset a|x|^n the left-hand side is nn times the energy in that term, which returns the result above; for a system of particles held together by forces, summing it over every coordinate gives the virial theorem.

Every minimum is a parabola near enough to the bottom, which is why quadratic terms are so common and why the half a kTkT is so nearly universal. It is also where the universality stops. A Morse potential and a Lennard-Jones potential each depart from their fitted parabola at an amplitude that can be stated, and above that amplitude the exponent equipartition sees is no longer 2 — the mode is anharmonic, its share is no longer half a kTkT, and the departure grows with temperature.

There is a neat way to see why the exponent alone survives. The Boltzmann weight of a term axna|x|^n has a natural width: the coordinate ranges over whatever satisfies axnkTa|x|^n \lesssim kT, so its typical excursion is (kT/a)1/n(kT/a)^{1/n}. Raise the temperature and that excursion grows, and the energy in the term grows in proportion to kTkT however steep the potential is — but the number of ways the coordinate can arrange itself grows differently for different exponents, and it is that count, not the energy scale, which fixes the share. A steep well gives the coordinate very little room to be anywhere but the bottom, so it carries little; a shallow one gives it a great deal, so it carries more.

So a measured heat capacity is a measurement of an exponent. A crystal whose atoms sit in strongly anharmonic wells has a lattice heat capacity above the Dulong–Petit value of 3k3k per atom rather than below it, because the potential is softer than quadratic and a softer potential holds more. That is the classical anharmonic correction, and it has no quantum mechanics in it at all.

The case where the exponent is one

The interesting failure is not a slightly anharmonic spring. It is a particle whose energy is pcpc.

For a particle moving at nearly the speed of light, E=p2c2+m2c4pcE = \sqrt{p^2c^2 + m^2c^4} \to pc, which is linear in each component of the momentum rather than quadratic. Each component therefore carries a whole kT, and the particle carries three.

A gas crossing from three halves to three. The mean kinetic energy per particle, in units of kT, against temperature in units of the particle's own rest energy, integrated over the relativistic Maxwell distribution. It is 1.502 kT at the cold end — the three halves of a non-relativistic gas, which is three quadratic momentum components at half a kT each — and 2.999 kT at the hot end, where the energy is pc and each component is linear rather than quadratic, so each carries a whole kT. The second curve is 1 + P/u, the ratio of specific heats the energy budget of a star uses: it falls from 5/3 to 4/3 across the same crossing. Nothing in between is a mixture of two gases; it is one gas whose particles are neither slow nor fast, and the drift between the two plateaus takes about four decades of temperature.
Fig. 3 A gas taken across the crossing, with everything integrated over the relativistic Maxwell distribution rather than interpolated. The mean kinetic energy per particle goes from 1.5 kT to 3.0 kT, and the ratio 1 + P/u — the quantity a star’s energy budget uses — falls from five thirds to four thirds. Nothing about the number of coordinates has changed. The shape of the energy has.

The crossing is not sharp. It takes about four decades of temperature, centred on where kTkT equals the particle’s rest energy: for electrons that is 6×1096 \times 10^9 kelvin, for protons 101310^{13}. A gas in the middle is not a mixture of two gases; it is one gas of particles that are neither slow nor fast, and its heat capacity has no simple count behind it.

Underneath every one of these averages is the same counting. Enormous numbers of arrangements produce a quantity that can be printed on a dial, and the temperature is a parameter of the distribution rather than a property of any member of it. Equipartition is a statement about that distribution’s shape, and it says nothing whatever about any one particle: at any instant most of them hold something other than their share.

The linear term that has been in the room all along

The relativistic case is exotic, and there is an entirely everyday one that has the same exponent.

A molecule at height hh in a uniform gravitational field has a potential energy mghmgh — linear in the coordinate, not quadratic. So a column of gas tall enough for gravity to matter has, per molecule, three halves of a kT of kinetic energy and a whole kT of potential energy, and its heat capacity at constant volume is 52Nk\tfrac52 Nk rather than 32Nk\tfrac32 Nk.

The air thins with height by exactly this mechanism, and it is the cleanest linear coordinate in the world. Density falls as emgh/kTe^{-mgh/kT}, which is the Boltzmann weight of an energy linear in the coordinate, so the scale height is precisely the height at which a molecule’s potential energy is one kTkT — and the mean potential energy of the whole column is one kTkT per molecule, regardless of how tall the column is.

The mean height of a molecule in an isothermal atmosphere is one scale height, whatever the scale height is: eight and a half kilometres on Earth, eleven on Mars, twenty-seven on Titan. That the answer contains no property of the atmosphere is the same universality the half a kT has, arriving through the same integral with n=1n = 1 instead of n=2n = 2.

The correction is small in practice for a laboratory sample and enormous for a planet, and the arithmetic behind an atmosphere’s energy budget uses it constantly. It is also the cleanest available demonstration that the theorem is not counting motions: nothing is moving in that extra kT. It is a position in a field, and it earns twice what a position in a spring earns.

The exponent and the adiabatic index are one number

The four thirds that the next section is about is not a separate fact from the exponent this essay began with. It is the same number, converted.

For a dilute gas in three dimensions whose particles have energy going as the nn-th power of their momentum, the pressure and the energy density are related by

P=n3u,P = \frac{n}{3}\,u,

which is a statement about how much momentum each particle delivers to a wall relative to how much energy it carries. And the ratio of specific heats for such a gas is γ=1+P/u\gamma = 1 + P/u, so

γ=1+n3.\gamma = 1 + \frac{n}{3}.

Put n=2n = 2, the ordinary quadratic case, and γ\gamma is five thirds — the value every monatomic gas has. Put n=1n = 1, the ultrarelativistic case, and it is four thirds. The two numbers that stellar structure treats as separate regimes are the two values of one exponent, and the crossing between them is the crossing this essay’s figure draws.

Which makes the chain complete and rather short. The shape of a particle’s energy in its own momentum fixes how much of a kT that momentum carries; that fixes the ratio of pressure to energy density; that fixes the adiabatic index; and the adiabatic index fixes whether a self-gravitating ball has energy to spend on resisting a squeeze. One exponent, four steps, and a star.

It also explains why photons give exactly four thirds without any argument about temperature. A photon has E=pcE = pc identically, at every energy, with no crossing to be made and no rest mass to compare against — so radiation pressure is one third of the radiation energy density exactly, and a star supported by it sits on the boundary rather than approaching it.

What four thirds does to a star

A self-gravitating ball of gas is held up by its own pressure, and whether it can be is decided by one number.

Three energies, one of which is a mirror of another. The gravitational, kinetic and total energies of a uniform self-gravitating sphere of 1 solar mass, against its radius, in units of the total energy it has at one solar radius. The kinetic energy is exactly minus half the gravitational one at every radius, which is the virial theorem, and the total is exactly minus the kinetic. So the curve that says how much energy the body has and the curve that says how hot it is are the same curve upside down: the mean temperature at one solar radius is 2.78 million K, and shrinking the body raises it.
Fig. 4 The virial theorem applied to one self-gravitating body: twice the kinetic energy plus the gravitational energy is zero, so the total energy is minus the kinetic energy and taking energy away makes the body hotter. That is the negative heat capacity a star lives by, and it holds for a non-relativistic gas.

The derivation of that result uses Etotal=U+KE_{\text{total}} = U + K with K=32NkTK = \tfrac32 NkT — the non-relativistic value. Redo it with a general ratio: for a gas obeying P=(γ1)uP = (\gamma - 1)u, the virial theorem gives

Etotal=3γ43(γ1)Ugrav,E_{\text{total}} = \frac{3\gamma - 4}{3(\gamma - 1)}\,U_{\text{grav}},

which is negative for γ>4/3\gamma > 4/3 and exactly zero at γ=4/3\gamma = 4/3.

The line that slopes the wrong way. Mean temperature against total energy, for a self-gravitating sphere at nine radii. The points fall on a straight line of negative slope: taking energy away makes the body hotter. The heat capacity read off the drawing is -4.10e+34 J/K against the -4.10e+34 J/K the theorem gives, and the sign is the whole content. The dashed line is an ordinary gas in a rigid box, whose temperature rises when energy is added, as everything one can put a thermometer in does.
Fig. 5 The heat capacity that follows, and why the sign of the total energy is the whole story. A body with negative total energy is bound: squeezing it releases energy, some of which becomes heat and pushes back. A body with zero total energy releases nothing when squeezed, so there is nothing for it to resist with, and it is on the boundary between stable and not.

That is why the number four thirds appears everywhere in stellar structure. A star supported by radiation pressure has γ=4/3\gamma = 4/3 exactly, since photons are the ultrarelativistic gas par excellence; a star hot enough that its electrons are relativistic has it approximately; and a white dwarf near the Chandrasekhar mass has it because its degenerate electrons are moving at nearly the speed of light.

There is another way equipartition fails and it is worth naming beside this one, because the causes are unrelated. When the exclusion principle stops most of the particles changing state at all, only the fraction within kTkT of the top of the distribution can absorb anything, and the heat capacity falls a hundredfold below what any count of coordinates predicts. That failure is quantum. The one this essay is about is not: it is a statement about the exponent in a classical energy.

Half a kT in a capacitor

Nothing in the derivation is mechanical, and the clearest demonstration of that is a coordinate with no mass, no spring and no motion in it: the charge on a capacitor.

The energy stored on a capacitor of capacitance CC carrying charge qq is q2/2Cq^2/2C — quadratic in the charge, so the charge is a coordinate of exactly the kind the theorem is about. Connect the capacitor to anything at temperature TT and equipartition applies at once:

q22C=12kBTVrms=kBTC.\left\langle \frac{q^2}{2C} \right\rangle = \tfrac12 k_BT \quad\Longrightarrow\quad V_{\text{rms}} = \sqrt{\frac{k_BT}{C}}.

The voltage across a capacitor at room temperature fluctuates, by an amount that depends on nothing but the capacitance. A picofarad gives 64 microvolts; a femtofarad gives 640; and no amount of care with the circuit reduces it, because the fluctuation is not noise picked up from anywhere but the thermal share of a quadratic coordinate.

Three things about that make it the sharpest illustration in this essay. The resistance does not appear — which is initially startling, since a resistor is what supplies the noise, and the resolution is that a larger resistance gives more noise per unit bandwidth over a narrower bandwidth, and the two cancel exactly. Nothing is moving anywhere, so the description of equipartition as sharing energy among “motions” fails completely. And the answer is a hard floor on a measurement: sampling a voltage onto a capacitor cannot be done more precisely than this, which is why the smallest capacitor in an image sensor’s pixel is chosen against a noise requirement rather than against a size one.

The floor under every instrument

The same argument, applied to a mechanical coordinate rather than an electrical one, gives the number every precision instrument is designed against.

Any sensor with a restoring force is a quadratic coordinate: a cantilever, a torsion fibre, a suspended mirror, a pressure diaphragm. Equipartition gives it 12kBT\tfrac12k_BT, so its mean square displacement is

x2=kBTk,\langle x^2 \rangle = \frac{k_BT}{k},

the thermal energy over the stiffness — and again with nothing else in it. A cantilever of one newton per metre wanders by 64 picometres at room temperature, which is comparable with the size of an atom and is the reason a scanning probe’s resolution is what it is.

The dependence is the useful part. Reducing the wander means raising the stiffness or lowering the temperature, and nothing else is available: not a better readout, not a longer average of the position itself, not a quieter room. And raising the stiffness costs sensitivity in exact proportion, since a stiffer sensor deflects less under the force being measured. That trade is why the smallest measurable force with such an instrument is a fixed quantity rather than something a better design improves, and why the improvements that have been made are almost all in cooling.

What it costs

Everything above is classical. The generalised theorem is a statement about a Boltzmann average over a continuous phase space, and it is exactly as good as that description. Where the spacing of a system’s levels exceeds kTkT the average is over a sum rather than an integral and the answer is smaller — which is the freeze-out that gives hydrogen its three plateaus.

The staircase equipartition cannot climb. The heat capacity of hydrogen at constant volume, in units of R, against temperature on a logarithmic axis. The three translational directions contribute 3/2 at every temperature. The two rotations switch on near 85.4 K — the temperature at which kT matches the first rotational step — taking the total to 5/2, computed here by summing the rigid rotor's partition function over two hundred levels rather than by assuming the plateau. The vibration switches on near 6332 K, taking it to 7/2. At 50 K the value is 2.469; at 300 K the value is 2.502; at 5000 K the value is 3.376. Classical equipartition predicts 7/2 at every temperature, including at four kelvin, and the size of the steps is set by ħ — which is how a heat capacity measures a quantum constant.
Fig. 6 The staircase equipartition cannot explain. Between the plateaus the count of quadratic terms is not an integer and is not anything: modes are partly available, and the value is set by the ratio of a level spacing to a temperature rather than by a count of anything.

The energy has been assumed separable. Writing the energy as a sum of terms, each depending on one coordinate, is what lets each term be averaged on its own. A gas of interacting particles does not separate, and its heat capacity contains a contribution from the interaction that no counting produces.

And a gas has been assumed ideal. The relation P=nkTP = nkT used above holds for a dilute gas of any speed, which is a stronger statement than it looks — it is true for photons in the sense that P=u/3P = u/3, but the pressure of a photon gas is not nkTnkT because the number of photons is not conserved.

Where the model stops

The relativistic crossing is drawn for a gas of fixed particle number. A real gas hot enough for its electrons to be relativistic is hot enough to make electron–positron pairs, and once pair production starts the particle count is a function of temperature. The heat capacity then acquires a term far larger than anything here, because energy goes into making mass rather than into motion — which is the instability that ends the life of a very massive star.

Nothing here says how fast anything happens. Equipartition describes an equilibrium and is silent about whether it is reached. A mode that exchanges energy with the rest of the system very slowly is out of equilibrium at any practical timescale and holds whatever it was given, which is why a shock-heated gas has a different electron temperature and ion temperature for a long time.

One curve for every solid, once the temperature is measured in its own units. The molar heat capacity of a solid in Debye's model, in units of the gas constant, against temperature divided by that solid's own Debye temperature. All four fall on one curve, which is the model's whole claim: a solid has one parameter and no others. It climbs to 3.000R, the Dulong and Petit value that every solid reaches when every mode is excited, and it falls at low temperature as the cube of the temperature with a coefficient of 233.782, both computed from the integral rather than quoted. The four solids reach half of Dulong and Petit at 26 K for lead, 85 K for copper, 160 K for silicon, 555 K for diamond — a spread of a factor of twenty, from one number each. The cube is the part the third law needs. Entropy is the integral of C/T from absolute zero, and an integrand going as T² converges there; a heat capacity that stayed at 3R all the way down would make that integral diverge logarithmically and there would be no absolute entropy to speak of.
Fig. 7 And where the counting is replaced rather than corrected. Four solids’ molar heat capacities fall on one curve once each temperature is measured in that solid’s own units — the whole claim of Debye’s model, which is that a solid has one parameter and no others. At the top the curve reaches 3R, which is equipartition’s answer for three quadratic terms twice over; at the bottom it goes as T³, which equipartition has no way of producing.

Two coupled oscillators are the simplest system that has equipartition arriving, and watching it arrive is the point: energy passes back and forth between them and their time-averages equalise, but the equalising takes as long as the coupling is weak. The theorem asserts the average and says nothing about the exchange, and in a system with a weak enough coupling the average is not reached within any time anybody has watched.

And the four thirds is a boundary rather than a verdict. A body at exactly γ=4/3\gamma = 4/3 is neutrally stable in this analysis, and what decides its fate is whatever has been neglected: general relativity, which makes things worse; electrostatic corrections, which help; rotation, which helps. A real star near the boundary is decided by the corrections, and the boundary only says which corrections matter.

The weight every average here is taken against is an exponential in the energy over kTkT, and that is the whole reason the exponent nn comes out of the average so cleanly. A power law inside an exponential integrates to a gamma function, and the ratio of two gamma functions differing by one in their argument is that argument. It is arithmetic rather than physics, which is why the result holds for coordinates that have nothing physically in common.

What the pictures cannot show

The hero figure plots a mean energy against a continuous exponent, and no real system has a continuously adjustable one. The points at n=1n = 1 and n=2n = 2 are physical; the curve between them is an interpolation through cases nobody can build, drawn because it makes the shape of the dependence legible and not because a system at n=1.5n = 1.5 is available.

Nor can any of these figures show what is being averaged. Every quantity here is an expectation over a distribution across 102310^{23} particles, and the object that the theorem is about — a single coordinate’s share, fluctuating wildly from instant to instant and from particle to particle — appears nowhere. What is drawn is the mean of something whose typical departure from the mean is of the same size as the mean itself.

The result is a century older than its use

The generalised form is Clausius’s, from 1870, and the phrase he used for the quantity xiH/xi\sum x_i \,\partial H/\partial x_ithe virial, from the Latin for force — has attached itself to two rather different results that share the same algebra.

Applied to one coordinate in contact with a heat bath it gives the theorem above and is a statement about temperature. Applied to a whole bound system with no bath at all, and time-averaged rather than ensemble-averaged, it gives the relation between kinetic and potential energy that the star figures use. The same expression, two averages, two subjects: one belongs to thermodynamics and the other to mechanics, and a great deal of confusion has come from the shared name.

What is common to both is the reason the answer is so free of detail. The quantity being averaged is homogeneous of some degree in its coordinate, and the average of a homogeneous function against an exponential of itself depends only on the degree. Everything else integrates away. That is why kT/nkT/n contains no spring constant, no mass and no volume — and why the same arithmetic serves a molecule in a bottle and a galaxy in a cluster.

Where this ladder goes next

Two rungs stand on equipartition. The first found the theorem’s quantum failure: modes with a level spacing above kT are not available, and hydrogen’s heat capacity climbs a staircase because of it. This one finds a failure with no quantum mechanics in it — the theorem is about quadratic terms, and a term that is not quadratic pays a different share.

The habit worth carrying away concerns shorthands. A rule stated as a count is usually a rule about a shape, restated for the case where every item has the same shape. Equipartition counts quadratic terms and was taught as counting motions because in the systems it was invented for the two agreed. The same substitution is behind half the misapplications in thermodynamics, and the way to detect it is to look for a case where the shapes differ.

What is left on this ladder is the theorem’s other half. The virial relation used here to get kTkT out of one coordinate, summed over every coordinate of a bound system, gives a relation between average kinetic and potential energies that holds for a molecule, a star cluster and a galaxy alike — and turning it into a measurement of a mass that cannot be weighed is what it has mostly been used for.

Part 2 of 7

This essay is one argument about Equipartition. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Adiabatic indexThe Boltzmann factorDegrees of freedomEquipartitionGravitational collapseHeat capacityInternal energyKinetic theoryPhase spaceRelativistic gasTemperatureVirial theorem