The share that is not half a kT
Assumes: Half a kT for every way of moving · The speeds in a still room
The rule is taught as a count. Work out how many ways a molecule can move, give each half a kT, and the heat capacity follows: three halves for a monatomic gas, five halves for a diatomic one at room temperature, three for a solid.
The rule is not about ways of moving. It is about the shape of the energy in each coordinate, and the half comes from one particular shape.
Where the half comes from
Take a coordinate whose contribution to the energy is , and average that energy over the Boltzmann weight . Substituting turns both the numerator and the denominator into gamma functions, and their ratio is .
Nothing about the coordinate survives into that answer — not , not whether is a position or a momentum, not what the coordinate is for. Only the exponent.
Put and the familiar half appears. A momentum contributes ; a spring contributes ; a rotation contributes . Those are quadratic, so each is worth half a kT, and a count of them is the same as a count of quadratic terms — which is why the shorthand works for every system a nineteenth-century laboratory contained.
The distribution of speeds in a still room is what that average is taken over, and its shape comes from three quadratic momentum components and nothing else. The mean kinetic energy it gives is three halves of a — the count and the shape agreeing, because the shape is the count: a Gaussian in each of three momenta is exactly what three quadratic terms in a Boltzmann weight produce.
The exponent as a measurement
The theorem in its general form is Clausius’s, and it has a version that is easier to use than the one above: for any coordinate ,
That is where the name virial comes from, and it holds whatever the Hamiltonian is. For the left-hand side is times the energy in that term, which returns the result above; for a system of particles held together by forces, summing it over every coordinate gives the virial theorem.
Every minimum is a parabola near enough to the bottom, which is why quadratic terms are so common and why the half a is so nearly universal. It is also where the universality stops. A Morse potential and a Lennard-Jones potential each depart from their fitted parabola at an amplitude that can be stated, and above that amplitude the exponent equipartition sees is no longer 2 — the mode is anharmonic, its share is no longer half a , and the departure grows with temperature.
There is a neat way to see why the exponent alone survives. The Boltzmann weight of a term has a natural width: the coordinate ranges over whatever satisfies , so its typical excursion is . Raise the temperature and that excursion grows, and the energy in the term grows in proportion to however steep the potential is — but the number of ways the coordinate can arrange itself grows differently for different exponents, and it is that count, not the energy scale, which fixes the share. A steep well gives the coordinate very little room to be anywhere but the bottom, so it carries little; a shallow one gives it a great deal, so it carries more.
So a measured heat capacity is a measurement of an exponent. A crystal whose atoms sit in strongly anharmonic wells has a lattice heat capacity above the Dulong–Petit value of per atom rather than below it, because the potential is softer than quadratic and a softer potential holds more. That is the classical anharmonic correction, and it has no quantum mechanics in it at all.
The case where the exponent is one
The interesting failure is not a slightly anharmonic spring. It is a particle whose energy is .
For a particle moving at nearly the speed of light, , which is linear in each component of the momentum rather than quadratic. Each component therefore carries a whole kT, and the particle carries three.
The crossing is not sharp. It takes about four decades of temperature, centred on where equals the particle’s rest energy: for electrons that is kelvin, for protons . A gas in the middle is not a mixture of two gases; it is one gas of particles that are neither slow nor fast, and its heat capacity has no simple count behind it.
Underneath every one of these averages is the same counting. Enormous numbers of arrangements produce a quantity that can be printed on a dial, and the temperature is a parameter of the distribution rather than a property of any member of it. Equipartition is a statement about that distribution’s shape, and it says nothing whatever about any one particle: at any instant most of them hold something other than their share.
The linear term that has been in the room all along
The relativistic case is exotic, and there is an entirely everyday one that has the same exponent.
A molecule at height in a uniform gravitational field has a potential energy — linear in the coordinate, not quadratic. So a column of gas tall enough for gravity to matter has, per molecule, three halves of a kT of kinetic energy and a whole kT of potential energy, and its heat capacity at constant volume is rather than .
The air thins with height by exactly this mechanism, and it is the cleanest linear coordinate in the world. Density falls as , which is the Boltzmann weight of an energy linear in the coordinate, so the scale height is precisely the height at which a molecule’s potential energy is one — and the mean potential energy of the whole column is one per molecule, regardless of how tall the column is.
The mean height of a molecule in an isothermal atmosphere is one scale height, whatever the scale height is: eight and a half kilometres on Earth, eleven on Mars, twenty-seven on Titan. That the answer contains no property of the atmosphere is the same universality the half a kT has, arriving through the same integral with instead of .
The correction is small in practice for a laboratory sample and enormous for a planet, and the arithmetic behind an atmosphere’s energy budget uses it constantly. It is also the cleanest available demonstration that the theorem is not counting motions: nothing is moving in that extra kT. It is a position in a field, and it earns twice what a position in a spring earns.
The exponent and the adiabatic index are one number
The four thirds that the next section is about is not a separate fact from the exponent this essay began with. It is the same number, converted.
For a dilute gas in three dimensions whose particles have energy going as the -th power of their momentum, the pressure and the energy density are related by
which is a statement about how much momentum each particle delivers to a wall relative to how much energy it carries. And the ratio of specific heats for such a gas is , so
Put , the ordinary quadratic case, and is five thirds — the value every monatomic gas has. Put , the ultrarelativistic case, and it is four thirds. The two numbers that stellar structure treats as separate regimes are the two values of one exponent, and the crossing between them is the crossing this essay’s figure draws.
Which makes the chain complete and rather short. The shape of a particle’s energy in its own momentum fixes how much of a kT that momentum carries; that fixes the ratio of pressure to energy density; that fixes the adiabatic index; and the adiabatic index fixes whether a self-gravitating ball has energy to spend on resisting a squeeze. One exponent, four steps, and a star.
It also explains why photons give exactly four thirds without any argument about temperature. A photon has identically, at every energy, with no crossing to be made and no rest mass to compare against — so radiation pressure is one third of the radiation energy density exactly, and a star supported by it sits on the boundary rather than approaching it.
What four thirds does to a star
A self-gravitating ball of gas is held up by its own pressure, and whether it can be is decided by one number.
The derivation of that result uses with — the non-relativistic value. Redo it with a general ratio: for a gas obeying , the virial theorem gives
which is negative for and exactly zero at .
That is why the number four thirds appears everywhere in stellar structure. A star supported by radiation pressure has exactly, since photons are the ultrarelativistic gas par excellence; a star hot enough that its electrons are relativistic has it approximately; and a white dwarf near the Chandrasekhar mass has it because its degenerate electrons are moving at nearly the speed of light.
There is another way equipartition fails and it is worth naming beside this one, because the causes are unrelated. When the exclusion principle stops most of the particles changing state at all, only the fraction within of the top of the distribution can absorb anything, and the heat capacity falls a hundredfold below what any count of coordinates predicts. That failure is quantum. The one this essay is about is not: it is a statement about the exponent in a classical energy.
Half a kT in a capacitor
Nothing in the derivation is mechanical, and the clearest demonstration of that is a coordinate with no mass, no spring and no motion in it: the charge on a capacitor.
The energy stored on a capacitor of capacitance carrying charge is — quadratic in the charge, so the charge is a coordinate of exactly the kind the theorem is about. Connect the capacitor to anything at temperature and equipartition applies at once:
The voltage across a capacitor at room temperature fluctuates, by an amount that depends on nothing but the capacitance. A picofarad gives 64 microvolts; a femtofarad gives 640; and no amount of care with the circuit reduces it, because the fluctuation is not noise picked up from anywhere but the thermal share of a quadratic coordinate.
Three things about that make it the sharpest illustration in this essay. The resistance does not appear — which is initially startling, since a resistor is what supplies the noise, and the resolution is that a larger resistance gives more noise per unit bandwidth over a narrower bandwidth, and the two cancel exactly. Nothing is moving anywhere, so the description of equipartition as sharing energy among “motions” fails completely. And the answer is a hard floor on a measurement: sampling a voltage onto a capacitor cannot be done more precisely than this, which is why the smallest capacitor in an image sensor’s pixel is chosen against a noise requirement rather than against a size one.
The floor under every instrument
The same argument, applied to a mechanical coordinate rather than an electrical one, gives the number every precision instrument is designed against.
Any sensor with a restoring force is a quadratic coordinate: a cantilever, a torsion fibre, a suspended mirror, a pressure diaphragm. Equipartition gives it , so its mean square displacement is
the thermal energy over the stiffness — and again with nothing else in it. A cantilever of one newton per metre wanders by 64 picometres at room temperature, which is comparable with the size of an atom and is the reason a scanning probe’s resolution is what it is.
The dependence is the useful part. Reducing the wander means raising the stiffness or lowering the temperature, and nothing else is available: not a better readout, not a longer average of the position itself, not a quieter room. And raising the stiffness costs sensitivity in exact proportion, since a stiffer sensor deflects less under the force being measured. That trade is why the smallest measurable force with such an instrument is a fixed quantity rather than something a better design improves, and why the improvements that have been made are almost all in cooling.
What it costs
Everything above is classical. The generalised theorem is a statement about a Boltzmann average over a continuous phase space, and it is exactly as good as that description. Where the spacing of a system’s levels exceeds the average is over a sum rather than an integral and the answer is smaller — which is the freeze-out that gives hydrogen its three plateaus.
The energy has been assumed separable. Writing the energy as a sum of terms, each depending on one coordinate, is what lets each term be averaged on its own. A gas of interacting particles does not separate, and its heat capacity contains a contribution from the interaction that no counting produces.
And a gas has been assumed ideal. The relation used above holds for a dilute gas of any speed, which is a stronger statement than it looks — it is true for photons in the sense that , but the pressure of a photon gas is not because the number of photons is not conserved.
Where the model stops
The relativistic crossing is drawn for a gas of fixed particle number. A real gas hot enough for its electrons to be relativistic is hot enough to make electron–positron pairs, and once pair production starts the particle count is a function of temperature. The heat capacity then acquires a term far larger than anything here, because energy goes into making mass rather than into motion — which is the instability that ends the life of a very massive star.
Nothing here says how fast anything happens. Equipartition describes an equilibrium and is silent about whether it is reached. A mode that exchanges energy with the rest of the system very slowly is out of equilibrium at any practical timescale and holds whatever it was given, which is why a shock-heated gas has a different electron temperature and ion temperature for a long time.
Two coupled oscillators are the simplest system that has equipartition arriving, and watching it arrive is the point: energy passes back and forth between them and their time-averages equalise, but the equalising takes as long as the coupling is weak. The theorem asserts the average and says nothing about the exchange, and in a system with a weak enough coupling the average is not reached within any time anybody has watched.
And the four thirds is a boundary rather than a verdict. A body at exactly is neutrally stable in this analysis, and what decides its fate is whatever has been neglected: general relativity, which makes things worse; electrostatic corrections, which help; rotation, which helps. A real star near the boundary is decided by the corrections, and the boundary only says which corrections matter.
The weight every average here is taken against is an exponential in the energy over , and that is the whole reason the exponent comes out of the average so cleanly. A power law inside an exponential integrates to a gamma function, and the ratio of two gamma functions differing by one in their argument is that argument. It is arithmetic rather than physics, which is why the result holds for coordinates that have nothing physically in common.
What the pictures cannot show
The hero figure plots a mean energy against a continuous exponent, and no real system has a continuously adjustable one. The points at and are physical; the curve between them is an interpolation through cases nobody can build, drawn because it makes the shape of the dependence legible and not because a system at is available.
Nor can any of these figures show what is being averaged. Every quantity here is an expectation over a distribution across particles, and the object that the theorem is about — a single coordinate’s share, fluctuating wildly from instant to instant and from particle to particle — appears nowhere. What is drawn is the mean of something whose typical departure from the mean is of the same size as the mean itself.
The result is a century older than its use
The generalised form is Clausius’s, from 1870, and the phrase he used for the quantity — the virial, from the Latin for force — has attached itself to two rather different results that share the same algebra.
Applied to one coordinate in contact with a heat bath it gives the theorem above and is a statement about temperature. Applied to a whole bound system with no bath at all, and time-averaged rather than ensemble-averaged, it gives the relation between kinetic and potential energy that the star figures use. The same expression, two averages, two subjects: one belongs to thermodynamics and the other to mechanics, and a great deal of confusion has come from the shared name.
What is common to both is the reason the answer is so free of detail. The quantity being averaged is homogeneous of some degree in its coordinate, and the average of a homogeneous function against an exponential of itself depends only on the degree. Everything else integrates away. That is why contains no spring constant, no mass and no volume — and why the same arithmetic serves a molecule in a bottle and a galaxy in a cluster.
Where this ladder goes next
Two rungs stand on equipartition. The first found the theorem’s quantum failure: modes with a level spacing above kT are not available, and hydrogen’s heat capacity climbs a staircase because of it. This one finds a failure with no quantum mechanics in it — the theorem is about quadratic terms, and a term that is not quadratic pays a different share.
The habit worth carrying away concerns shorthands. A rule stated as a count is usually a rule about a shape, restated for the case where every item has the same shape. Equipartition counts quadratic terms and was taught as counting motions because in the systems it was invented for the two agreed. The same substitution is behind half the misapplications in thermodynamics, and the way to detect it is to look for a case where the shapes differ.
What is left on this ladder is the theorem’s other half. The virial relation used here to get out of one coordinate, summed over every coordinate of a bound system, gives a relation between average kinetic and potential energies that holds for a molecule, a star cluster and a galaxy alike — and turning it into a measurement of a mass that cannot be weighed is what it has mostly been used for.
Part 2 of 7
This essay is one argument about Equipartition. The others:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
Adiabatic indexThe Boltzmann factorDegrees of freedomEquipartitionGravitational collapseHeat capacityInternal energyKinetic theoryPhase spaceRelativistic gasTemperatureVirial theorem
- Half a kT in a piece of wire degrees of freedom, equipartition, temperature
- The correction that took a century degrees of freedom, equipartition, heat capacity
- A boiling point is a pressure, not a temperature the boltzmann factor, temperature
- Pressure is a rate of arrival, and the gas law falls out of counting equipartition, temperature
- The curve that would not come down equipartition, temperature
- The engine that pays back more than it takes heat capacity, temperature