Thermodynamics

Weighing what cannot be put on a scale

Summed over every coordinate of a bound system, equipartition stops being a statement about temperature and becomes a relation between two averages: twice the kinetic energy equals n times the potential energy for a potential going as the nth power. For gravity that fixes a bound system's total energy from how fast its parts move — so a Doppler shift and an angular size return a mass, and for the Coma cluster the mass they return is fifty times the mass that shines.

Assumes: The share that is not half a kT · Half a kT for every way of moving

The share that is not half a kT ends by naming what is left: the theorem’s other half. Equipartition in its familiar form assigns half a kBTk_BT to each quadratic coordinate. Summed over every coordinate of a bound system, the same derivation gives something with no temperature in it at all.

Twice the kinetic energy, and what it equals. Twice the time-averaged kinetic energy of a bound orbit, divided by its time-averaged potential energy, against the power with which that potential depends on separation. Each point is a measurement: an eccentric orbit integrated for more than a hundred radial periods, with the two averages accumulated along it, and the radius checked to vary by at least a fifth so that the orbit is not trivially circular. The line is the exponent itself, and the points miss it by at most 6.6e-4. Two cases carry everything. At n = 2, a harmonic well, the two energies are equal — which is the ordinary equipartition statement, half a kT to the kinetic term and half a kT to the potential one. At n = −1, which is gravity and the Coulomb force, twice the kinetic energy equals minus the potential energy, so the total energy of a bound system is minus its kinetic energy. Nothing about temperature entered, and nothing about equilibrium: the relation holds for one orbit averaged over time as well as for a crowd averaged over members.
Fig. 1 Twice the time-averaged kinetic energy of a bound orbit, divided by its time-averaged potential energy, against the power with which that potential depends on separation. Each point is an eccentric orbit integrated for more than a hundred radial periods with both averages accumulated along it. The line is the exponent itself. At n = 2 the two energies are equal; at n = −1, which is gravity, twice the kinetic energy is minus the potential energy.

The relation, and where it comes from

The derivation is one line of integration by parts, and it is the same line that gives equipartition. Take the quantity rp\sum \mathbf{r}\cdot\mathbf{p} for a bound system, average its rate of change over a long time, and note that it must average to zero — the system is bound, so rp\sum\mathbf{r}\cdot\mathbf{p} cannot grow without limit. Expanding the derivative gives twice the kinetic energy on one side and rV\sum \mathbf{r}\cdot\nabla V on the other, and for a potential going as rnr^n that second quantity is nn times the potential energy.

So

2T=nV2\langle T\rangle = n\langle V\rangle

and there is no temperature anywhere in it. The averages are over time for one orbit, or over members for a crowd, and neither requires a heat bath, a thermometer or an equilibrium. That independence is the first thing worth checking, and the hero figure checks it: the points come from single eccentric orbits, integrated, with the radius verified to vary by at least a fifth so that no orbit is trivially circular.

Two of the exponents carry everything that follows.

n=2n = 2, a harmonic well. The two averages are equal, which is exactly the ordinary equipartition statement in disguise: half a kBTk_BT in the kinetic term and half a kBTk_BT in the potential one, so a solid’s heat capacity counts two quadratic terms per direction and comes out at 3kB3k_B per atom rather than 32\tfrac{3}{2} — and fails at low temperature for the reason every mode count fails there.

n=1n = -1, gravity and the Coulomb force. Twice the kinetic energy is minus the potential energy, so the total energy is

E=T+V=TE = \langle T\rangle + \langle V\rangle = -\langle T\rangle

— minus the kinetic energy. A bound gravitating system’s total energy is fixed by how fast its parts are moving, and that is the fact the rest of this essay is about.

What a coordinate is worth, by the shape of its energy. The mean energy stored in one coordinate, in units of kT, against the power with which that coordinate enters the energy. The curve is 1/n and the points are the same average obtained by integrating the Boltzmann weight numerically, agreeing to 3.1e-3 per cent at worst. an energy going as |x|^0.5 holds 2.0001 kT; an energy going as |x|^1 holds 1.0000 kT; an energy going as |x|^2 holds 0.5000 kT; an energy going as |x|^3 holds 0.3333 kT; an energy going as |x|^6 holds 0.1667 kT; an energy going as |x|^10 holds 0.1000 kT. The familiar half a kT is the case n = 2 and nothing more general than that: a coordinate whose energy is linear in it — the momentum of an ultrarelativistic particle, or a field with no restoring force but a constant tension — carries a whole kT, and one confined by a very steep wall carries almost nothing. Equipartition is a theorem about quadratic terms, and calling it a theorem about degrees of freedom is the substitution that makes it fail.
Fig. 2 The generalised equipartition theorem the relation comes from: the mean energy in one coordinate, in units of kT, against the power with which that coordinate enters the energy. It is kT/n, so a quadratic coordinate holds a half and a coordinate in a steep well holds almost nothing. The virial relation is this statement summed over every coordinate of a bound system instead of applied to one at a time, and the same integration by parts produces both.

The same exponent, inside an atom

Before the astronomical use, it is worth noticing that the Coulomb force has the same exponent and therefore the same relation — and that it produces a statement about chemistry which almost nobody expects.

A hydrogen atom in its ground state has a total energy of 13.6-13.6 electronvolts. By the relation, E=TE = -\langle T\rangle, so its average kinetic energy is +13.6+13.6 electronvolts and its average potential energy is 27.2-27.2. Neither of those is a small correction to the total; they are both twice as large as it, and they cancel to within a factor of two.

That has a consequence for bonding that reverses the usual picture. When two hydrogen atoms form a molecule the total energy falls by 4.7 electronvolts. By the same relation the potential energy must fall by 9.4 and the kinetic energy must rise by 4.7. So a chemical bond does not form because the electrons have settled into a lower-energy, calmer arrangement: the electrons end up moving faster, and the bond exists because the electrostatic energy fell by twice as much.

Counting terms against measuring them. Measured heat capacities at constant volume at room temperature, in units of R, against the equipartition prediction of half a unit for every quadratic term in the energy. Monatomic gases have three translations and nothing else, and the prediction is exact. Diatomic gases have two rotations as well and the prediction is right if — and only if — the vibration is left out of the count, which nothing in classical physics licenses. The largest disagreement in the table is 0.30R, and it belongs to the molecules with the softest vibrations, which are exactly the ones whose vibrational steps are small enough for room temperature to reach.
Fig. 3 Equipartition in its familiar form, for comparison: the heat capacity of four gases against the number of quadratic terms each molecule’s energy has. That count and the virial relation come out of the same integration by parts — one applied to a single coordinate at a temperature, the other summed over every coordinate of a bound system. The first predicts how much energy a molecule takes to warm; the second fixes the ratio between the kinetic and electrostatic energies that hold it together.

The relation is also used as a quality test. Any approximate wavefunction can be asked whether it satisfies 2T=V2\langle T\rangle = -\langle V\rangle, and a poor one does not — the ratio is reported routinely in electronic-structure calculations as a check that the basis is adequate and the geometry optimised. A theorem that must hold exactly is a test that costs nothing to apply, and it catches errors that no comparison with experiment would identify as errors of that kind.

The instrument

A cloud of things orbiting each other has a kinetic energy that can be measured and a potential energy that contains its mass. That is enough.

The velocities come from Doppler shifts, which give only the component along the line of sight — so the measurable quantity is a one-dimensional velocity dispersion σ\sigma, and taking the motions as isotropic makes the kinetic energy 32Mσ2\tfrac{3}{2}M\sigma^2. The potential energy of a cloud of extent RR is of order GM2/R-GM^2/R. Setting twice the first equal to minus the second gives

M3σ2RGM \approx \frac{3\sigma^2 R}{G}

Weighing a thing that cannot be put on a scale. The mass a self-gravitating cloud must have, in solar masses, against the spread in its members' line-of-sight velocities, for three cluster radii — both axes logarithmic. The relation is 3σ²R/G, which follows from twice the kinetic energy equalling minus the potential energy and from nothing else: there is no assumption about what the cluster is made of, whether the mass emits light, or how it is distributed beyond its extent. The marked point is the Coma cluster at a thousand kilometres per second across one and a half megaparsecs, which gives 1.0e+15 solar masses. Its galaxies' stars come to about 2.0e+13, so the mass the motions require is 52 times the mass the light accounts for. Zwicky did this arithmetic in 1933 on worse data, got a larger factor, called what was missing dunkle Materie, and was largely ignored for forty years. The steepness of the axis is why the conclusion is robust: the mass goes as the square of the dispersion, so a velocity measurement wrong by thirty per cent moves the mass by less than a factor of two and cannot close a factor of fifty.
Fig. 4 The mass a self-gravitating cloud must have, against the spread in its members’ line-of-sight velocities, for three radii. Coma at a thousand kilometres per second across one and a half megaparsecs requires 10¹⁵ solar masses; the stars in its galaxies come to two per cent of that. The mass goes as the square of the dispersion, so a velocity wrong by thirty per cent cannot close a factor of fifty.

Nothing in that expression is a property of the matter. It does not ask what the cluster is made of, whether the mass emits light, whether it is gas or stars or something else, or how it is arranged beyond its extent. It counts what does not shine exactly as readily as what does, which is what makes it an instrument rather than an inventory. In that respect it is the opposite of a photometric mass, which counts only what radiates and has to be corrected by a mass-to-light ratio nobody can measure directly.

Fritz Zwicky applied it to the Coma cluster in 1933 with worse data than the figure uses, found the mass far above what the galaxies’ light accounted for, called what was missing dunkle Materie, and was largely ignored for forty years. Vera Rubin’s rotation curves in the 1970s made the same point on individual galaxies, by reading a field’s falloff as a statement about the shape of its source, and the two lines of evidence are independent: one measures a bound crowd’s total energy and the other measures a field at a radius.

Why the conclusion survives the errors

The estimate looks fragile — an order-of-magnitude potential energy, an assumed isotropy, a radius that has to be chosen — and it is worth seeing why it is not.

The mass goes as σ2\sigma^2. A dispersion measured thirty per cent too high inflates the mass by seventy per cent, which against a factor of fifty is nothing. The geometric factor of three depends on the cluster’s density profile and varies between two and about five across plausible profiles, which is a factor of two and a half. The radius enters linearly and is known to a factor of two at worst.

Multiply every error in the same direction and the estimate moves by less than an order of magnitude. A factor of fifty is not reachable from a factor of ten. That insensitivity is the ordinary situation for a result obtained from a conservation law rather than from a model: the answer carries the uncertainty of its inputs and no more, and none of the inputs is uncertain by the amount needed.

What the argument is sensitive to is not a number at all. It is an assumption, and it is the third figure.

The condition underneath

The relation holds for a system whose averages have settled. A cloud released from rest does not satisfy it at first — all its energy is potential — and it takes a few traverses for the orbits to mix and the averages to stop drifting.

How many times the parts have crossed. How many crossing times each self-gravitating system has had since the beginning, against its own size in megaparsecs, on logarithmic axes. A crossing time is its radius divided by the spread in its members' speeds, and the virial relation holds for a system that has had several of them — long enough for the orbits to have mixed and the averages to have settled. a globular cluster has had 14,114, a galaxy has had 188, a galaxy group has had 8.47, the Coma cluster has had 9.41, a supercluster has had 0.14. The line at one is where the method stops meaning anything, and a supercluster is below it: its parts have not had time to complete a single traverse, so they are still moving apart with the expansion rather than orbiting anything, and a velocity dispersion measured across it is not a virial dispersion. That is the systematic error the mass estimates of the previous figure have to survive, and it is why the dark-matter argument is made on clusters rather than on the largest structures there are.
Fig. 5 How many crossing times each system has had since the beginning, against its size. A crossing time is the radius divided by the speed spread. A globular cluster has had fourteen thousand; Coma has had nine; a supercluster has had a seventh of one. Below the line the parts have not completed a single traverse, are still separating with the expansion, and their velocity spread is not a virial dispersion at all.

That draws the method’s domain of validity, and it draws it from two numbers anybody quoting a virial mass already has. A globular cluster is virialised beyond any doubt. A galaxy is. A cluster is, marginally, which is why the estimates for clusters carry the systematic uncertainty they do and why the ones for clusters still in the act of merging are worse. A supercluster is not, and a “virial mass” for one is a number with no theorem behind it.

The value of a stated condition is that it can be checked before the measurement rather than argued about afterwards. Coma’s nine crossing times is a fact about Coma, available from its size and its dispersion, and it licenses the mass estimate in a way that no amount of care with the spectroscopy would.

What the dispersion is made of

There is a question the mass estimate raises and does not answer, and it deserves a section because answering it is what turned a discrepancy into a subject.

A velocity dispersion is a spread in speeds. For it to mean what the virial relation needs, the things whose speeds are being measured must be tracers — objects moving in the cluster’s potential and nothing else, sampling it fairly. Galaxies are reasonable tracers and they are also the only things whose individual velocities can be measured, which is a coincidence the method depends on.

Two later measurements avoid the coincidence, and their agreement is what makes the conclusion difficult to argue with.

The hot gas. A cluster is filled with gas at ten million kelvin, which radiates X-rays, and the gas’s temperature is a direct reading of the depth of the potential well it is sitting in: a gas in hydrostatic equilibrium in a potential has a temperature set by that potential and its own density profile. That is the virial relation again, applied to a fluid instead of to a crowd, and the mass it returns agrees with the galaxies’ for relaxed clusters.

And the lensing. Mass bends light, so the distortion of the shapes of galaxies behind a cluster maps its mass directly, with no assumption about equilibrium, tracers or dynamics whatever. That is a measurement of the same quantity by a route that shares not one step with the other two, and it agrees.

Three methods, three sets of assumptions, one answer. The virial estimate is the oldest and the crudest, and it is the one that had to be trusted for forty years before the other two existed. What makes it worth teaching is not its precision but how little it assumes: a theorem about bound systems, a Doppler shift and an angular size.

The half that runs the other way

There is a second use of the relation that has nothing to do with weighing, and it is the one gravitation needs.

If E=TE = -\langle T\rangle for a gravitating system, then taking energy away from such a system raises its kinetic energy. That is a negative heat capacity, nothing stable has one, and it is the reason a ball of gas heats up as it cools: a star radiating into cold space is not cooling down, it is contracting and getting hotter, and its whole life is a slow fall it cannot stop. That same sign is what eventually puts a massive core past the mass cold matter can hold up.

The same sign explains why a self-gravitating system cannot come to equilibrium with a heat bath, why star clusters evaporate rather than settling, and why the statistical mechanics of gravitating systems is a subject with its own pathologies — including that no configuration of such a system maximises the count an entropy is, so the usual route to a distribution has nothing to maximise. Every one of those follows from the exponent being 1-1, which is the leftmost point of the hero figure, and from nothing else about gravity.

A gas crossing from three halves to three. The mean kinetic energy per particle, in units of kT, against temperature in units of the particle's own rest energy, integrated over the relativistic Maxwell distribution. It is 1.501 kT at the cold end — the three halves of a non-relativistic gas, which is three quadratic momentum components at half a kT each — and 3.000 kT at the hot end, where the energy is pc and each component is linear rather than quadratic, so each carries a whole kT. The second curve is 1 + P/u, the ratio of specific heats the energy budget of a star uses: it falls from 5/3 to 4/3 across the same crossing. Nothing in between is a mixture of two gases; it is one gas whose particles are neither slow nor fast, and the drift between the two plateaus takes about four decades of temperature.
Fig. 6 And the case where the relation changes: the mean kinetic energy per particle of a gas, in units of kT, as the particles become relativistic. A slow particle’s energy is quadratic in its momentum and carries three halves of a kT; a fast one’s is linear and carries three. The virial coefficient follows the same crossing, so a cloud held up by relativistic particles satisfies a different relation from one held up by slow ones — and the difference is exactly what puts a massive star on the edge of being unable to hold itself up.

The inequality, which is a criterion for collapse

The relation is an equality for a system in a steady state. Written as an inequality it becomes a criterion for whether a cloud can be in a steady state at all, and that is its other large use.

A cloud with 2T+V<02T + V < 0 has too little kinetic energy to support itself and must contract. Written out with T=32NkBTT = \tfrac{3}{2}Nk_BT for a gas of NN particles and V=3GM2/5RV = -3GM^2/5R, that inequality gives a mass above which a cloud of given temperature and density collapses — the Jeans mass, which is the virial theorem rearranged and which sets the scale on which a molecular cloud fragments into stars.

The same comparison decides whether a disturbance in a self-gravitating medium grows instead of travelling: a sound wave longer than the Jeans length has gravity winning over pressure, and it collapses rather than propagating. So the criterion and the dispersion relation are the same statement, one written as an energy balance and one as a frequency that has gone imaginary.

There is a third reading that is worth having because it is the one that applies to an accelerating universe. A structure is bound if its own gravity beats the expansion, which is the same inequality with the expansion’s kinetic energy on the other side — and the crossing-time figure below is that comparison made with two numbers instead of an integral.

A steady state is not an equilibrium, and a factor is not a shape

The potential energy is taken as GM2/R-GM^2/R to within a factor. The exact coefficient depends on how the mass is distributed: a uniform sphere gives 3GM2/5R-3GM^2/5R, a more concentrated profile gives more. The standard estimates use a profile fitted to the galaxy distribution, which assumes the dark mass follows the light — an assumption the measurement is partly about.

The velocities are taken as isotropic. They need not be: a cluster whose orbits are preferentially radial has a line-of-sight dispersion that overstates the total near its centre and understates it at the edge. Separating the mass profile from the orbital anisotropy from projected data alone is not possible without an extra assumption, and the ambiguity is a real limit on precision.

The dark mass is assumed to be collisionless and in the same potential. It need not be distributed like the galaxies, and if it is not, the geometric factor is wrong by the difference — which is one of the things reading the field’s falloff at a radius settles independently, because a falloff reports a shape rather than a total.

Members have to be identified. A galaxy projected onto the cluster but not bound to it inflates the dispersion, and since the mass goes as the square of that, a few interlopers matter. The standard remedies are iterative and they are the largest source of disagreement between published masses of the same cluster.

And the relation is for a system in a steady state, not an equilibrium. Nothing in it says the system is at a temperature, and a cluster is not: its galaxies, its hot gas and its dark mass each have their own velocity dispersion and there is no reason for the three to agree. They can be compared, and the comparison is one of the tests of the whole picture — the gas temperature and the galaxy dispersion give consistent masses for relaxed clusters, which is a genuine check that nothing has been assumed into the answer.

A converged ratio that hides how long it took to converge

The first figure draws a ratio of two time-averages and hides the orbits they came from. The averages converge slowly and non-monotonically — an eccentric orbit spends most of its time near apoapsis, so the running average of the kinetic energy oscillates with the radial period and only settles after many of them. A plot of the converged ratio says nothing about how long convergence took, which is the quantity the third figure is about in a different guise.

The cluster figure draws a required mass and cannot show what that mass is. Everything in this essay is consistent with the missing mass being ordinary matter that happens not to shine, and the reason it is not is evidence from elsewhere entirely — the abundances of the light elements, the fluctuations in the microwave background, the ratio of hot gas to total mass in clusters. The virial argument establishes that mass is missing and is silent about what it is made of, which is exactly the strength that makes it useless for the second question.

And the crossing-time figure draws a ratio of two times and cannot show what happens between them. A system at three crossing times is not either virialised or not; it is in the middle of becoming so, with its inner parts settled and its outer parts still falling in, and the virial relation applied to the whole of it returns something between the right answer and no answer. Where a real cluster sits on that continuum is a question about its individual history.

Still open: what the theorem is worth when the averages do not exist

Everything here assumes the time-averages converge. For a bound orbit in a smooth potential they do, and the hero figure watches them do it. For a system with many bodies interacting directly, it is less obvious.

A gravitating system of NN bodies has no equilibrium state in the ordinary sense. Its energy can always be lowered by tightening a binary and flinging a third body out, so there is no maximum-entropy configuration to relax to, and the system evolves indefinitely: the core contracts, the halo evaporates, and nothing settles. Over long enough times a star cluster does not virialise, it dissolves.

What the virial relation describes is therefore a quasi-steady state — long-lived compared with a crossing time and short-lived compared with the evaporation time — and whether that separation of timescales is clean enough depends on the system. For a globular cluster the two differ by a factor of a thousand and the description is excellent. For a small group of galaxies they differ by much less. And for the largest structures the relation is not applicable at all, which is where the mass estimates have to come from gravitational lensing or from the microwave background instead, by arguments with no virial theorem in them.

The habit worth carrying away is the one this whole essay is. A theorem that relates two averages is a measuring instrument whenever one of them is observable and the other contains what is wanted. Nothing about the virial relation was designed for weighing anything; it is a consequence of a system being bound. Its use as the only way to weigh most of the mass in the universe came from noticing that one of its two averages is a Doppler shift.

Part 7 of 7

This essay is one argument about Equipartition. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

Dark matterDegrees of freedomEquipartitionGravitational collapseKinetic energyMeasurementPhase spacePotential energyStatistical mechanicsVirial theorem