Weighing what cannot be put on a scale
Assumes: The share that is not half a kT · Half a kT for every way of moving
The share that is not half a kT ends by naming what is left: the theorem’s other half. Equipartition in its familiar form assigns half a to each quadratic coordinate. Summed over every coordinate of a bound system, the same derivation gives something with no temperature in it at all.
The relation, and where it comes from
The derivation is one line of integration by parts, and it is the same line that gives equipartition. Take the quantity for a bound system, average its rate of change over a long time, and note that it must average to zero — the system is bound, so cannot grow without limit. Expanding the derivative gives twice the kinetic energy on one side and on the other, and for a potential going as that second quantity is times the potential energy.
So
and there is no temperature anywhere in it. The averages are over time for one orbit, or over members for a crowd, and neither requires a heat bath, a thermometer or an equilibrium. That independence is the first thing worth checking, and the hero figure checks it: the points come from single eccentric orbits, integrated, with the radius verified to vary by at least a fifth so that no orbit is trivially circular.
Two of the exponents carry everything that follows.
, a harmonic well. The two averages are equal, which is exactly the ordinary equipartition statement in disguise: half a in the kinetic term and half a in the potential one, so a solid’s heat capacity counts two quadratic terms per direction and comes out at per atom rather than — and fails at low temperature for the reason every mode count fails there.
, gravity and the Coulomb force. Twice the kinetic energy is minus the potential energy, so the total energy is
— minus the kinetic energy. A bound gravitating system’s total energy is fixed by how fast its parts are moving, and that is the fact the rest of this essay is about.
The same exponent, inside an atom
Before the astronomical use, it is worth noticing that the Coulomb force has the same exponent and therefore the same relation — and that it produces a statement about chemistry which almost nobody expects.
A hydrogen atom in its ground state has a total energy of electronvolts. By the relation, , so its average kinetic energy is electronvolts and its average potential energy is . Neither of those is a small correction to the total; they are both twice as large as it, and they cancel to within a factor of two.
That has a consequence for bonding that reverses the usual picture. When two hydrogen atoms form a molecule the total energy falls by 4.7 electronvolts. By the same relation the potential energy must fall by 9.4 and the kinetic energy must rise by 4.7. So a chemical bond does not form because the electrons have settled into a lower-energy, calmer arrangement: the electrons end up moving faster, and the bond exists because the electrostatic energy fell by twice as much.
The relation is also used as a quality test. Any approximate wavefunction can be asked whether it satisfies , and a poor one does not — the ratio is reported routinely in electronic-structure calculations as a check that the basis is adequate and the geometry optimised. A theorem that must hold exactly is a test that costs nothing to apply, and it catches errors that no comparison with experiment would identify as errors of that kind.
The instrument
A cloud of things orbiting each other has a kinetic energy that can be measured and a potential energy that contains its mass. That is enough.
The velocities come from Doppler shifts, which give only the component along the line of sight — so the measurable quantity is a one-dimensional velocity dispersion , and taking the motions as isotropic makes the kinetic energy . The potential energy of a cloud of extent is of order . Setting twice the first equal to minus the second gives
Nothing in that expression is a property of the matter. It does not ask what the cluster is made of, whether the mass emits light, whether it is gas or stars or something else, or how it is arranged beyond its extent. It counts what does not shine exactly as readily as what does, which is what makes it an instrument rather than an inventory. In that respect it is the opposite of a photometric mass, which counts only what radiates and has to be corrected by a mass-to-light ratio nobody can measure directly.
Fritz Zwicky applied it to the Coma cluster in 1933 with worse data than the figure uses, found the mass far above what the galaxies’ light accounted for, called what was missing dunkle Materie, and was largely ignored for forty years. Vera Rubin’s rotation curves in the 1970s made the same point on individual galaxies, by reading a field’s falloff as a statement about the shape of its source, and the two lines of evidence are independent: one measures a bound crowd’s total energy and the other measures a field at a radius.
Why the conclusion survives the errors
The estimate looks fragile — an order-of-magnitude potential energy, an assumed isotropy, a radius that has to be chosen — and it is worth seeing why it is not.
The mass goes as . A dispersion measured thirty per cent too high inflates the mass by seventy per cent, which against a factor of fifty is nothing. The geometric factor of three depends on the cluster’s density profile and varies between two and about five across plausible profiles, which is a factor of two and a half. The radius enters linearly and is known to a factor of two at worst.
Multiply every error in the same direction and the estimate moves by less than an order of magnitude. A factor of fifty is not reachable from a factor of ten. That insensitivity is the ordinary situation for a result obtained from a conservation law rather than from a model: the answer carries the uncertainty of its inputs and no more, and none of the inputs is uncertain by the amount needed.
What the argument is sensitive to is not a number at all. It is an assumption, and it is the third figure.
The condition underneath
The relation holds for a system whose averages have settled. A cloud released from rest does not satisfy it at first — all its energy is potential — and it takes a few traverses for the orbits to mix and the averages to stop drifting.
That draws the method’s domain of validity, and it draws it from two numbers anybody quoting a virial mass already has. A globular cluster is virialised beyond any doubt. A galaxy is. A cluster is, marginally, which is why the estimates for clusters carry the systematic uncertainty they do and why the ones for clusters still in the act of merging are worse. A supercluster is not, and a “virial mass” for one is a number with no theorem behind it.
The value of a stated condition is that it can be checked before the measurement rather than argued about afterwards. Coma’s nine crossing times is a fact about Coma, available from its size and its dispersion, and it licenses the mass estimate in a way that no amount of care with the spectroscopy would.
What the dispersion is made of
There is a question the mass estimate raises and does not answer, and it deserves a section because answering it is what turned a discrepancy into a subject.
A velocity dispersion is a spread in speeds. For it to mean what the virial relation needs, the things whose speeds are being measured must be tracers — objects moving in the cluster’s potential and nothing else, sampling it fairly. Galaxies are reasonable tracers and they are also the only things whose individual velocities can be measured, which is a coincidence the method depends on.
Two later measurements avoid the coincidence, and their agreement is what makes the conclusion difficult to argue with.
The hot gas. A cluster is filled with gas at ten million kelvin, which radiates X-rays, and the gas’s temperature is a direct reading of the depth of the potential well it is sitting in: a gas in hydrostatic equilibrium in a potential has a temperature set by that potential and its own density profile. That is the virial relation again, applied to a fluid instead of to a crowd, and the mass it returns agrees with the galaxies’ for relaxed clusters.
And the lensing. Mass bends light, so the distortion of the shapes of galaxies behind a cluster maps its mass directly, with no assumption about equilibrium, tracers or dynamics whatever. That is a measurement of the same quantity by a route that shares not one step with the other two, and it agrees.
Three methods, three sets of assumptions, one answer. The virial estimate is the oldest and the crudest, and it is the one that had to be trusted for forty years before the other two existed. What makes it worth teaching is not its precision but how little it assumes: a theorem about bound systems, a Doppler shift and an angular size.
The half that runs the other way
There is a second use of the relation that has nothing to do with weighing, and it is the one gravitation needs.
If for a gravitating system, then taking energy away from such a system raises its kinetic energy. That is a negative heat capacity, nothing stable has one, and it is the reason a ball of gas heats up as it cools: a star radiating into cold space is not cooling down, it is contracting and getting hotter, and its whole life is a slow fall it cannot stop. That same sign is what eventually puts a massive core past the mass cold matter can hold up.
The same sign explains why a self-gravitating system cannot come to equilibrium with a heat bath, why star clusters evaporate rather than settling, and why the statistical mechanics of gravitating systems is a subject with its own pathologies — including that no configuration of such a system maximises the count an entropy is, so the usual route to a distribution has nothing to maximise. Every one of those follows from the exponent being , which is the leftmost point of the hero figure, and from nothing else about gravity.
The inequality, which is a criterion for collapse
The relation is an equality for a system in a steady state. Written as an inequality it becomes a criterion for whether a cloud can be in a steady state at all, and that is its other large use.
A cloud with has too little kinetic energy to support itself and must contract. Written out with for a gas of particles and , that inequality gives a mass above which a cloud of given temperature and density collapses — the Jeans mass, which is the virial theorem rearranged and which sets the scale on which a molecular cloud fragments into stars.
The same comparison decides whether a disturbance in a self-gravitating medium grows instead of travelling: a sound wave longer than the Jeans length has gravity winning over pressure, and it collapses rather than propagating. So the criterion and the dispersion relation are the same statement, one written as an energy balance and one as a frequency that has gone imaginary.
There is a third reading that is worth having because it is the one that applies to an accelerating universe. A structure is bound if its own gravity beats the expansion, which is the same inequality with the expansion’s kinetic energy on the other side — and the crossing-time figure below is that comparison made with two numbers instead of an integral.
A steady state is not an equilibrium, and a factor is not a shape
The potential energy is taken as to within a factor. The exact coefficient depends on how the mass is distributed: a uniform sphere gives , a more concentrated profile gives more. The standard estimates use a profile fitted to the galaxy distribution, which assumes the dark mass follows the light — an assumption the measurement is partly about.
The velocities are taken as isotropic. They need not be: a cluster whose orbits are preferentially radial has a line-of-sight dispersion that overstates the total near its centre and understates it at the edge. Separating the mass profile from the orbital anisotropy from projected data alone is not possible without an extra assumption, and the ambiguity is a real limit on precision.
The dark mass is assumed to be collisionless and in the same potential. It need not be distributed like the galaxies, and if it is not, the geometric factor is wrong by the difference — which is one of the things reading the field’s falloff at a radius settles independently, because a falloff reports a shape rather than a total.
Members have to be identified. A galaxy projected onto the cluster but not bound to it inflates the dispersion, and since the mass goes as the square of that, a few interlopers matter. The standard remedies are iterative and they are the largest source of disagreement between published masses of the same cluster.
And the relation is for a system in a steady state, not an equilibrium. Nothing in it says the system is at a temperature, and a cluster is not: its galaxies, its hot gas and its dark mass each have their own velocity dispersion and there is no reason for the three to agree. They can be compared, and the comparison is one of the tests of the whole picture — the gas temperature and the galaxy dispersion give consistent masses for relaxed clusters, which is a genuine check that nothing has been assumed into the answer.
A converged ratio that hides how long it took to converge
The first figure draws a ratio of two time-averages and hides the orbits they came from. The averages converge slowly and non-monotonically — an eccentric orbit spends most of its time near apoapsis, so the running average of the kinetic energy oscillates with the radial period and only settles after many of them. A plot of the converged ratio says nothing about how long convergence took, which is the quantity the third figure is about in a different guise.
The cluster figure draws a required mass and cannot show what that mass is. Everything in this essay is consistent with the missing mass being ordinary matter that happens not to shine, and the reason it is not is evidence from elsewhere entirely — the abundances of the light elements, the fluctuations in the microwave background, the ratio of hot gas to total mass in clusters. The virial argument establishes that mass is missing and is silent about what it is made of, which is exactly the strength that makes it useless for the second question.
And the crossing-time figure draws a ratio of two times and cannot show what happens between them. A system at three crossing times is not either virialised or not; it is in the middle of becoming so, with its inner parts settled and its outer parts still falling in, and the virial relation applied to the whole of it returns something between the right answer and no answer. Where a real cluster sits on that continuum is a question about its individual history.
Still open: what the theorem is worth when the averages do not exist
Everything here assumes the time-averages converge. For a bound orbit in a smooth potential they do, and the hero figure watches them do it. For a system with many bodies interacting directly, it is less obvious.
A gravitating system of bodies has no equilibrium state in the ordinary sense. Its energy can always be lowered by tightening a binary and flinging a third body out, so there is no maximum-entropy configuration to relax to, and the system evolves indefinitely: the core contracts, the halo evaporates, and nothing settles. Over long enough times a star cluster does not virialise, it dissolves.
What the virial relation describes is therefore a quasi-steady state — long-lived compared with a crossing time and short-lived compared with the evaporation time — and whether that separation of timescales is clean enough depends on the system. For a globular cluster the two differ by a factor of a thousand and the description is excellent. For a small group of galaxies they differ by much less. And for the largest structures the relation is not applicable at all, which is where the mass estimates have to come from gravitational lensing or from the microwave background instead, by arguments with no virial theorem in them.
The habit worth carrying away is the one this whole essay is. A theorem that relates two averages is a measuring instrument whenever one of them is observable and the other contains what is wanted. Nothing about the virial relation was designed for weighing anything; it is a consequence of a system being bound. Its use as the only way to weigh most of the mass in the universe came from noticing that one of its two averages is a Doppler shift.
Part 7 of 7
This essay is one argument about Equipartition. The others:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
Dark matterDegrees of freedomEquipartitionGravitational collapseKinetic energyMeasurementPhase spacePotential energyStatistical mechanicsVirial theorem
- Half a kT in a piece of wire degrees of freedom, equipartition, measurement, statistical mechanics
- The temperature a molecule does not have degrees of freedom, equipartition, statistical mechanics
- The correction that took a century degrees of freedom, equipartition
- The heat that changes no temperature, and where it actually goes equipartition, potential energy
- The plane in which three bodies are flat measurement, phase space