The turn that two pushes leave behind
Assumes: Speeds that refuse to add, and the quantity that does · Now is a choice of slicing
Velocities that point the same way compose by a rule that is awkward and completely understood. Velocities that do not point the same way compose by a rule that is awkward and leaves something behind.
Each curve here is computed by multiplying two boost matrices and pulling the rotation out of the product, rather than by evaluating a formula. The closed form for perpendicular boosts is used to check the extraction and appears nowhere in the drawing, which is worth saying because the extraction is easy to get wrong in a way that looks right — reading the composite velocity off the matrix’s first row rather than its first column leaves a residue that is not a rotation at all, and gives an angle about ten per cent out.
The one-dimensional case, which hides it
Along a line, speeds compose by
with . It is the rule that keeps everything below the speed of light, and it is famous.
The awkwardness of that formula disappears in the right variable. Writing makes the composition , ordinary addition, and — the rapidity — is the natural parameter for a boost in the same way an angle is the natural parameter for a rotation.
The reason it stops working is that rapidity in more than one dimension is a vector, and boosts in different directions do not commute — so their “rapidities” cannot simply add, any more than rotations about different axes add.
The rotation, extracted
Write a boost in three dimensions on the coordinates as a matrix, compose two of them, and look at the result. A pure boost matrix is symmetric; the product of two boosts in different directions is not. That asymmetry is the rotation, present before any decision about how to factor the answer.
A single boost is the familiar scissor of a moving frame’s axes. A second boost in a perpendicular direction cannot be drawn on that diagram at all — and that is part of why the effect went unremarked for so long. The standard picture of special relativity has one space dimension in it, and the Wigner rotation needs two.
Extracting it is mechanical. Read the composite velocity off the product’s first column, apply the inverse of the pure boost with that velocity, and what remains is a matrix whose time row and time column are — a pure spatial rotation, whose angle is read off its own entries. For two equal boosts at right angles the answer is
and the numerical extraction agrees with it to twelve decimal places, which is what the generator checks before drawing anything.
The consequence is stated most sharply in the language of groups. The Lorentz boosts do not form a group. Compose two and the result is not in the set. What is a group is the set of boosts together with the rotations, which is the Lorentz group proper — so rotations are not an optional extra bolted on to relativity but something the boosts generate whether anybody wants them or not.
There is a way to see that something like this had to happen, without any matrices. A boost is a hyperbolic rotation in a plane containing the time axis, and two rotations in different planes generally do not commute — the same statement as for ordinary rotations in three dimensions, where turning a book about two different axes in the two possible orders leaves it in two different orientations. The commutator of two ordinary rotations is a rotation; the commutator of two boosts in different planes turns out to be a rotation as well, in the plane spanned by the two boost directions. That commutator is not small in any sense that lets it be dropped: it is the same order in the speeds as the boosts’ own second-order terms, which are the terms relativity is about.
The size of the effect is set entirely by how far the two boosts are from parallel. Two boosts a degree apart leave essentially nothing; two at right angles leave nearly the maximum; and the peak in the figure moves past 90° as the speeds rise, reaching 122° at β = 0.95, because at high speed the composite velocity is itself far from either component and the geometry stops being symmetric.
Going round in a circle
A single pair of boosts leaves a fixed rotation. A body moving in a circle is being boosted continuously, in a direction that turns, and the rotations accumulate.
A velocity changed by a perpendicular acceleration is a sequence of boosts at right angles to the current velocity, which is exactly the configuration that leaves the largest rotation behind. Going once round a circle means applying an unbroken sequence of them whose directions cover the full turn, and the rotations do not cancel: each is second order in and all of them have the same sign.
Summing the infinitesimal Wigner angles round one revolution gives
per revolution, in the sense opposite to the orbital motion. It is second order in the speed, so it is small; the point is that it is not zero, and it is not zero for a reason that has nothing to do with any force.
The whole of the effect is . For small speeds that is , so the turn per orbit is — and everything about Thomas precession is the observation that a quantity usually discarded as a second-order correction is here the entire phenomenon, because the first-order term is zero. There is nothing for it to be small compared with.
One further consequence of the rate form deserves a sentence. Because the precession is times the orbital frequency, it grows without bound as the speed approaches that of light, while the orbital frequency itself is bounded — so a highly relativistic particle in a magnetic ring has its spin turning many times per revolution relative to its momentum. That is not a curiosity: it is the basis of every measurement of a particle’s anomalous magnetic moment, where the quantity measured is precisely the difference between the spin precession rate and the momentum rotation rate, and the Thomas term is one of the pieces that has to be subtracted before what is left can be compared with theory.
The factor of two
In 1925 the electron’s spin was proposed to explain the fine structure of atomic spectra — the splitting of lines into close pairs. The calculation is straightforward. An electron orbiting a nucleus sees, in its own frame, the nucleus orbiting it, which is a current, which makes a magnetic field; the electron’s magnetic moment has an energy in that field which depends on whether the spin is aligned with it; and the two orientations therefore have different energies, splitting the line.
The fine structure of hydrogen is a splitting of its levels by parts in a hundred thousand, far too small to draw on any scale that shows the levels themselves. It was measured before it was explained, which is the usual order and is the reason a factor of two in the explanation was immediately visible: the spin–orbit calculation gave twice the observed splitting, and every ingredient in it had been checked.
Every ingredient in that calculation is independently known and independently checkable. The electron’s magnetic moment had been measured; the field it sees follows from the Coulomb field of the nucleus transformed into the moving frame, which is the transformation the electromagnetic essays derive; the orbital radius and speed come from the Bohr model. There is no free parameter anywhere and nothing to adjust, which is what makes the result so awkward.
The calculation gave twice the measured splitting. Not approximately twice, and not twice for some values of the quantum numbers — exactly twice, everywhere, which is the signature of a missing factor rather than a missing mechanism.
A beam of atoms split by a field gradient into two is the direct demonstration that an electron’s magnetic moment takes two values along any chosen axis. Everything in the fine-structure calculation depends on that moment and on the field the electron sits in, and neither of those was wrong — which is what left the frame the calculation was done in as the only remaining suspect.
Thomas supplied the factor in 1926, and it is on this page. The calculation is done in the electron’s rest frame, and that frame is being carried round a closed path in velocity space, so it is rotating — at the rate the figure above draws. The spin, which is not being torqued by anything as it goes round, therefore precesses relative to that frame — a direction carried round a loop and coming back turned — and the extra precession subtracts from the one the magnetic energy predicts. The subtraction is a factor of exactly a half in the limit of small speeds, and it is a half because .
It is worth being precise about which frame is doing what, because this is the step that is usually waved through. The magnetic-energy calculation is set up in the instantaneous rest frame of the electron, and it computes a torque on the spin from the field in that frame. But “the instantaneous rest frame” is a different frame at every instant, and the sequence of them is exactly a sequence of boosts in turning directions. Comparing the spin’s orientation at the start of an orbit with its orientation at the end therefore requires knowing how those frames’ axes relate, and they do not relate by nothing.
Nothing was added to the physics. No new interaction, no new particle, no new constant: the repair was the observation that a sequence of boosts is not a boost, applied to a frame nobody had thought to ask about. Uhlenbeck said afterwards that he and Goudsmit had considered abandoning the spin hypothesis entirely over the discrepancy.
The size of the correction is worth writing down as a rate rather than as a turn per orbit, because that is the form the atomic calculation uses. The precession frequency is times the orbital frequency, which for small speeds is — and the magnetic-energy calculation predicts a spin precession of in the opposite sense. The two differ by exactly the factor drawn, and the sum is half the naive answer. It is one of the very few places in physics where a discrepancy of exactly two was resolved by a term nobody had thought to include rather than by an error in one that was.
The angle is an area
The closing remark about holonomy has a precise form that makes the whole effect look like elementary geometry, and it is worth stating because it turns a matrix computation into a picture.
Velocities in relativity do not live in an ordinary flat space. Composing them is not vector addition, the “distance” between two velocities is the relative rapidity, and the space they inhabit — with that distance — is a hyperbolic space, of constant negative curvature.
Now take two boosts and draw them as a triangle in that space: a geodesic from rest to the first velocity, a geodesic from there to the composite, and a geodesic back to rest. In a flat space the three interior angles of such a triangle would sum to a half turn. In a hyperbolic space they sum to less, and the shortfall — the angular defect — is equal to the triangle’s area.
That defect is the Wigner rotation. Not proportional to it, not approximately it: the angle a frame is turned through by two boosts is exactly the area of the hyperbolic triangle their rapidities span.
Which explains everything the matrix calculation shows without any matrices. Two parallel boosts give a degenerate triangle with no area, so no rotation. Small boosts give a small triangle, and the area of a small triangle goes as the product of two sides — which is why the effect is second order in the speeds. Boosts at right angles give the largest triangle for given side lengths, which is where the rotation peaks. And the peak moves past a right angle at high speed because hyperbolic triangles with long sides behave unlike Euclidean ones.
It also explains why the accumulated rotation round a closed orbit is what it is. Carrying a frame round a closed loop in velocity space encloses a region, and the total turn is that region’s area — which for a circular orbit at speed is the area of a hyperbolic circle of the corresponding rapidity, and works out to exactly .
A rotation that appears from a composition of boosts is therefore the same phenomenon as a vector failing to return to itself after being carried round a loop on a sphere, with the sphere replaced by a hyperbolic space and the loop drawn in velocities rather than in positions.
One third of a gyroscope’s drift
The boundary this essay draws between the kinematic effect and the gravitational one is worth crossing briefly, because the two appear together in a measured number and the split between them is instructive.
A gyroscope in a circular orbit around a massive body precesses, relative to the distant stars, at a rate proportional to the mass and inversely to the orbital radius. That is the geodetic precession, and it was predicted in 1916.
Its usual decomposition has two pieces. One is exactly the Thomas precession of this essay: the gyroscope is being carried round a closed path in velocity space, and it comes back turned by per orbit, in the sense opposite to the motion. The other is a genuinely gravitational contribution from the curvature of space around the mass — the fact that the spatial geometry a gyroscope is parallel-transported through is not flat.
The two are in a ratio of one to minus four, and their sum is three halves of the naive orbital term. So the kinematic piece is a third of the total and points the other way, and neither piece alone predicts what is measured.
That decomposition is not gauge-invariant — how much is called “Thomas” and how much “curvature” depends on the coordinates — and the sum is. Which is a useful thing to know about the split: it is a way of understanding where the answer comes from, and not a division of the effect into two separately measurable parts.
The measurement was made by a satellite carrying four gyroscopes in polar orbit, comparing their axes against a guide star over a year. The predicted geodetic drift was about 6.6 arcseconds a year and the measurement agreed to a fraction of a per cent; a second, far smaller drift from the Earth’s rotation dragging spacetime with it was also detected, at about a tenth of an arcsecond a year.
So the effect on this page is not merely an ingredient in an atomic calculation from 1926. It is a third of a number measured in orbit, and the fact that the total came out right is evidence that the kinematic piece is there — since removing it would change the prediction by a third and the measurement rules that out.
Where the model stops
The precession formula assumes circular motion at constant speed. The general case — the Bargmann–Michel–Telegdi equation — carries the acceleration explicitly and reduces to per revolution only for a circle. For a general path there is no “per orbit” to speak of, and the accumulated rotation depends on the path in velocity space rather than on its endpoints.
The electron is treated as a classical orbiting body with a spin attached. It is not, and everything about the Bohr picture used to set this problem up is wrong in detail. What survives into the proper quantum treatment is the factor, which appears in the Dirac equation automatically — the equation contains the spin, the magnetic coupling and the Thomas factor together, with no separate step, which is one of the strongest arguments for it. The classical derivation is a way of understanding where the factor comes from, not the reason it is there.
And the rotation is kinematic, which is not the same as fictitious. It has no cause in the sense of a force, and it produces a measurable frequency shift in a real spectrum. Its status is exactly that of the composition law itself: not a dynamical effect, and not therefore an unreal one.
There is one more boundary worth naming. Everything here is flat spacetime — special relativity, no gravity anywhere — and the same word, precession, is used for a related effect that is not this one. A gyroscope in orbit around a massive body precesses by an amount that has a Thomas-like kinematic part and a genuinely gravitational part from the curvature of spacetime, and separating the two is the whole difficulty of measuring either. The kinematic part is on this page; the other belongs to the essays that treat gravity as geometry.
What the pictures cannot show
The Wigner figure draws an angle against an angle, and what is turning is a set of axes rather than an object. Nothing physical rotates in the sense a wheel rotates — what happens is that two observers who agree about a sequence of boosts disagree about the orientation of the axes at the end of it, and neither is wrong.
The thomas figure draws a turn per orbit, and an electron does not have an orbit. Every number on it is a classical stand-in for a quantum expectation value, and the reason the stand-in gets the right answer is that the effect is kinematic and cares only about the velocity, which the quantum treatment also has.
Where the ladder goes next
This ladder began with speeds that refuse to add, which is the one-dimensional law and the change of variable that tames it. This rung is what happens when the direction changes. The rungs after it: the general Wigner rotation for unequal, non-perpendicular boosts, of which the equal case drawn here is a slice; the aberration of starlight, which is the composition law applied to a direction rather than a magnitude; four-velocity, which composes by ordinary matrix multiplication and makes all of this bookkeeping rather than surprise; and the geometric reading, in which the rotation is a holonomy — the angle a vector is turned by parallel transport round a closed loop on a curved surface, the surface here being the hyperbolic space of velocities.
That last reading is the one worth carrying forward. A quantity that comes back changed after a round trip is measuring the curvature of the space it was carried through, and the space of velocities in relativity is not flat. Thomas precession is the first place in physics where that fact has a number attached to it, and the number was found because a spectral line refused to be twice as wide as it was.
Part 2 of 5
This essay is one argument about Velocity addition. The others:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.
Atomic spectraThe Lorentz transformationNon commutativityPrecessionRapidityReference framesRelativityRotationSpinVelocity addition
- The centre that is not a place the lorentz transformation, reference frames, spin
- Charge and current are one thing the lorentz transformation, rapidity
- Six numbers, one object the lorentz transformation, rapidity
- The collision that wastes most of the energy the lorentz transformation, reference frames
- The contraction no photograph shows the lorentz transformation, reference frames
- The diagram a ruler cannot read the lorentz transformation, reference frames