Waves

The ripple that counts the neighbours

An absorption edge is drawn as a step and it is a step with a ripple on it — a modulation of eleven per cent in copper, five in a zinc site buried in a protein. The ripple is the ejected electron's own wave, scattered back onto the atom that emitted it, so its period is a distance. It is the only way of measuring where an atom's neighbours are that does not need a crystal.

Assumes: The steps in an absorption curve · Where the loudness goes

Where absorption acquires thresholds takes the photon energy up until absorption acquires thresholds, and reads three things off a threshold: its height, which counts how many new absorbers became available; its position, which shifts by a few electronvolts with the oxidation state of the atom; and its use, which is that a threshold is the one feature in an attenuation curve a subtraction can isolate.

It stops at the edge, as the essay reading the same coefficient as a force stops at the coefficient’s mechanical consequences. Everything above it is drawn there as a smooth fall, because on an axis spanning three decades that is what it looks like.

Expand the axis and it is not smooth. For the first several hundred electronvolts above an edge the absorption oscillates, by a few per cent of itself, in a pattern that is perfectly reproducible and that has nothing to do with the absorbing atom.

The ripple that sits on top of an absorption edge. The Cu K absorption of copper foil, 293 K through its edge and for seven hundred electronvolts above it, with the smooth atomic background it would have if the absorbing atom were alone drawn beneath it. The difference between the two is the fine structure: a modulation reaching 11 per cent, dying away as the photon energy rises, and entirely absent from a free atom. It is there because the ejected electron is a wave that the neighbouring atoms scatter back onto the atom that emitted it, so the absorption depends on whether the returning wave arrives in step with the outgoing one — which depends on the distance to the neighbour and on nothing else about the sample.
Fig. 1 Copper’s K edge at 8,979 electronvolts, and seven hundred electronvolts above it, with the absorption a lone copper atom would have drawn dashed beneath. The difference between them reaches eleven per cent just above the edge and is still visible at the right-hand end. Nothing about the copper atom changed; what the dashed curve lacks is neighbours.

The absorption depends on what the electron finds when it comes back

The photoelectric absorption of an X-ray is a transition: a photon is destroyed, a core electron is ejected, and the probability of that happening depends on how well the initial state — an electron tightly bound near the nucleus — matches the final one.

The final state is an electron travelling outward — the same ejection the photoelectric threshold is about, seen from far enough above the threshold that the electron leaves with real speed. In a free atom that is a spherical wave going away and never returning, and the transition probability is a smooth function of energy, because nothing about the outgoing wave has any structure in it.

Put an atom twelve neighbours away and the outgoing wave meets them. Each neighbour scatters a little of it, and some of what it scatters comes straight back to the atom that emitted it. The final state is then not a spherical wave going away; it is a spherical wave going away plus a small returning wave, and the two add at the absorbing atom the way any two waves add anywhere else. One electron is interfering with itself, which is the same statement a single particle through two slits makes and is here being used as a ruler.

So the amplitude of the final state at the nucleus — which is what the transition probability depends on — is larger when the returning wave arrives in step and smaller when it arrives out of step.

Whether it arrives in step is decided by the phase it picked up going out and coming back, which for a neighbour at distance RR is 2kR2kR, where kk is the wavenumber of the electron’s own matter wave. Raise the photon energy and kk rises; the round-trip phase passes through multiple after multiple of 2π2\pi; the absorption goes up and down.

That is the whole mechanism, and its one remarkable feature is that a distance has been written into a curve of absorption against energy, with no image formed and no diffraction pattern recorded.

The natural variable is therefore not the photon energy but the electron’s wavenumber, which is the square root of the energy above the edge. Plotted against kk the modulation is periodic, and its period is π/R\pi/R.

The fine structure with the edge subtracted away. The modulation alone, weighted by the square of the photoelectron's wavenumber and plotted against that wavenumber rather than against photon energy — which is the variable the interference is periodic in, because the phase accumulated on a round trip to a neighbour at distance R and back is 2kR. The sample drawn is a zinc site in a protein, which has one shell: 4 O at 1.98 Å. One distance gives one frequency, and the envelope falls because the neighbour's ability to scatter the electron back does. The weighting by k² is there because the raw modulation dies away steeply; it makes the far end of the range visible without changing where any of the zeros are.
Fig. 2 The modulation alone for the simplest case there is: a zinc atom with four oxygens at 1.98 ångströms and nothing else close enough to matter, which is the arrangement at a metal site in many proteins. One distance gives one frequency. The envelope dies away because oxygen is a light atom and stops sending the electron back by about eight inverse ångströms.

Why the period is a distance and the amplitude is a count

Writing the modulation as χ\chi, the fractional change in absorption, the single-scattering expression is a sum over shells of neighbours:

χ(k)=jNjS02Fj(k)kRj2e2σj2k2e2Rj/λ(k)sin ⁣(2kRj+ϕj(k)).\chi(k) = \sum_j \frac{N_j S_0^2 F_j(k)}{k R_j^2}\,e^{-2\sigma_j^2 k^2}\,e^{-2R_j/\lambda(k)}\,\sin\!\left(2kR_j + \phi_j(k)\right).

Every factor in it is doing a separate and legible job, which is unusual for an expression of that length.

The sine carries the distance. Its argument is the round-trip phase, so the frequency of the oscillation in kk is 2R2R and the period is π/R\pi/R. For copper’s first shell at 2.556 ångströms that is 1.23 inverse ångströms; for the second at 3.615 it is 0.87. Two shells therefore beat against one another with a period of 2.97, which is the slow envelope visible in the copper spectrum and is the difference of the two distances read directly off the picture.

NjN_j carries the count. Twice as many neighbours at the same distance give twice the modulation, so the amplitude of a frequency component is how many atoms are at that distance. Coordination numbers come out of this measurement, and they are the reason it is used on catalysts, where the question is usually how many atoms a metal particle’s surface site has left.

1/R21/R^2 is the inverse-square law, arriving here as the dilution of a spherical wave on its way out, exactly as it arrives in every other spreading problem. A neighbour twice as far away is reached by a quarter as much wave.

Fj(k)F_j(k) is how well that neighbour scatters, and it is the part that says which element the neighbour is. A light atom has a backscattering amplitude concentrated at low kk; a heavy one scatters right across the range and has structure in it. Two shells at the same distance made of different elements are distinguishable for that reason alone, which nothing about a diffraction pattern reports at all.

S02S_0^2 is an embarrassment, and it is worth naming rather than hiding. It is an overall reduction factor, between about 0.7 and 1.0, that accounts for the other electrons of the absorbing atom relaxing while the core hole is made — the final state is not quite the state the one-electron picture assumes. It is usually fitted rather than computed, and it multiplies the coordination number, so it is exactly the parameter that makes a coordination number the least reliable thing the technique reports.

The fine structure with the edge subtracted away. The modulation alone, weighted by the square of the photoelectron's wavenumber and plotted against that wavenumber rather than against photon energy — which is the variable the interference is periodic in, because the phase accumulated on a round trip to a neighbour at distance R and back is 2kR. The sample drawn is copper foil, 293 K, which has 2 shells: 12 Cu at 2.56 Å, 6 Cu at 3.61 Å. The sample drawn is copper foil, 20 K, which has 2 shells: 12 Cu at 2.56 Å, 6 Cu at 3.61 Å. Two distances beat against one another, and the beat period is the difference between them. The weighting by k² is there because the raw modulation dies away steeply; it makes the far end of the range visible without changing where any of the zeros are.
Fig. 3 Copper foil at room temperature and at twenty kelvin, the same two shells in both. The fast oscillation is the first shell and the slow modulation of its amplitude is the beat against the second. Cooling changes nothing about where the atoms are and a great deal about how much signal survives at high wavenumber: at k=12k = 12 the first shell retains 8.9 per cent of its amplitude at 293 kelvin and 44.6 at twenty, a factor of five for a change that costs only liquid helium.

The two exponentials are why the measurement is local and why it is made cold

The two damping terms decide, between them, everything about what the technique can and cannot see.

e2σ2k2e^{-2\sigma^2 k^2} is a spread in the distance. The neighbours are not at RR; they are at RR give or take a vibration, and σ2\sigma^2 is the mean-square spread of that distance. Waves returning from a distribution of distances arrive with a distribution of phases and partially cancel, and because the phase is proportional to kk the cancellation worsens as the square of kk. The functional form is the same Gaussian that damps a Bragg peak with temperature, for the same reason.

Two things are folded together in σ2\sigma^2 and the measurement cannot separate them: the thermal vibration, which can be removed by cooling, and the static disorder, which cannot. A cold spectrum that is still damped is reporting a genuine spread of bond lengths, and that is often the point of taking it.

e2R/λ(k)e^{-2R/\lambda(k)} is whether the electron survives the trip. An electron travelling through matter loses energy to plasmons and to other electrons, and an electron that has lost energy has the wrong wavenumber and contributes nothing but background. The distance it travels before that happens is its inelastic mean free path, and it is small.

How far a photoelectron can report from. The distance a photoelectron travels before an inelastic collision destroys its phase, against its wavenumber, with the round-trip distances to three coordination shells marked across it. The curve has a minimum of about four and a half ångströms near four inverse ångströms — slow electrons excite plasmons, fast ones have less time to — and it is that minimum, not anything about the sample, that decides how much of a structure the measurement can see. A neighbour at two ångströms means a round trip of four, which survives across the whole range. One at six means a round trip of twelve, and at the minimum of the curve fewer than one electron in ten arrives with its phase intact. That is why the technique reports the first two or three shells of any material and is silent about the rest, and why it works at all on a sample that has no long-range order to diffract from.
Fig. 4 The photoelectron’s mean free path against its wavenumber, with the round-trip distances to three shells drawn across it. The curve has a minimum near four inverse ångströms — slow electrons excite plasmons efficiently, fast ones have less time in which to — and the minimum is about four and a half ångströms. A first shell at two ångströms means a round trip of four, and four in ten electrons survive it. A shell at six ångströms means a round trip of twelve, and seven in a hundred do.

This is the most consequential single number in the subject and it is nothing to do with the sample. The mean free path is a property of an electron moving through condensed matter of any kind, it has a minimum of four or five ångströms somewhere in the middle of the useful range, and it is what makes the measurement local.

Locality is usually a limitation. Here it is the reason the technique exists.

A diffraction pattern is a sum over an entire crystal, and a sum over an entire crystal requires an entire crystal. A glass has none, a solution has none, and a single metal atom held at the active site of an enzyme has none. All three are transparent to diffraction and all three give a perfectly good absorption spectrum, because a measurement that cannot see past five ångströms does not care whether anything beyond five ångströms is ordered.

The surprising consequence is that the technique is element-selective by construction. The edge belongs to one element, so the ripple on it reports only on that element’s own surroundings — one iron atom in ten thousand atoms of protein, at a concentration where no other structural method has anything to say.

Reading the distances off, and the half-ångström nobody should trust

A sum of sinusoids in kk whose frequencies are distances is exactly the shape a Fourier transform was invented for, and taking one turns the spectrum into a picture of the neighbourhood. It is the same manoeuvre that turns an interferogram into a spectrum, run in the other direction.

The neighbours, read off a transform. The magnitude of the Fourier transform of the weighted modulation, against distance. Each coordination shell appears as a peak, and the area under a peak carries how many atoms are in that shell. The dashed lines mark where the atoms actually are: every peak sits about half an ångström SHORT of its shell, because the electron's phase is shifted by the potentials of the atom it left and the atom it bounced off, and that shift is very nearly linear in k. Correcting for it is the whole of turning this picture into a distance, and the correction is computed rather than fitted. Beyond about five ångströms there is nothing, whatever the sample: an electron that has travelled further than its own mean free path has lost its energy and cannot contribute.
Fig. 5 The transform of the copper spectrum against distance. Two peaks, for the two shells, with the true distances marked. Both peaks sit short — 2.56 ångströms of copper appears at about 2.1 — and the shortfall is not noise, not resolution and not a mistake. It is the phase shift, and it is present in every such transform ever taken.

The peaks are in the wrong place, always, by roughly half an ångström. The reason is the last term in the expression, ϕj(k)\phi_j(k), which has been ignored so far.

An electron leaving a copper atom does not leave a bare Coulomb potential; it climbs out of the atom’s own screened potential and its phase is advanced. When it reaches a neighbour it does not bounce off a hard sphere; it penetrates that atom’s electron cloud, and its phase is advanced again. The total shift is ϕ(k)\phi(k), it is roughly linear in kk, and a term linear in kk added to 2kR2kR is indistinguishable, inside the sine, from a change in RR.

So the transform’s peak lands at Rϕ/2R - \phi'/2 rather than at RR. The offset is around a quarter to half an ångström for most combinations, and since a bond length is two or three ångströms it is an error of fifteen to twenty per cent — enormous by the standards of what the measurement is otherwise capable of, which is a hundredth of an ångström.

The repair is not a calibration. The phase shifts are computed, from the atomic potentials, by the same codes that compute the backscattering amplitudes; the computed phase is put into the fit and the distance that comes out is the real one. That is why an analysis is a fit to a theoretical spectrum rather than a reading off a transform, and it is why the transform is best understood as a picture of the neighbourhood rather than a measurement of it.

The neighbours, read off a transform. The magnitude of the Fourier transform of the weighted modulation, against distance. Each coordination shell appears as a peak, and the area under a peak carries how many atoms are in that shell. The dashed lines mark where the atoms actually are: every peak sits about half an ångström SHORT of its shell, because the electron's phase is shifted by the potentials of the atom it left and the atom it bounced off, and that shift is very nearly linear in k. Correcting for it is the whole of turning this picture into a distance, and the correction is computed rather than fitted. Beyond about five ångströms there is nothing, whatever the sample: an electron that has travelled further than its own mean free path has lost its energy and cannot contribute.
Fig. 6 A zinc site with two oxygens at 1.98 ångströms and two sulphurs at 2.31 — a third of an ångström apart, from a spectrum whose usable range is thirteen inverse ångströms. The two shells are not resolved into two peaks. What separates them is not the picture but the fit: sulphur’s backscattering amplitude peaks at a higher wavenumber than oxygen’s and its phase shift is different, so the two contribute differently at different kk even where their peaks overlap.

The resolution of the transform is the usual one for a truncated Fourier measurement — the same trade between the length of a record and the sharpness of what can be got out of it that decides how finely an interferometer separates two lines — and it is π/2Δk\pi/2\Delta k, so a range of thirteen inverse ångströms resolves about 0.12 ångströms. That is what separates two peaks in the picture. Two shells closer together than that are not resolved in the picture and can still be separated in a fit, because the fit uses the amplitude and phase functions as well as the frequency — three pieces of evidence rather than one.

Claiming two shells a twentieth of an ångström apart from a fit is a different kind of statement from reading two peaks off a transform, and the difference between those two statements is most of what is argued about in the literature of the technique.

Forty years of a signal nobody could read

The oscillations were not discovered by anybody looking for a way to measure bond lengths. They were reported in 1931, within a few years of the first decent X-ray spectrometers, as an unexplained wiggle above an edge, and the argument about what caused them ran for four decades.

The argument was between short range and long range. One account held that the structure came from the electron being scattered by the crystal as a whole — from the band structure of the final state, which is a property of the lattice — and predicted that the oscillations should depend on crystal orientation and vanish in a liquid. The other held that it came from the immediate neighbours and should survive melting. Measurements were made of both kinds, they were hard, and they were not decisive.

What settled it was not a better measurement. It was a change of variable. In 1971 the spectrum was replotted against the electron’s wavenumber rather than against photon energy, and its Fourier transform was taken — and peaks appeared at the known interatomic distances of the sample. A sum of sinusoids whose frequencies are distances had been sitting in the data since 1931, and it is not visible against photon energy because the frequency drifts continuously as the square root.

That is worth pausing on, because it is not a story about instrumentation. The forty years were spent arguing about a mechanism using data that contained the answer, in a coordinate system that hid it. The transform did not add information; it made the information that was already there legible, and the technique went from a curiosity to a standard method within a decade of somebody plotting it differently.

There is a second reason it took so long, and it is the ordinary one: a spectrum over a usable range took a day on a laboratory X-ray tube and takes a second on a storage ring. A method that needs an argument settled before it is worth using, and a source that does not exist yet, waits for both.

The dilute case, which is what it is mostly used for

Everything so far has been copper metal: twelve identical neighbours, a strong signal, a spectrum that can be measured in minutes.

The ripple that sits on top of an absorption edge. The Zn K absorption of a zinc site in a protein through its edge and for seven hundred electronvolts above it, with the smooth atomic background it would have if the absorbing atom were alone drawn beneath it. The difference between the two is the fine structure: a modulation reaching 5 per cent, dying away as the photon energy rises, and entirely absent from a free atom. It is there because the ejected electron is a wave that the neighbouring atoms scatter back onto the atom that emitted it, so the absorption depends on whether the returning wave arrives in step with the outgoing one — which depends on the distance to the neighbour and on nothing else about the sample.
Fig. 7 The same construction for a zinc atom with four oxygen neighbours. The modulation reaches five per cent rather than eleven, because there are four neighbours rather than twelve and oxygen is a poorer backscatterer than copper. A measurement has to resolve that five per cent against the absorption of everything else in the sample, and in a protein the zinc may be one atom in ten thousand.

That ratio is the practical problem the whole apparatus of the field is built around. At a concentration of one in ten thousand, the zinc’s own absorption is a small fraction of the total, and the ripple is five per cent of that small fraction — so the quantity being measured is a few parts in a hundred thousand of what the detector sees.

It is measured anyway, by not measuring the transmitted beam at all. The atom that absorbed the photon is left with a core hole, the hole is filled, and a fluorescent photon comes out at an energy characteristic of that element and no other. Counting those photons counts absorption events in the chosen element only, and the enormous background from everything else in the sample never enters the count. What limits the measurement then is photon statistics and the brightness of the source, which is why the technique was a curiosity in the 1970s and became routine when storage rings were built.

Single scattering, parameterised amplitudes, and a symmetric spread

The expression above is single scattering. It assumes the electron goes out, bounces once and comes back. It can also bounce off two atoms in succession, and when three atoms lie nearly in a line the forward scattering off the middle one is strong enough that the multiple path dominates the single one. That is a nuisance if distances are wanted and a gift if angles are: the intensity of a multiple-scattering path is a measure of a bond angle, which nothing in the single-scattering expression can supply.

The backscattering amplitudes and phase shifts drawn here are smooth parameterised functions, not the tabulated atomic calculations a real analysis uses. Their shapes are right — the amplitude’s peak moves to higher wavenumber with atomic number, the phase is nearly linear and puts every peak short — and their values are not. A spectrum computed from them would not fit a measurement, and no number here should be read as one.

The Gaussian spread assumes a symmetric distribution of distances. A real bond length distribution is skewed, because the potential is anharmonic — which is the same anharmonicity that makes a solid expand when heated — and the skew biases the fitted distance downward with temperature. Treating σ2\sigma^2 as a single number is an approximation that fails first, and exactly where thermal expansion is largest.

And the first thirty or so electronvolts above the edge are not described by any of this. There the electron is slow, its mean free path is long, it scatters many times, and the structure that results is a different measurement with a different name reporting on symmetry and oxidation state rather than on distances. The expression here applies above about thirty electronvolts and below about a kilo­electronvolt, and the two ends are both soft.

The noise, the phase thrown away, and the hole that is still there

They cannot show the noise. Every curve drawn here is a computed function evaluated exactly, and a real spectrum at k=13k = 13 is a few counts above a background, so the high-wavenumber end of a measurement looks nothing like the clean oscillation drawn. The range a spectrum is usable over is decided by where the noise overtakes the signal, and that boundary is invisible in a calculation.

Nor can they show why a transform peak is asymmetric. The transform drawn is the magnitude only, which is what is conventionally plotted; the transform is complex, its real and imaginary parts oscillate underneath the magnitude, and a fit uses both. A magnitude plot discards the phase information that does most of the work, which is one more reason the picture is not the measurement.

And they cannot show the core hole. The whole process happens while the absorbing atom has a hole in its innermost shell, which lives for about a femtosecond and therefore has a width of an electronvolt. That width smears the spectrum in energy, it is what ultimately limits how far out in kk any measurement can reach even with a perfect source, and it is set by the atom rather than by the apparatus.

Still open: how much of a structure a local probe can be made to report

The measurement’s reach is set by the mean free path, and the mean free path is not adjustable. What is adjustable is how much is extracted from what does come back.

Multiple-scattering paths reach further than single ones, because a path that visits three atoms samples geometry that no pair does, and modern analysis fits dozens of such paths rather than two or three shells. How far that can be pushed — whether a spectrum of finite range and real noise contains enough independent information to determine a structure of several shells rather than to check one already proposed — is a question about information content rather than about physics, and the honest answer depends on the sample. The usual estimate of how many independent parameters a spectrum can support is 2ΔkΔR/π2\Delta k\Delta R/\pi, which for a good measurement over four ångströms of range is about thirty, and a structure of dozens of paths has more parameters than that.

That estimate is itself worth treating carefully, because it is the Nyquist count for a band-limited signal and a spectrum is not one. It assumes the noise is uniform across the range, and it is not — the signal is dying while the noise is not, so the far end of the range carries fewer independent numbers than its width suggests. It also assumes the parameters are independent of one another, and the coordination number and the disorder term are notoriously not: raising one and lowering the other leaves the amplitude nearly unchanged, which is why those two are the pair that fits report as strongly correlated. A structure determined with thirty nominal degrees of freedom may be determined with rather fewer real ones.

The direction taken instead is to bring outside information in. A fit that is constrained by a molecular-dynamics simulation, or by a plausible chemical geometry, or by spectra taken at several temperatures and fitted together, has more evidence than one spectrum and correspondingly more to say. Whether that is a determination of a structure or a test of one somebody already believed is a question about the constraints rather than about the data, and it is the honest place to leave it.

The habit worth carrying away is about what a small oscillation on a large background is worth. A modulation of a few per cent is not a correction to the thing it sits on; it can carry information the large term does not contain at all. The edge height counts electrons and the ripple on it locates atoms, and the second is a thousand times smaller than the first. The same reading applies wherever a measurement is dominated by a term that varies smoothly: the smooth part is usually the least interesting thing in it, and the part worth having is what is left after the smooth part is subtracted away.

Part 5 of 6

This essay is one argument about Attenuation. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

AbsorptionAttenuationBinding energyCoordination numberDebye waller factorFourier transformInterferenceLocal structureMatter wavesMean free pathPhase shiftX-rays