Quantum

The correlation no instructions can produce

A pair of gloves in two boxes agrees perfectly and needs no physics, because the answers were settled at packing. What no packing can imitate is the shape that appears as the two analysers are turned relative to each other, and the shape is a number — 2.828 where every list of pre-agreed answers is stuck at 2.

Assumes: The answer that was not there before · Sharpness has to be paid for

Two particles are prepared together and sent in opposite directions. Each meets an analyser with a dial on it, and each analyser reports one of two outcomes. Nothing in that description is unusual. What is unusual is what happens to the agreement between the two reports as the dials are turned.

The correlation, and the best a shared list of answers can do. The coincidence correlation between two polarisation analysers against the angle between them, over two full turns of the correlation — a polariser turned through 180° is the same polariser, so the picture repeats. The singlet gives −cos 2Δ, drawn through −1.00 at 0°, 1.00 at 90°, −1.00 at 180°, 1.00 at 270°. Beside it is the best correlation any shared list of pre-agreed answers can produce: straight lines between the same four extremes, with corners where the cosine is smooth. The two agree exactly at the multiples of 45° and nowhere else, and they are furthest apart — by 0.2105 — at 19.77° and 70.23°, which is ½ arcsin(2/π) from either end of the quarter turn. The difference is not a matter of degree: it is a curve against a shape with a corner in it, and no list can be bent into the curve.
Fig. 1 The coincidence correlation against the angle between the two analysers, for a polarisation-entangled pair. The smooth curve is −cos 2Δ; the straight lines beside it are the best any list of answers agreed before the pair separated can manage. The two meet exactly at the multiples of 45° and are furthest apart by 0.2105, at 19.77° and 70.23°. One is a curve, the other has corners, and that difference is the whole of what follows.

Everything below is where those two shapes come from, what the difference between them measures, and what four decades of experiment had to remove before the measurement meant anything.

Two gloves in two boxes

Entanglement is usually introduced as a correlation so strong it looks like telepathy, and that is not what is strange about it.

Take a pair of gloves, put one in each of two boxes, and post the boxes to opposite ends of the country. Open one and find a left glove: the other box holds a right glove, with certainty, instantly, at any separation. No physics is involved and nothing travels. The two answers were settled when the boxes were packed, and opening a box reveals a fact rather than creating one.

A singlet pair, measured with both analysers set to the same angle, behaves exactly like the gloves. The two outcomes are always opposite. Written as ±1 apiece, the average of their product is −1, and a perfect anticorrelation is the least mysterious statistic in physics.

The correlation between two analysers, against the angle between them. The coincidence correlation between two polarisation analysers against the angle between them, over two full turns of the correlation — a polariser turned through 180° is the same polariser, so the picture repeats. The singlet gives −cos 2Δ, drawn through −1.00 at 0°, 1.00 at 90°, −1.00 at 180°, 1.00 at 270°. Nothing about the curve is unusual on its own — a correlation of −1 at equal settings is what any pair of matched objects gives. What no list of instructions can produce is its shape.
Fig. 2 The same correlation with the comparison taken away, which is how the result is usually first shown. Nothing about the curve is remarkable on its own: it passes through −1.00 at 0°, +1.00 at 90°, −1.00 at 180° and +1.00 at 270°, and a pair of matched objects gives every one of those values. The strangeness is in the shape between the marked points, and this drawing does not contain enough to see it.

So the interesting question is not how strong the agreement is at one setting, but what the whole family of agreements looks like across all settings — because a list packed into the boxes must answer every question either analyser might ask, and must do so before the boxes part company.

What turning the analysers does

The convention matters here and is worth stating before any number is quoted, because it is where an argument about entanglement most often goes quietly wrong.

The figures use polarisation-entangled photons, since that is what every experiment since 1972 has actually used. A polarising analyser turned through an angle α turns the measurement direction it defines by 2α, so the coincidence correlation for analysers at a and b is

E(a,b)=cos(2(ab)),E(a, b) = -\cos\bigl(2(a - b)\bigr),

with a and b the physical dial readings. It is still minus the cosine of the angle between the two measurement directions, exactly as for a pair of spins; only the gearing is 2:1. Two consequences follow. A polariser has a period of 180°, so the picture above repeats twice across a full turn of the dial. And the settings that matter most sit at 22.5° rather than at 45° — a fact the site’s generator enforces by refusing the spin-½ set 0°, 90°, 45°, 135°, which under this convention gives a combination of exactly zero and would draw a figure whose whole subject was a violation of nothing.

Now the packed list. Suppose each photon carries a complete set of answers, one for every angle its analyser might be set to, fixed at the source. The most economical version of that is a single hidden angle λ drawn afresh for each pair, with the instruction: report plus if the analyser’s measurement direction lies in the half turn beginning at λ, and minus otherwise, with one side’s answer inverted so that equal settings always disagree.

That model reproduces the perfect anticorrelation at Δ = 0. It gives each analyser, taken alone, a fair coin. And its correlation rises in a straight line from −1 at equal settings to +1 a quarter turn away, because the two photons disagree exactly when λ falls in the wedge between their measurement directions, and that wedge grows in proportion to the angle. The result is the sawtooth in the hero figure: exact at 0°, 45°, 90° and every multiple of 45°, too flat everywhere in between.

It has corners, and it must. Any pre-agreed list is a fixed assignment averaged over some distribution, so the correlation it produces is an average of straight-line pieces pinned to the same fixed points. Straight lines between fixed points cannot be bent into a cosine. The widest disagreement is 0.2105, and the generator locates it by search at 19.77° and 70.23°, then checks that against the closed form: the sawtooth’s slope matches the cosine’s where sin2Δ=2/π\sin 2\Delta = 2/\pi, so the extremum sits at 12arcsin(2/π)\tfrac12\arcsin(2/\pi) from either end of the quarter turn.

That angle is worth separating from the famous one. The place where two correlations differ most is not the place where a four-term combination of them peaks, and the second is what an experiment is built around; conflating the two is common, and the figures keep them apart.

Malus's law. The fraction of polarised light passing a filter, against the angle between the light's own direction of shaking and the filter's axis. It is the cosine squared: half at 45 degrees, nothing at 90.
Fig. 3 Where the factor of two comes from, on a bench. A single filter passes the cosine squared of the angle between the light’s own direction of shaking and its axis: 100% at 0°, 75% at 30°, 50% at 45° and nothing at 90°. Squaring a cosine turns it into a cosine of the doubled angle, which is why a direction of shaking has period 180° while the correlation built from it has period 90°.

Four settings, and a number an experiment counts

A curve against a shape with corners is not yet a measurement. Turning it into one takes four numbers, and the recipe is Clauser, Horne, Shimony and Holt’s, from 1969.

Give the first analyser two settings a and a′, the second two settings b and b′, run the pair at all four combinations, and form

S=E(a,b)E(a,b)+E(a,b)+E(a,b).S = E(a,b) - E(a,b') + E(a',b) + E(a',b').

For any assignment of pre-agreed answers whatever, three of those four terms can be made to agree and the fourth cannot, and the algebra bounds S|S| by 2. That is Bell’s theorem in the form an experimentalist can use: not a statement about one correlation but about a sum of four, holding for every distribution over every list of instructions, including lists nobody has thought of.

The singlet reaches 2√2 = 2.82843, and the excess is 0.8284.

The CHSH combination, against what any instruction list can reach. The CHSH combination |S| for analysers set to 0°, θ, 2θ and 3θ, plotted against θ. For the singlet it rises from 2 to a maximum of 2.82843 — 2√2, Tsirelson's bound — at θ = 22.500°, which is the setting every Bell experiment is built around, and it stays above 2 for every θ up to 34.26°. The straight lines are the same combination for the best shared instruction list: exactly 2 while 3θ is still inside the first quarter turn, then falling away. It never exceeds 2 anywhere on the sweep, and no list of pre-agreed answers can — that is Bell's inequality. The gap at the optimum is 0.8284, which is what an experiment measures.
Fig. 4 The combination for the one-parameter chain of settings 0°, θ, 2θ, 3θ, plotted against the step θ. The quantum curve leaves 2 immediately, peaks at 2.82843 — Tsirelson’s bound — at θ = 22.500°, and stays above 2 for every θ below 34.26°. The straight lines are the same combination for the best shared list, and they do not fall short of the classical bound: they sit on it, then fall away.

That the local model saturates the bound is the sharpest way to hold the result. A weaker model — one that gave 1.8 — would make the comparison look like a matter of quality, as though better instructions might close the gap. The list drawn here is the best there is: it matches the state’s perfect anticorrelation, matches every one-sided statistic, and reaches exactly 2. There is nothing left to improve, and the cosine is still 0.83 above it.

The experiment is then a tally. Pairs arrive, the two dials are set, two outcomes are recorded, and the four correlations are estimated as running averages of ±1 products. The figures below run that tally against both models side by side, on the same pairs, with each estimate’s own standard error drawn as a band.

The experiment run, and where its estimate settles. The CHSH combination estimated from 5,000 simulated coincidences at the four settings 0°, 45°, 22.5°, 67.5°, plotted against the number of pairs collected so far with the shaded band its own standard error. The estimate settles on 2.8528 ± 0.0396, which is 0.6 standard errors from 2√2 = 2.82843 and 22 above the 2 that no local theory can pass. The lower curve is a shared instruction list — one hidden angle per pair, deterministic answers, sampled the same number of times at the same settings — and it gives 2.0304 ± 0.0487, on the classical bound rather than below it, because this is the best list there is. The two are 13 combined standard errors apart. Both models give each analyser a "+" half the time; the difference is only in the coincidences.
Fig. 5 Five thousand simulated coincidences, 1,250 at each of the four settings. The quantum estimate settles at 2.8528 ± 0.0396, which is 0.6 standard errors from 2√2 and 22 above the classical 2; the shared list gives 2.0304 ± 0.0487. The two are 13 combined standard errors apart, and the left-hand end of the plot shows what a few hundred pairs are worth — the estimate wanders across half a unit before the counting settles it.
The experiment run, and where its estimate settles. The CHSH combination estimated from 100,000 simulated coincidences at the four settings 0°, 45°, 22.5°, 67.5°, plotted against the number of pairs collected so far with the shaded band its own standard error. The estimate settles on 2.8202 ± 0.0090, which is 0.9 standard errors from 2√2 = 2.82843 and 91 above the 2 that no local theory can pass. The lower curve is a shared instruction list — one hidden angle per pair, deterministic answers, sampled the same number of times at the same settings — and it gives 2.0048 ± 0.0109, on the classical bound rather than below it, because this is the best list there is. The two are 58 combined standard errors apart. Both models give each analyser a "+" half the time; the difference is only in the coincidences.
Fig. 6 The same run continued to a hundred thousand pairs, which is why the two figures share their left halves exactly. Twenty times the data halves the error bar twice over: 2.8202 ± 0.0090, now 91 standard errors above 2, against the list’s 2.0048 ± 0.0109. Neither curve moves anywhere new: what accumulates is not the effect but the confidence in it.

The error bars are not decoration. They are computed from the run’s own counts — each estimate is a mean of ±1 values, so its variance is (1E2)/n(1-E^2)/n — which is exactly the error bar an experiment would quote about itself, without reference to the answer it was expecting.

The strangeness that belongs to one particle already

Entanglement is routinely explained with a phenomenon that has nothing to do with two particles, and the confusion is worth clearing out, because this site has already drawn that phenomenon — in a chain of analysers, and in dots behind two slits.

The strangeness that belongs to one particle is worth separating from the strangeness that needs two, because they are usually run together. A single analyser returns one of a countable set of answers, and asking a second question afterwards destroys the settled value of the first. That is odd, it is well established, and it is not what Bell’s theorem is about — a local hidden-variable model reproduces all of it without difficulty.

What no such model reproduces is the correlation between two distant analysers as their angles are varied together. The single-particle oddness is a statement about measurement; Bell’s is a statement about what could have been arranged in advance. Keeping them apart is what makes the theorem’s conclusion narrow enough to be worth having.

3 filters at 0°, 45°, 90°: 12.5% gets through. Unpolarised light passing through 3 polarising filters with axes at 0 degrees, 45 degrees, 90 degrees. The first removes half whatever its angle; each one after it passes the cosine squared of the turn from the filter before. 12.5 per cent of the original intensity survives.
Fig. 7 The one-particle result that a bench can show in ten seconds. Two crossed filters pass nothing; slide a third in between them at 45° and 12.5% of the original light comes through. Adding a filter increases the transmission, which is impossible if a filter only removes — and it is a fact about single photons passing through glass, not about any pair.

Both of those are genuinely strange, and neither of them is Bell’s result. A chain of analysers, or a stack of filters, can be reproduced by a model in which each particle carries a list of answers and the middle element rewrites the list. What such a model cannot reproduce is the shape drawn at the top of this page. The one-particle demonstrations show that a measured value need not pre-exist; the two-particle result shows that no pre-existing values, however assigned, can produce the numbers — a strictly stronger statement, and one that needs the second particle.

What the result does not say

Three readings overshoot, and the reasoning against each is short.

No signal is sent. Each analyser on its own reports plus half the time and minus half the time, whatever the distant dial is set to. The figures assert this rather than assuming it: at all four settings, in both models, a run is rejected if either analyser’s rate of plus outcomes strays more than four standard errors from one half. The correlation lives in the joint record, and a joint record exists only once the two sets of results are brought to the same place over an ordinary classical channel, at the speed of light or slower. Without it, the distant record is indistinguishable from coin flips.

Measuring one particle does not change the other. There is no experiment that reads a difference at one analyser according to what was done at the other. What the measurement changes is the description available to somebody who holds both records, and a description is not a thing at a place. This is also why the result puts no strain on the slicing of now: the two measurements are spacelike separated, so their order depends on the frame, and any account in which one caused the other would have to choose a frame no experiment can identify.

What is ruled out is an assumption, not a picture. The assumption is that the outcomes were determined in advance by properties the particles carried, with each outcome depending only on its own particle and its own analyser. That is a specific, reasonable, and until 1964 nearly universal supposition, and it is what the measurement kills.

Not as non-local as it could be

Here is the strangest fact on this page, and it is not the violation.

Quantum mechanics does not go as far as it is allowed to. Tsirelson showed in 1980 that no quantum state of any dimension can push the combination past 2√2, so the value the figures reach is a ceiling and not a coincidence. But 2√2 is not the ceiling that causality imposes. Popescu and Rohrlich asked in 1994 what the no-signalling condition alone permits, and the answer is 4 — the logical maximum, where the four correlations are perfect in the three places the sum adds them and perfect in the one place it subtracts. A hypothetical device with that behaviour, since christened a PR box, has flat marginals at both ends and transmits nothing. It obeys relativity as strictly as a pair of photons does.

So there is room between what quantum mechanics does and what causality forbids, and quantum mechanics sits well inside it, at 2.828 out of 4. Nobody knows what principle picks out that number. Candidates exist — information causality, macroscopic locality — and each recovers 2√2 from an assumption that looks reasonable and has no independent standing.

A theory being too well behaved is a stranger fact than a theory misbehaving. The usual story is that quantum mechanics broke a bound that classical physics had assumed; the harder question is why it stopped where it did, when stopping was not required.

The same combination is a game, which is where the number turns up again. Two players are separated, each handed a random bit, and each must answer with a bit; they win if their answers differ exactly when both input bits are 1. With any strategy agreed in advance, including any shared random data, the best possible win rate is 75%. Sharing an entangled pair and measuring it at the angles above raises it to cos²22.5° = 85.36%, which is the number the second analyser chain prints as 0.427 out of 0.500. Nothing in the game mentions physics, and the bound on it is the same bound.

What it costs, and how each loophole was shut

An inequality is only violated once every ordinary way of faking the result has been removed, and each removal cost something.

Detection efficiency is a threshold, not a quality. If detectors miss most pairs, the recorded sample need not represent the whole, and a local model that decides when to be detected can fake any correlation. Closing this needs efficiency above 2(√2 − 1) = 82.8% for a maximally entangled pair — a hard number, not a target — and Eberhard showed in 1993 that a deliberately unbalanced state lowers it to 66.7% at the price of a smaller violation. Delft’s 2015 experiment paid differently, entangling two electron spins in diamond 1.3 km apart and reading them out electrically at near-unit efficiency; 245 trials gave 2.42 ± 0.20.

Speed sets the geometry. The locality loophole is the possibility that one analyser’s setting reached the other end in time to matter, and closing it means choosing the setting and finishing the measurement inside the light-crossing time. Aspect, Dalibard and Roger did it in 1982 by switching the light between two analysers with acousto-optic deflectors every 10 ns, over a 12 m baseline that light needs 40 ns to cross. Every subsequent experiment has had to buy nanoseconds in the same way.

Freedom of choice is bought from somewhere far away. The remaining gap is that the settings themselves might correlate with the source. In 2018 two groups attacked it from opposite directions: one took its settings from the colour of light emitted by quasars 7.8 billion years ago, so any common cause would predate the Earth; the other from about a hundred thousand people pressing keys in a video game across thirteen laboratories, tens of millions of bits of human whim. Neither closes the hole; both push its lid further out.

The property became a product. A violation certified without trusting the apparatus is a supply of randomness no adversary can have precomputed, and NIST ran a Bell test in 2018 as a randomness beacon. The same argument underwrites device-independent key distribution, whose security follows from the measured S rather than from any claim about the boxes — the commercial descendant of a state that cannot be read without disturbing it.

The 2022 Nobel Prize went to Clauser, Aspect and Zeilinger, which is best read as the close of the argument rather than its opening: the prize was for having removed the alternatives one at a time over fifty years.

Where the model stops

Every figure here assumes perfect analysers, perfect detection and exactly the settings written on the axis, and each of those is an idealisation with a measurable cost.

Visibility is the whole budget. A real source produces a state slightly mixed, and the combination scales with the visibility V as 22V2\sqrt2 \, V. Violation therefore needs V above 1/2=0.70711/\sqrt2 = 0.7071, which sounds generous until the other losses are counted; good modern sources run near 0.97, giving 2.744 rather than 2.828. The loophole-free runs came in lower still, near 2.4, because closing a loophole costs visibility.

A sampled run is a statistical statement. “The bound was broken” is the wrong sentence. The right one names the separation: at 5,000 pairs the estimate is 2.8528 ± 0.0396, which is 22 standard errors above 2, and at 100,000 it is 91. Both are enormous; neither is infinite, and the number of standard errors is the claim.

A tally is not the same as a fact, and it is worth saying so about the numbers above. The same double slit shows nothing after twenty arrivals and unmistakable fringes after a thousand; a Bell measurement is a statement of exactly that kind, accumulated over many runs. That is what the error bars are for, and it is why the experiments are quoted in standard deviations rather than as a single decisive event.

Freedom of choice is an assumption about the world. Every other loophole concerns the apparatus and can be engineered shut. This one concerns whether the settings are independent of the source’s state, and a determined superdeterminism can always place the correlation earlier. Quasars and video games push the required conspiracy to absurdity without eliminating it, and honesty requires saying that the theorem’s conclusion is conditional on this and cannot be made otherwise.

The local model drawn is one model. The sawtooth is the best pre-agreed list, not the only one; the theorem’s force comes from covering every list at once, which is algebra rather than something a drawing can display.

What the picture cannot show

It cannot show the state. Nothing in these figures draws the pair. The correlations are what the state produces, and the object producing them — an equal superposition of two joint possibilities, with a relative sign — has no picture, because it is not a property of either particle separately. That is the actual content of the word entanglement, and it is invisible here in the way that phase is invisible in an interference pattern.

It cannot show the separation. The horizontal axes are angles and counts. The whole locality condition is a statement about spacetime — that the two measurement events lie outside each other’s light cones — and needs a diagram with axes of a different kind to state at all. A correlation plot is silent about whether the correlation was measured across a laboratory bench or across 1.3 km.

It cannot show a single pair. Every curve here is an average over thousands of coincidences, and no individual pair is anomalous or informative. A single record reads as two coin flips, exactly as a single decay reads as an event with no cause. What is being drawn is a property of an ensemble, and the ensemble is the only place the property exists.

It cannot show what makes an outcome come out. The sampler draws from the state’s own joint probabilities and says nothing about why a given pair gave plus and minus rather than the reverse. Quantum mechanics supplies the weights and no mechanism beneath them, and the figures decline to invent one — the same silence the double slit keeps about which path was taken.

The ladder from here

The rungs above this one: the GHZ state, where three particles turn the statistical argument into a single run that a local model must get flatly wrong; entanglement swapping, which entangles two particles that never met and is the working part of a quantum repeater; the monogamy of entanglement, which says that a maximal correlation with one partner forbids any with a third and is what makes eavesdropping detectable; teleportation, which moves a state using a shared pair and two classical bits and so shows exactly how much the pair carries; and decoherence, where entanglement with an uncontrolled environment reproduces everything collapse was invented to explain.

Identical particles and the symmetry of a two-particle state is the nearest neighbouring ladder, since exclusion is what a particular entangled state looks like when the two labels cannot be told apart. Beside it sit what a measurement is, the price of sharpness, the arrival of light in lumps, the counting that underlies every statistical law, and the boundary where the quantum account hands back the classical one — which is where the question of why the world does not look like this page belongs.

The claim to carry forward is a claim about shape. A correlation’s strength can always be arranged in advance; its form cannot. The gloves in the two boxes agree perfectly and explain nothing, and the only thing that distinguishes a pair of photons from a pair of gloves is what happens to the agreement when somebody turns a dial.

Part 1 of 5

This essay is one argument about Entanglement. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

CausalityDecoherenceEntanglementInterferenceMalus's lawPhotonPolarisationProbability densitySpinSuperposition