Optics

The crystal that answers twice

Lay a piece of calcite on a printed page and the print appears twice. One image sits still when the crystal is turned and the other goes round it. Nothing has been done to the light except pass it through a material whose response to a field is not a number.

Assumes: The direction of the shaking, and the filter that only asks about it · The bend at the boundary, and what it is really about

A refractive index is written as a number, and a number is what it is for glass and water and air — enough to fix the whole of refraction when the medium has no preferred direction. It is not what it is for most solids. A crystal is an arrangement of atoms that is not the same in every direction, and its response to an applied electric field need not be parallel to that field — so the quantity relating the two is a matrix, and the speed of light in the material depends on which way the light is going and which way it is shaking.

The index that depends on which way the light is going. The two refractive indices of 3 uniaxial crystals, against the angle between the wave normal and the crystal's optic axis. The flat lines are the ordinary index, which is the same in every direction because the ordinary wave's field is always perpendicular to the axis. The curves are the extraordinary index, which runs from the ordinary value along the axis — where the two waves are identical and the crystal behaves like glass — to its extreme value at right angles to it. calcite (CaCO₃) has n_o = 1.6584 and n_e = 1.4864, so n_e − n_o = -0.1720; quartz (SiO₂) has n_o = 1.5443 and n_e = 1.5534, so n_e − n_o = 0.0091; lithium niobate has n_o = 2.3005 and n_e = 2.2075, so n_e − n_o = -0.0930. The sign of that difference is what makes a crystal positive or negative, and it decides which of the two images in a double-refracting crystal is the one that moves.
Fig. 1 The two refractive indices of three crystals, against the angle between the wave normal and the crystal’s optic axis. The flat lines are the ordinary index, which is the same in every direction. The curves are the extraordinary index, which runs from the ordinary value along the axis to its extreme value at right angles to it. Along the axis the two coincide and the crystal is indistinguishable from glass.

For the simplest case — a uniaxial crystal, with one special direction — that costs exactly one extra number. One wave, called ordinary, always has its field perpendicular to the optic axis whatever direction it travels, so it sees the same index non_o always. The other, called extraordinary, sees an index that depends on the angle θ\theta between its wave normal and the axis:

1n(θ)2=cos2θno2+sin2θne2.\frac{1}{n(\theta)^2} = \frac{\cos^2\theta}{n_o^2} + \frac{\sin^2\theta}{n_e^2}.

That is the equation of an ellipse, which is why the effect is described by an index ellipsoid rather than by a pair of numbers.

Which wave is which

Two waves can travel in a given direction through a crystal, and which of them a particular beam becomes is decided by its polarisation. The rule is geometric. For a wave normal at angle θ\theta to the optic axis, there is a plane containing both the normal and the axis — the principal plane — and the two allowed states are: field perpendicular to that plane, and field in it. The first never has any component along the axis, whatever θ\theta is, so it always sees non_o; that is the ordinary wave, and it is ordinary in the precise sense that Snell’s law with a single index describes it completely. The second has a component along the axis that grows with θ\theta, so the index it sees is a mixture, and it is the extraordinary one.

Unpolarised light entering a crystal is therefore not split by anything doing sorting. It is resolved, in the same sense a polariser resolves it: the incoming field is written as a sum of the two allowed states, each of those propagates according to its own index, and what emerges is two beams with orthogonal polarisations and no light lost. Putting a polariser in front of the crystal and turning it makes each image vanish in turn, which is the standard demonstration that the two beams are polarised and the standard way of proving that this is a resolution and not a filtering.

What is not the reason for the double image

Calcite’s two images are the effect everybody meets first, and the usual explanation — two indices, so two refracted rays — is wrong in an instructive way.

It is worth ruling out the obvious explanation first. Ordinary refraction at a boundary bends a ray by an amount set by the ratio of the indices — and at zero incidence it bends nothing, whatever the ratio. So a beam entering a crystal face square on should not split at all, and in calcite it does. Whatever produces the second image, it is not the boundary doing something twice.

So a beam entering a calcite slab at normal incidence has both of its wave normals going straight through. Snell’s law has nothing to say. And yet the two images are separated, and the separation is a millimetre in a centimetre of crystal.

Two rays out of one, 1.091 mm apart. A beam entering a 10 mm slab of calcite (CaCO₃) at normal incidence, with the optic axis at 45° to the surface. Both waves travel straight — at normal incidence the wave normals are not refracted at all, because Snell's law with a zero angle of incidence gives a zero angle of refraction for any index whatever. The extraordinary wave's energy does not follow its wave normal: the electric displacement and the electric field are not parallel in a crystal, so the Poynting vector tilts by 6.22° and the ray emerges 1.091 mm to one side. That is the double image, and it is why one of the two images rotates when the crystal is turned while the other stays put. The walk-off is largest at 41.9°, where it reaches 6.26°, and it is exactly zero along the axis and across it.
Fig. 2 The real mechanism. Both wave normals go straight; the extraordinary wave’s energy does not follow its own wave normal, because in a crystal the electric displacement and the electric field are not parallel, so the Poynting vector tilts away by 6.22° here. The ray emerges 1.09 mm to one side of where it entered, and that is the double image. The walk-off is exactly zero along the optic axis and at right angles to it, and largest at 41.9°.

This is the one place in elementary optics where the distinction between a wave normal and a ray has consequences a reader can see with their own eyes. In an isotropic medium the two coincide and the distinction is pedantry. In a crystal they differ by up to six degrees, and turning the crystal turns the extraordinary ray’s offset with it while the ordinary ray goes straight — which is why one image goes round the other.

Two rays out of one, 0.205 mm apart. A beam entering a 5 mm slab of lithium niobate at normal incidence, with the optic axis at 40° to the surface. Both waves travel straight — at normal incidence the wave normals are not refracted at all, because Snell's law with a zero angle of incidence gives a zero angle of refraction for any index whatever. The extraordinary wave's energy does not follow its wave normal: the electric displacement and the electric field are not parallel in a crystal, so the Poynting vector tilts by 2.34° and the ray emerges 0.205 mm to one side. That is the double image, and it is why one of the two images rotates when the crystal is turned while the other stays put. The walk-off is largest at 43.8°, where it reaches 2.36°, and it is exactly zero along the axis and across it.
Fig. 3 The same effect in lithium niobate, whose birefringence is half as large as calcite’s but whose indices are much higher. The walk-off is smaller and the crystal is used for it anyway: the material is transparent, electrically controllable and available in optical quality, and every one of these numbers is set by two measured indices and an angle a technician chooses when the boule is cut.

Where the anisotropy comes from

The index tensor has been taken as given, and it is worth one paragraph on why a crystal has one, because the answer explains why calcite’s birefringence is so much larger than everything else’s.

A refractive index is a statement about how readily the material’s charges are displaced by a field — the polarisability, averaged over the structure. In an isotropic material that displacement is parallel to the field and its size is one number. In a crystal, the electrons are held by bonds that point in particular directions, so a field along a direction the bonds favour displaces them further than a field across it, and the induced polarisation is neither parallel to the field nor the same size in every direction.

Calcite is the extreme case among common minerals, and the reason is visible in its structure. It is built of flat, triangular carbonate groups, all lying in parallel planes stacked perpendicular to the optic axis. The electrons in such a group are easy to push around within the plane of the triangle and hard to push out of it, so a field lying in those planes sees a large polarisability and a field along the axis a much smaller one. The result is an ordinary index of 1.658 and an extraordinary one of 1.486 — a difference of 0.172, which is two hundred times quartz’s.

The sign follows from the same picture. Calcite is called negative uniaxial because ne<non_e < n_o, and that is the direct consequence of the axis being the hard direction; a crystal whose special direction is the easy one is positive, and quartz, whose structure is a helix of linked tetrahedra rather than a stack of flat groups, is.

Which makes the whole of this essay a consequence of the dielectric response being a tensor rather than a scalar, and the tensor being a tensor because a crystal’s bonds point somewhere.

Two components, one delay

The other use of a crystal is not to separate the two waves in space but to delay one against the other in time. Cut a slab with its optic axis lying in the surface, and light entering at normal incidence has both waves travelling straight down the same path — no walk-off, because the wave normal is at right angles to the axis and the walk-off vanishes there — at different speeds. They emerge together, out of step.

The phase difference accumulated in a thickness dd is

Γ=2πdnoneλ,\Gamma = \frac{2\pi d\,|n_o - n_e|}{\lambda},

and choosing dd makes Γ\Gamma whatever is wanted.

A wave plate of quartz (SiO₂), in thickness. How much light a slab of quartz (SiO₂) passes between crossed polarisers, against its thickness, at 589 nm with the plate's axes at 45° to both. The two components travel at different speeds and arrive with a phase difference proportional to the thickness, so the transmission is the square of a sine and returns to zero every time the difference reaches a whole cycle. The quarter-wave thickness is 16.18 μm and the half-wave thickness is 32.36 μm; the first turns linear polarisation into circular and passes half the light, and the second turns it through 90° and passes all of it. Nothing is absorbed at any thickness: the light that does not emerge is the light the second polariser sends the other way.
Fig. 4 A slab of quartz between crossed polarisers, against its thickness. The transmission is the square of a sine and returns to zero every time the delay reaches a whole cycle. The quarter-wave thickness is 16.2 μm and the half-wave thickness twice that; nothing is absorbed at any thickness, and the light that does not emerge is the light the second polariser sends the other way. Quartz is used for this rather than calcite because calcite’s birefringence is twenty times larger and its quarter-wave plate would be 0.86 μm thick — thinner than a soap film. A retardation can also be had with no crystal in it at all, from two total internal reflections, and that version is achromatic where this one is not.

What the delay does to the light is best seen by tracing the electric field.

One input, five outputs, and only the delay is different. Light polarised at 45° to a crystal's fast axis, drawn as the path its electric field traces in a plane over one cycle, after passing through plates of five different thicknesses. The two components are unchanged in amplitude — a wave plate absorbs nothing and rejects nothing — and the only thing that differs between these panels is how far one component has been delayed against the other. At no delay the field oscillates along a line. At a quarter of a cycle it goes round a circle: the same two oscillations, the same amplitudes, and a state that has no direction of oscillation at all. At half a cycle it is a line again, turned through twice the angle between the input and the axis. Circular polarisation is not a third kind of light; it is two of the first kind, out of step.
Fig. 5 The path the field traces over one cycle, for five different delays, with the input polarised at 45° to the crystal’s fast axis. The two components are unchanged in amplitude throughout — a wave plate absorbs nothing — and the only thing that differs between the panels is how far one has been delayed against the other. At a quarter cycle the field goes round a circle: the same two oscillations, the same amplitudes, and a state with no direction of oscillation at all.

The last of those panels is the half-wave plate, which turns linear polarisation through twice the angle between the input and the crystal’s axis. Set the plate at 45° and the polarisation is turned by 90°, which is the cheapest way to rotate a beam’s polarisation without rotating anything mechanical through the beam.

What circular polarisation is not

The natural reading of that circle is that a third kind of light has been produced. It has not.

Polarisation is a statement about the electric field’s behaviour over a cycle, not about a second field or a second wave. In an electromagnetic wave the electric and magnetic fields are at right angles to each other and to the direction of travel, and what distinguishes linear from circular is only how the electric vector moves as the wave goes by — steady in one direction, or turning. There is one wave in both cases.

Circular light is two linear oscillations of equal amplitude, a quarter cycle apart, in perpendicular directions. It could equally be described as a single state in a basis of left- and right-circular components, in which case linear light is the sum of two circular ones. Which description is the real one is not a question a measurement can answer, and both are used: chemists working with sugars think in circular components because a chiral molecule treats them differently, and everybody working with polarisers thinks in linear ones because a polariser is a linear device.

Malus's law. The fraction of polarised light passing a filter, against the angle between the light's own direction of shaking and the filter's axis. It is the cosine squared: half at 45 degrees, nothing at 90.
Fig. 6 Malus’s law: the intensity through a second polariser as the cosine squared of the angle between them. A polariser takes a component rather than sorting light into kinds, which is the claim the rung below this one was about. A wave plate does something a polariser cannot: it changes the relative phase of two components without changing either amplitude, so nothing is discarded and no light is lost.
3 filters at 0°, 45°, 90°: 12.5% gets through. Unpolarised light passing through 3 polarising filters with axes at 0 degrees, 45 degrees, 90 degrees. The first removes half whatever its angle; each one after it passes the cosine squared of the turn from the filter before. 12.5 per cent of the original intensity survives.
Fig. 7 The three-filter experiment: two crossed polarisers pass nothing, and inserting a third between them at 45° makes light appear. A quarter-wave plate inserted in the same place does the same thing by a completely different route — it passes all the light rather than a quarter of it, because it rotates the state instead of projecting it. Comparing the two is the cleanest way to see that a projection loses something and a retardation does not.

The instruments this makes possible

Almost every use of birefringence is one of the two effects above put to work.

Separating two polarisations in space uses the walk-off. A calcite beam displacer is a slab with its axis cut at 45°, and it takes one beam in and gives two parallel beams out, laterally displaced, orthogonally polarised and both at full brightness. A polarising beamsplitter built out of a polariser throws half the light away; one built out of a crystal throws none of it away, and that difference matters wherever the light is scarce. The Nicol prism, which was the standard polariser for a century, goes further: it splits the two rays inside calcite and then gets rid of one of them by total internal reflection at a cemented joint cut so that the ordinary ray exceeds the critical angle and the extraordinary one does not.

Changing a polarisation state in time uses the retardance. A quarter-wave plate in front of a mirror is an optical isolator — though not a non-reciprocal one, which needs a rotation a return trip doubles: light goes in linear, comes back circular, passes the plate again and returns linear at 90° to where it started, so a polariser that let it out will not let it back in. That arrangement protects lasers from their own reflections, and it is why a plate whose whole function is to delay one component by a quarter of a cycle turns up in equipment that has nothing else optical about it.

And making the retardance controllable turns the plate into a display. A liquid-crystal cell is a layer whose birefringence responds to an applied voltage, sandwiched between crossed polarisers: at one voltage it is a half-wave plate and the pixel is bright, at another it retards nothing and the pixel is dark. Every screen this essay might be read on is several million of the arrangement in the retardance figure, addressed individually.

Thirty micrometres, and why

The colours between crossed polarisers are usually treated as a decoration, and one whole science uses them as a measurement instead.

A geological thin section is a slice of rock ground to a standard thickness of thirty micrometres, mounted on a slide and examined between crossed polarisers on a rotating stage. Every mineral in it retards light by its own birefringence times that thickness, and the retardation decides the colour — the same subtraction of one wavelength from a white spectrum a thin film performs: white where every wavelength gets through, and successively more saturated colours as particular wavelengths are extinguished and the transmission curve of the retardance figure is climbed.

The standard thickness exists because it makes the colours a reference. Quartz has a birefringence of about nine thousandths, so at thirty micrometres it retards by around 270 nanometres — a pale grey at the top of the first order, which is instantly recognisable and present in almost every rock. The section is ground until the quartz looks right, and every other mineral in the field is then read against a chart of colour against retardation, which converts a colour directly into a birefringence and a birefringence into an identification.

That is an unusually direct instrument. A property of a material — the difference of two indices — is read off as a colour by eye, at a standardised thickness, with no photometry and no calibration beyond the grinding. And the sensitivity is remarkable: minerals differing by a few thousandths in birefringence are told apart at a glance, which is a measurement of an index difference to a part in a thousand made without measuring an index at all.

Rotating the stage adds the second half of the information. As the crystal turns, its axes sweep past the polariser’s direction, and the pixel goes dark four times per revolution — at the angles where one of the crystal’s own axes lines up with the polariser and there is nothing to resolve into two components. Where those extinctions fall relative to a visible crystal face is a statement about the orientation of the optic axes in the crystal, which is a further identification.

Birefringence without a crystal

Nothing in the argument requires a crystal, and two cases where there is none are worth having.

The first is a material with no anisotropy in its molecules at all, made anisotropic by its geometry. A stack of alternating layers, each isotropic and each much thinner than a wavelength, behaves as one medium — the light averages over many layers and cannot resolve them — and that averaged medium has different indices for a field along the layers and across them. It is form birefringence, and it is why some biological tissues, which are stacks of membranes, show between crossed polarisers even though nothing in them is crystalline.

The second is stress. Squeeze an isotropic solid and its bonds are no longer equivalent in every direction, so it becomes weakly birefringent, with the axes along the principal stresses and the retardation proportional to the difference between them. A transparent model of a structure, loaded and viewed between crossed polarisers, therefore displays a map of its own internal stress as a pattern of fringes, each fringe marking a contour of constant stress difference.

That was the standard method of stress analysis before anything could be computed, and it is still the quickest way to see where a load concentrates. It also means that a moulded plastic object viewed through polarised light shows the stresses frozen into it when it cooled — which is why a clear plastic lid seen through polarised sunglasses is covered in colours nobody put there.

Where the model stops

Uniaxial is the easy case. Most crystals are biaxial: three different principal indices, two optic axes rather than one, and a surface of wave normals that has four conical points where extremely odd things happen — a single ray entering along one emerges as a hollow cone of light, which Hamilton predicted in 1832 and Lloyd found within two months. None of that appears in the equation above, which has one special direction in it by assumption.

The indices depend on wavelength. Every number quoted here is at the sodium D line, and a quarter-wave plate cut for 589 nm is not a quarter-wave plate for 450 nm. That is why a plate between crossed polarisers shows colours in white light rather than simple brightness: the transmission is the square of a sine whose argument carries 1/λ1/\lambda, so different colours are at different points on the curve.

The index that depends on which way the light is going. The two refractive indices of 3 uniaxial crystals, against the angle between the wave normal and the crystal's optic axis. The flat lines are the ordinary index, which is the same in every direction because the ordinary wave's field is always perpendicular to the axis. The curves are the extraordinary index, which runs from the ordinary value along the axis — where the two waves are identical and the crystal behaves like glass — to its extreme value at right angles to it. quartz (SiO₂) has n_o = 1.5443 and n_e = 1.5534, so n_e − n_o = 0.0091; sapphire (Al₂O₃) has n_o = 1.7681 and n_e = 1.7599, so n_e − n_o = -0.0082; muscovite mica has n_o = 1.5980 and n_e = 1.5936, so n_e − n_o = -0.0044. The sign of that difference is what makes a crystal positive or negative, and it decides which of the two images in a double-refracting crystal is the one that moves.
Fig. 8 Three weakly birefringent materials on the same axes as the hero figure. Quartz’s two indices differ by nine parts in ten thousand, mica’s by four, sapphire’s by eight — against calcite’s difference of 0.172, which is two hundred times larger. Everything in this essay applies to all of them; what changes is only how thick a slab has to be to do anything, and that is why mica sheets are used for wave plates and calcite for beam separators.

The plate is assumed thin enough to ignore the beam’s own spread. A wave plate delays one component against the other by a fixed phase only for a beam travelling exactly along the design direction. A converging or diverging beam contains rays at a range of angles, each seeing a slightly different index difference, so a plate placed in a focused beam retards different parts of it differently. This is why high-quality plates are made from two pieces of crystal with their axes crossed, so that the bulk of the retardance cancels and what is left is a small difference that is far less sensitive to angle.

And the crystal is assumed transparent and unstressed. Absorption that depends on polarisation — dichroism — is a different effect, and it is what a sheet polariser is made of. Stress makes an isotropic material birefringent, which is what turns a plastic ruler between crossed filters into a map of its own stresses, and it makes the effect a function of position rather than a property of the material.

One more limit is worth naming because it is the reason the equations here have any use at all. Everything above treats the crystal as a continuous medium with a smooth dielectric tensor, and a crystal is an arrangement of discrete atoms. The continuum description works because the wavelength of light is thousands of times the spacing between atoms, so the light samples an enormous number of them and sees only their average. At X-ray wavelengths the ratio is reversed, the averaging stops, and the same crystal is no longer a birefringent medium but a diffraction grating in three dimensions — a different phenomenon, from the same atoms, because the probe got smaller than the structure.

What the pictures cannot show

The polarisation-state figure draws the field’s path over one cycle, and a real optical cycle is 5×10145\times10^{14} per second. Nothing has ever seen that path. What is measured is always an average — an intensity through an analyser, at several analyser angles, from which the state is reconstructed — so the ellipse in the figure is an inference, and a very secure one, rather than an observation.

The walk-off figure draws two rays and no fields. What is actually different between the two rays is the direction the electric field points, and that is what selects which index each one sees; the drawing shows the consequence and cannot show the cause.

Where the ladder goes next

The polarisation ladder began with the direction of the shaking and the filter that asks about it and continued with the angle at which a reflection picks a side. This rung makes the medium itself directional. The rungs above it: optical activity, where a solution of chiral molecules rotates the plane continuously with no crystal axes anywhere; the electro-optic effect, where an applied voltage changes the indices and a wave plate becomes a switch running at gigahertz; the Fresnel rhomb, which makes circular light out of two total internal reflections and no birefringence at all; and the Jones and Mueller calculus, which turns all of this into matrix multiplication and handles partial polarisation, which none of the pictures here can.

The habit worth carrying away is what happens when a constant becomes a matrix. A scalar relation assumes the response is parallel to what caused it, and dropping that assumption is not a small correction — it produces two waves where there was one, an energy flow that is not along the wave normal, and a device whose whole function is that the two do not keep in step.

Part 3 of 8

This essay is one argument about Polarisation. The others:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the reading path: the things themselves, and every essay that touches each one.

AnisotropyBirefringenceCircular polarisationCrystalPhasePolarisationRefractive indexSnell's lawSuperpositionWave plate