What the eye does

Three numbers

A spectrum has as many degrees of freedom as anyone cares to give it. The eye reports three. Everything colour science can do, and every way it fails, follows from that one collapse.
16 min read 8 figures Three numbersComputed, not quoted

Light arriving at the eye is described by a function: how much power at each wavelength. Sample it every five nanometres across the visible range and that is eighty-one numbers; sample it every nanometre and it is four hundred; the underlying object has no finite dimension at all.

What leaves the retina, as far as colour is concerned, is three numbers.

A spectrum, weighted three ways, and the three numbers left overThe illuminant D65 above; below, the same spectrum multiplied by each matching function. The area under each product is one coordinate of XYZ. Everything else about the spectrum — its shape, its structure, all its remaining degrees of freedom — is discarded here.D65 spectrum400450500550600650700wavelength / nmx̄ → 95.04ȳ → 100.00z̄ → 108.90the area under each product is one coordinateCIE 1931 2° observer
Fig. 1 The illuminant D65 above; below, the same spectrum multiplied by each of the three colour-matching functions. The area under each product is one coordinate. Everything else about the curve — its structure, its remaining degrees of freedom, all of them — is discarded at this step and cannot be recovered. The handle moves the illuminant. The spectrum entering the collapse changes at every stop and the count coming out of it never does.

The same three weightings applied to a different spectrum give three different numbers, and nothing about the operation changes.

A spectrum, weighted three ways, and the three numbers left over. The illuminant A above; below, the same spectrum multiplied by each matching function. The area under each product is one coordinate of XYZ. Everything else about the spectrum — its shape, its structure, all its remaining degrees of freedom — is discarded here.
Fig. 2 Illuminant A put through the same three weightings. The spectrum bears no resemblance to the one above and the arithmetic is identical: multiply by three curves, integrate, keep three numbers.

A spectrum with no structure at all is the sharpest version of the same point, because there is nothing in it for the three curves to be selective about.

A spectrum, weighted three ways, and the three numbers left over. The illuminant E above; below, the same spectrum multiplied by each matching function. The area under each product is one coordinate of XYZ. Everything else about the spectrum — its shape, its structure, all its remaining degrees of freedom — is discarded here.
Fig. 3 And a flat spectrum, which is the clearest case of all. Every wavelength arrives with the same power, so the three numbers that come out are the areas under the three curves and nothing else — the collapse with the input taken out of it.

Four of them is enough to see what the four have in common, which is everything that matters and is not the spectrum.

A spectrum, weighted three ways, and the three numbers left over. The illuminant D50 above; below, the same spectrum multiplied by each matching function. The area under each product is one coordinate of XYZ. Everything else about the spectrum — its shape, its structure, all its remaining degrees of freedom — is discarded here.
Fig. 4 A fourth, for the sake of what the four have in common. Whatever goes in, three numbers come out, and the function that went in cannot be recovered from them.
A spectrum, weighted three ways, and the three numbers left over. The illuminant D65 above; below, the same spectrum multiplied by each matching function. The area under each product is one coordinate of XYZ. Everything else about the spectrum — its shape, its structure, all its remaining degrees of freedom — is discarded here.
Fig. 5 And the first spectrum again under the ten-degree observer. The three weighting functions have changed and the operation has not, so which three numbers come out is a property of the observer and the collapse is a property of the arithmetic.

This is the single most consequential fact in the subject, and almost everything else here is a consequence of it.

The three receptors

The retina carries three classes of cone, distinguished by which wavelengths they absorb. Each one integrates the incoming spectrum against its own sensitivity curve and reports a single number — how much it caught in total. It has no way to report where in the spectrum the light came from, because that information is gone the moment the photons are counted.

They are conventionally named L, M and S, for long, medium and short wavelength. They are frequently called red, green and blue cones, and that naming is wrong in three separate ways: the L cone peaks in the yellow-green rather than the red, the L and M curves overlap almost completely, and no cone corresponds to a colour, since colour is what the brain makes of the three responses together.

The measured gap between the L and M peaks is twenty-five nanometres, out of a visible range spanning three hundred and twenty. Two of the three receptors are looking at very nearly the same thing.

What follows immediately

The mapping is many-to-one, massively. An infinite-dimensional space maps onto three dimensions, so the pre-image of any colour is an infinite-dimensional family of spectra. This is metamerism and it is not an edge case; it is the generic situation.

Colour reproduction is possible at all. If the eye reported the full spectrum, a display would have to reproduce the full spectrum. Because it reports three numbers, three primaries suffice — and the entire industry of screens, printing, photography and paint rests on that one fact.

Colour reproduction is limited in a specific way. Three fixed primaries mixed in non-negative amounts reach a triangle, and the set of visible chromaticities is not a triangle, so some colours cannot be reproduced by any three-primary system whatever.

Some physical differences are invisible in principle. Not hard to see — invisible. Two lights differing by a spectrum in the null space of the three matching functions produce identical cone responses, and no amount of attention or training will separate them.

Why three, and not two or four

The number is not a law of physics. It is a fact about a particular species, and it varies.

Most mammals are dichromats, with two cone classes; dogs and cats see a two-dimensional colour space. Old-world primates including humans have three, the result of a gene duplication that split an ancestral long-wavelength pigment into the L and M pair — which explains why those two peaks sit so close together, since they are recent copies of one another.

Many birds, reptiles and fish are tetrachromats, with four classes, often including one sensitive into the ultraviolet. Their colour space has four dimensions, and a metameric pair for a human is generally not a metameric pair for them. This has a practical consequence that is easy to overlook: colour reproduction is species-specific. A photograph reproduces colour for humans and does not reproduce it for a bird, because it was engineered against the human matching functions and against nothing else.

Some human females carry four cone pigments as a consequence of the X-linked inheritance of the L and M genes. Whether that produces genuine tetrachromatic vision — a four-dimensional colour space rather than three — is a question about the visual cortex rather than the retina, and the evidence is that it usually does not.

The mathematics of the collapse

Write the spectrum as a vector ss with one component per sampled wavelength, and the three sensitivity curves as the rows of a matrix AA. The cone responses are

r=As\mathbf{r} = A\,s

and that is the whole of it. Everything downstream is coordinate changes applied to r\mathbf{r}.

Two properties of this equation carry most of the subject:

It is linear. Doubling the light doubles the responses; adding two lights adds their responses. This is why colours mix the way they do, why the chromaticity diagram’s mixing line is straight, and why the whole apparatus can be built out of matrices. It holds to good accuracy over a wide range of intensities and fails at the extremes.

Its null space is enormous. Any bb with Ab=0Ab = 0 is invisible. Since AA has three rows and eighty-one columns in this site’s sampling, the null space has seventy-eight dimensions.

That figure is the collapse seen from underneath. Seventy-eight of the eighty-one degrees of freedom in a sampled spectrum are invisible in principle, and no attention, training or viewing condition recovers any of them.

What the standard observer is

The three curves in the figure above are not measurements of cones. They are colour-matching functions: the amounts of three chosen primaries needed to match each pure wavelength, determined by asking people to adjust knobs until two halves of a field looked the same.

That is a behavioural experiment, not a physiological one, and it is worth being clear that the CIE system was built from matching data decades before anyone could measure a cone pigment directly. The functions are a model of what matches what, and the cone fundamentals used on this site are derived from them by a linear transformation rather than measured independently.

The consequences are dealt with in seventeen observers in 1931, and the short version is that the standard observer is an average over a small number of people, is known to be wrong in the blue, and is not anybody.

What the collapse does not do

Here is the error the collapse most often invites: treating the three numbers as a description of appearance.

They are not. They are a description of the stimulus — what arrived, integrated three ways. How a patch of that stimulus looks depends on what surrounds it, what preceded it, how bright the room is, and what the visual system has adapted to. Two patches with identical cone responses can look plainly different, and that is not a failure of the measurement but a category error about what was measured.

The three numbers predict when two things will match under identical conditions. That is a genuine and useful thing to be able to predict, and it is narrower than it sounds.

The collapse is why reproduction is possible

The practical consequence deserves stating on its own, because it is easy to treat metamerism as a defect rather than as the thing that makes an industry exist.

Two different spectra that are the same colourTwo reflectance curves differing by 92 per cent RMS, and the two patches they produce under D65: identical to ΔE00 = 6.2e-14, which is arithmetic noise rather than a small number. Both patches are inside the sRGB gamut, so neither has been clipped into agreement.400450500550600650700wavelength / nmthe two coloursΔE00 = 6.2e-14spectra differ by 92%reflectances, under D65CIE 1931 2° observer
Fig. 6 Two reflectance curves twenty per cent apart, and the two patches they produce. The patches agree to arithmetic noise. Nothing about the light has been reproduced; what has been reproduced is the triple of numbers, which is all that was ever going to leave the retina.

A screen showing a photograph of a leaf is not emitting the spectrum of a leaf. It is emitting some mixture of three primaries chosen so that the three integrals come out the same. The physical light is entirely different and the colour is the same, because colour is the three numbers and not the light.

This is why three primaries suffice for a display, three inks plus black for print, three channels for a camera, three numbers for a file format. Every one of those choices traces back to the count of cone classes, and would be different for an observer with four.

What the three numbers do not carry

The collapse discards more than most descriptions admit, and two of the losses are worth naming.

Spectral structure is gone entirely. A tungsten lamp and a white LED can be arranged to have the same chromaticity while having radically different spectra — one smooth, one a blue spike plus a phosphor hump. The three numbers cannot distinguish them, and neither can an observer looking at the lamps directly. Shine both on a set of coloured surfaces and the differences appear immediately, because the surfaces reweight the spectra before the eye integrates them.

Wavelength is not recoverable. There is no sense in which a colour “is” a wavelength. Most colours correspond to no wavelength at all: every purple, every brown, every pastel, white itself. The identification of colour with wavelength survives in casual usage and is wrong in both directions — most colours have no wavelength, and no wavelength can be displayed anyway.

Where the linearity fails

The equation r=As\mathbf{r} = As is a very good model over the range of intensities encountered in ordinary viewing, and it is worth knowing where it stops.

At low light the cones stop responding altogether and the rods take over. Rod vision is monochromatic — one receptor class, one number, no colour at all. The transition is gradual and there is a substantial range where both contribute, during which colour vision is real but distorted. The CIE’s standard observers are photopic, meaning they describe the cone-only regime, and applying them to dim scenes is simply outside their remit.

At high light the cones saturate and then bleach, which is the mechanism behind the afterimage left by a bright light.

And even within the linear regime, what is linear is the response, not the appearance. Doubling the light does not double the brightness; perceived lightness follows something much closer to a cube root, which is why CIELAB has a cube root in it and why a mid-grey code value carries about a fifth of white’s luminance rather than half.

Rods, and the fourth receptor nobody counts

The retina has a fourth photoreceptor class, and it is left out of every calculation on this site for a reason worth stating.

Rods are far more numerous than cones and far more sensitive, and they have a single spectral sensitivity. One receptor class means one number, which means no colour: rod vision is genuinely monochromatic. In dim light the cones stop responding, the rods take over, and colour vision simply stops — the familiar experience of a moonlit landscape having shapes and brightness but no hues.

Between the two regimes is a range where both contribute, which is called mesopic vision and is the lighting engineer’s nightmare. Colour vision exists there but is distorted, the effective luminous efficiency curve shifts toward the blue, and no standard observer describes it. The CIE’s 1931 and 1964 observers are both photopic — they describe the cone-only regime and nothing else — so every number on this site is implicitly a claim about a reasonably well-lit scene.

The shift in efficiency has a name and a visible consequence. The Purkinje effect: as light fades, blues appear to brighten relative to reds, because the rods peak at a shorter wavelength than the cones do. A red flower and a blue flower of equal apparent brightness in daylight will not be equal at dusk, and the blue will win.

Three numbers, and then what

The collapse to three is the retina’s contribution. What happens next is a recombination that changes the axes without changing the count.

The signals are converted almost immediately into opponent channels: something like light-versus-dark, red-versus-green, and blue-versus-yellow. The evidence is partly anatomical and partly phenomenological, and the phenomenological part is easy to check. There is no colour that is simultaneously reddish and greenish, or simultaneously bluish and yellowish, whereas reddish-yellow and bluish-green are ordinary. The four unique hues feel more fundamental than the mixtures between them despite having no special status in the cone responses.

This is why the afterimage of a red patch is green, and why the perceptual colour spaces are all built with an opponent structure — CIELAB’s aa^* and bb^* axes are red-green and blue-yellow, and Oklab’s are the same idea fitted more carefully.

Opponency and trichromacy were rival theories through the nineteenth century, Hering against Helmholtz, and the resolution was that they describe different stages of the same pathway. Three receptors, then three opponent signals computed from them. The dimension count survives; the axes do not.

What was computed here

The cone fundamentals in the second figure are not tabulated. They are computed by taking a monochromatic stimulus at each wavelength, integrating it against the CIE colour-matching functions to get XYZ, and transforming that into cone space with the Hunt–Pointer–Estévez matrix. Deriving them keeps them from drifting out of agreement with the rest of the site.

Two checks guard the result. The peaks must land in the right places — the computed values are 575, 550 and 445 nm, against expected ranges that would catch a transposed matrix, which otherwise produces three plausible bell curves in the wrong positions. And the L–M separation must come out small, since a transformation that pushed them apart would be describing a different eye.

The collapse figure is checked differently: the three shaded areas must reproduce the XYZ that the direct integral gives, so the picture and the number come from the same arithmetic rather than being computed twice with a chance to disagree.

The collapse does not care which lamp it is handed, and a tungsten spectrum under the ten-degree observer is the least daylight-like case available.

A spectrum, weighted three ways, and the three numbers left over. The illuminant A above; below, the same spectrum multiplied by each matching function. The area under each product is one coordinate of XYZ. Everything else about the spectrum — its shape, its structure, all its remaining degrees of freedom — is discarded here.
Fig. 7 Illuminant A above and the same spectrum multiplied by each of the 1964 matching functions below. The area under each product is one coordinate of XYZ, and everything else about the spectrum is gone.

Three numbers, and then two and a half

The collapse to three is exact. What happens to the three immediately afterwards is not a fourth collapse but something close to one, and it is easy to miss because the dimension count does not change.

The L and M fundamentals overlap heavily — their peaks are about 30 nm apart on curves several times that wide — so their outputs to any natural spectrum are very nearly the same number. Measured across a family of smooth reflectances under daylight, the two correlate at 0.999.

Sending three such signals down the optic nerve would spend two channels carrying one quantity. What is sent instead is a sum and two differences, and the differences are wildly unequal in how much they carry: decorrelating the cone responses gives an achromatic axis holding some 97% of the variance, a blue-yellow axis holding under 3%, and a red-green axis holding about a ten-thousandth.

So the space is three-dimensional and its three dimensions are nothing like the same size. The red-green axis — the one most people mean by “colour”, the one most colour blindness affects, the one CIELAB writes first — is the axis carrying least by three orders of magnitude, for exactly the reason it exists: L and M are nearly the same measurement, so their difference is nearly nothing, and nearly nothing was still worth extracting.

Three cones, two axes follows the recoding properly. What belongs here is that the count of three is a fact about dimensions and not about how much each one carries.

What the pictures cannot show

The cone curves are a model, and a contested one. The Hunt–Pointer–Estévez transformation is one of several in use — Smith & Pokorny, Stockman & Sharpe and CAT02 all give slightly different fundamentals, and any figure drawn from one of them inherits its assumptions. The site names the matrix wherever it matters.

More importantly, no figure here can show the collapse from the inside. The reader is looking at these spectra through the very apparatus under discussion, so the curve labelled “spectrum” is itself being delivered as three numbers per pixel. There is no vantage point from which the discarded information is visible.

Who found it, and when

Thomas Young proposed in 1802 that the eye must contain a small number of receptor types, on the grounds that it could not plausibly carry a separate mechanism for every wavelength. Hermann von Helmholtz developed the idea and the pair are usually credited together.

Maxwell made it quantitative in the 1850s with colour-matching experiments using spinning discs, and produced the first colour photograph in 1861 by the three-filter method — a direct application of the claim that three channels suffice. The cone pigments themselves were not measured until the 1960s, more than a century after the theory that required them.

The same operation on the other daylight illuminant, under the same observer, is what makes the point about the arithmetic rather than the light.

A spectrum, weighted three ways, and the three numbers left over. The illuminant D50 above; below, the same spectrum multiplied by each matching function. The area under each product is one coordinate of XYZ. Everything else about the spectrum — its shape, its structure, all its remaining degrees of freedom — is discarded here.
Fig. 8 D50 through the ten-degree observer. Three integrals, three numbers, and the shape of the curve — all of its remaining degrees of freedom — discarded by the same three multiplications.

What the fourth receptor does to the collapse

The collapse to three is exact for the cone system, and the cone system is not the whole retina.

Rods are the most numerous receptor by a factor of twenty and are the only ones working below about a hundredth of a candela. Their pigment peaks near 498 nm rather than at any of the three cone peaks, so they weight a spectrum differently — and between roughly a hundredth of a candela and five, both systems contribute at once.

So the three numbers are three numbers for a light-adapted observer. In the mesopic range the projection is onto something closer to four, weighted by an adaptation level nothing in the file records, and the whole apparatus of chromaticity quietly assumes it is not there. The eye that has no colour takes that up properly.

Where this goes next

The immediate consequence is metamerism, which is the collapse seen from the other side. The question of why the coordinate system is built the way it is leads to why colour is exactly three-dimensional, which deals with the awkward experimental fact that made XYZ necessary. And for what the three numbers do and do not entitle anyone to say, matching is not appearance.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 24 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Colour-matching functionsCone fundamentalsMetamerismNull spacePrimariesProjectionSpectral power distributionStandard observerTrichromacy