What the eye does

Three cones, two axes

The retina does not send three receptor signals down the optic nerve. It sends a sum and two differences, and the reason falls out of the statistics of natural light rather than out of anything the eye intended.

Assumes Three numbers and Why colour is exactly three-dimensional.

17 min read 8 figures Three numbersComputed, not quoted

Three receptors produce three numbers. The obvious thing to do with three numbers is send them somewhere, and the retina does not do the obvious thing.

What leaves the eye is a sum and two differences, and the transformation between them is the most consequential piece of arithmetic in the visual system — everything from the four unique hues to the shape of colour blindness to the axes of CIELAB descends from it.

The redundancy the recoding removes

The L and M cones peak near 564 and 534 nanometres. That is thirty nanometres apart on curves more than a hundred nanometres wide, so the two overlap almost completely.

Run a family of natural reflectances under daylight through both, and their responses correlate at 0.999. Knowing one is very nearly knowing the other.

Sending three such signals down the optic nerve would spend two channels carrying one quantity. The optic nerve is a genuine bottleneck — about a million fibres for a hundred and twenty million receptors, most of them rods — so the pressure to not waste it is real, and this is the classic case of it.

Deriving the axes rather than quoting them

The standard presentation gives the opponent channels as definitions: luminance is L+M, red-green is L−M, blue-yellow is S−(L+M). They are usually presented as facts about physiology, established by recording from cells.

They can be derived instead, and the derivation is more informative than the citation.

Take a family of smooth reflectances spanning what the world contains. Run them through the cone fundamentals. Take logarithms, because receptor response is compressive and the redundancy being removed is multiplicative. Compute the covariance and find its eigenvectors — the three directions along which the data varies independently.

What comes out is not chosen and not fitted:

  • An axis weighting the three cones at 0.576, 0.579, 0.577 — equal to two decimal places. An achromatic channel.
  • An axis at −0.475, −0.336, 0.813. S against the sum of L and M. Blue-yellow.
  • An axis at −0.665, 0.743, −0.081. L against M, with almost no S. Red-green.

Nothing in the calculation knows what a colour channel is. The opponent structure falls out of the statistics of natural spectra, and the eye’s job was to notice.

This is Buchsbaum and Gottschalk’s result from 1983, and it is one of the most satisfying in the subject: a physiological arrangement explained by the properties of the environment it evolved in, with the explanation checkable in a few lines of arithmetic.

The ordering, which is not the expected one

The three channels carry 97.1%, 2.85% and 0.012% of the variance.

The last of those is red-green, and it is the surprising one — because red-green is the channel everybody means by “colour”. It is the one colour blindness is usually about, the one CIELAB writes first, the one the word mostly refers to in ordinary speech. Across natural spectra it carries one part in ten thousand, three orders of magnitude below blue-yellow.

The first version of the check in this site’s gate demanded red-green second, on the strength of that intuition. It failed, and the measurement was right. The check now requires the ordering achromatic, blue-yellow, red-green, and requires red-green to be at least twenty times smaller than blue-yellow — asserted in that direction, because the ordering is the finding.

The reason is the correlation above. L and M are nearly the same measurement, so their difference is nearly nothing. A channel built from a difference between two near-identical signals has almost no signal in it.

What makes it worth building anyway is that “almost nothing” is not nothing, and the little that is there is useful. The L and M difference is what separates ripe fruit from leaves, and the standard account of why old-world primates re-evolved trichromacy is exactly that. A channel carrying one part in ten thousand of the variance was worth a gene duplication because of which part.

What a channel carrying one part in ten thousand is actually worth can be seen by removing it, which is what protanopia does and what the construction below simulates on a palette chosen to be distinguishable.

One palette under normal vision and protanopiaThe same 7 colours simulated by the Brettel–Viénot–Mollon construction at severity 1.0. protanopia collapses the red-green distinctions. This shows which discriminations survive, not what anybody sees.normal trichromatprotanopiaBrettel–Viénot–Mollon 1997, severity 1.0a simulation is not an experiencesRGB, D65
Fig. 1 Seven colours under normal vision and under protanopia at full severity, by the Brettel–Viénot–Mollon construction. The channel the variance measurement calls negligible is the one whose absence collapses several of these into pairs, which is the difference between how much information a channel carries and how much it is used for.

What the second stage explains

Several facts that make no sense at the receptor level become straightforward once the recoding is in place.

There is no reddish green. A single channel cannot be positive and negative simultaneously, so a colour cannot be at both ends of the red-green channel. It can be at one end of each channel, which is why reddish blue and yellowish green are ordinary. This is the strongest structural prediction opponent theory makes, and it is a prediction about a class of experiences rather than a number.

There are four unique hues. Two chromatic channels with two ends each. Where exactly they land is a harder question the theory answers less well.

Afterimages are complementary. Fix on red, look away, see green. Adaptation in an opponent channel drives it toward its opposite end, which a receptor-level account has to work considerably harder to produce.

Colour deficiency comes in the shapes it does. Losing L or M collapses the channel built from their difference, which is why the common deficiencies are red-green. Losing S — far rarer, and not sex-linked — collapses the other. The simulation is a projection onto the surviving channels.

Chromatic and achromatic resolution differ enormously. The achromatic channel carries nearly all the variance and is sampled finely; the chromatic channels are pooled far more heavily. This is why image and video compression can discard most of the chroma resolution and why nobody notices — chroma subsampling is an engineering exploitation of the second stage, arrived at empirically decades before the justification was available.

What compression already knew

The clearest practical evidence for the variance ordering is a piece of engineering that predates the measurement by decades.

Every image and video format in use discards most of the chromatic resolution and keeps the luminance resolution. JPEG converts to a luma-plus-two-chroma representation and subsamples the chroma; every video codec does the same, usually more aggressively, throwing away three quarters of the chroma samples. Nobody notices.

That works because the achromatic channel carries nearly all the variance and is sampled finely by the visual system, while the chromatic channels are pooled far more heavily in the retina and carry far less. A codec discarding chroma resolution is exploiting the second stage, and the arrangement was arrived at empirically — by people finding what could be thrown away without complaints — long before the information-theoretic justification existed.

The colour space every codec uses for this is a luma-chroma transform of RGB, and it is worth noticing that it is not the opponent basis derived above. It is a fixed matrix chosen for computational convenience and standardised in the 1950s for analogue television compatibility. It is opponent-shaped — one weighted sum and two differences — and it is not the decorrelating transform.

That it works nearly as well is the useful lesson. The decorrelation result says the achromatic direction dominates and the chromatic directions are small; any transform that separates a weighted sum from two differences captures most of the benefit, and the exact weights matter far less than the structure. Which is also why CIELAB’s opponent-shaped axes are useful while not being the opponent channels.

Hering and Helmholtz were both right

The nineteenth-century dispute is usually taught as a contest with a winner, and it was not.

Helmholtz, following Young, argued for three receptor types. Hering argued that colour experience is organised around opposed pairs — red against green, yellow against blue, black against white — and pointed at exactly the phenomena above: no reddish green, four elementary hues, complementary afterimages.

The arguments were treated as incompatible for eighty years. They describe consecutive stages. Three receptors feed two chromatic channels and one achromatic one, Helmholtz was right about the first stage, Hering about the second, and the resolution arrived in the 1950s when Hurvich and Jameson measured opponent responses directly.

The lesson generalises past colour. Two well-supported theories that appear to contradict each other are sometimes describing different levels of the same system, and the question “which is right” was the wrong question for eight decades.

Where CIELAB fits, and does not

CIELAB’s a* and b* axes are constantly described as the red-green and yellow-blue axes, and the description is an inheritance from this stage.

It is not accurate. If a* were the red-green axis, the four unique hues would sit at 0°, 90°, 180° and 270°; they sit at 25°, 88°, 161° and 250°. The axes were constructed to be roughly opponent-shaped and are not the opponent channels, and treating them as such — a quadrant test for hue category, a “warm or cool” heuristic — places boundaries twenty-odd degrees from where observers put them.

The two chromatic channels can also be removed one at a time and at partial strength, which separates the question of which channel is lost from the question of how completely.

One palette under normal vision and two dichromacies. The same 7 colours simulated by the Brettel–Viénot–Mollon construction at severity 0.7. deuteranopia collapses the red-green distinctions, and tritanopia leaves them and collapses blue against yellow instead. This shows which discriminations survive, not what anybody sees.
Fig. 2 Normal vision against deuteranopia and tritanopia at a severity of 0.7. The first row collapses red against green and the second leaves those and collapses blue against yellow, so the two rows are the two chromatic channels of the recoding, each removed on its own.

The one receptor that has nothing to decorrelate

A useful contrast: none of this applies to rod vision, and the reason is instructive about what the second stage is for.

Rods have one pigment. One receptor type produces one signal, one signal has no redundancy with anything, and there is nothing for an opponent stage to remove. Rod vision therefore has no chromatic channels, which is the same statement as saying it has no colour — not a deficiency of the rod pathway but a straightforward consequence of there being one of it.

This puts the trichromatic case in perspective. Three receptors are not three sources of information about colour — they are three heavily overlapping measurements of the same spectrum, from which a small amount of chromatic information can be extracted by differencing. The extraction is the second stage’s whole job, and the amount extracted is, by the measurement above, about 3% of what arrives.

The rest is luminance, and luminance is what a single receptor type already gives. Colour is the residual.

The loss is graded rather than binary, since an anomalous trichromat has a shifted pigment rather than a missing one, so the same construction at a third of full severity is closer to what most affected people carry.

One palette under normal vision and three dichromacies. The same 7 colours simulated by the Brettel–Viénot–Mollon construction at severity 0.3. protanopia and deuteranopia collapse the red-green distinctions, and tritanopia leaves them and collapses blue against yellow instead. This shows which discriminations survive, not what anybody sees.
Fig. 3 The same four rows at a severity of 0.3 rather than 1.0. Most of the palette survives, which is why an anomalous trichromat often reaches adulthood without discovering that their colour vision differs from anybody else’s.

What was computed here

The reflectances are constructed rather than measured, and the essay says so. They are a family of smooth curves with three degrees of freedom — level, slope and curvature — walked over a lattice.

That is an honest caricature of a real property: measured collections of natural reflectances turn out to need about three basis functions to reconstruct to within measurement error, which is a remarkable fact and not one the eye had any say in. A measured collection would be better and this site does not have one.

The eigendecomposition is by cyclic Jacobi rotation, written out rather than approximated, because the whole argument is that specific axes fall out of the data — an iterative method that stopped early would produce axes that were nearly right and prove nothing. Each eigenvector’s sign is fixed so its largest entry is positive, since an eigenvector is only defined up to a flip and nothing downstream would be comparable otherwise.

Each axis is classified from its own weights rather than from where it landed in the ordering. An eigenvector with three same-signed entries is achromatic wherever it came out; one setting L against M with little S is red-green. The classification is a test the vectors can fail, not a label applied to whichever axis came second.

Three things are asserted. L and M must correlate above 0.99, because without that redundancy there is nothing for the recoding to remove. The three axes must classify as achromatic, blue-yellow and red-green in that order. And the achromatic axis must weight the three cones within 0.05 of each other, since an achromatic channel that weighted them unequally would not be one.

The printed vectors are an orthonormal basis

The three eigenvectors are quoted to three decimals and it is worth checking that the printed numbers still are what they claim to be, because a reader has no other access to the computation.

Each has unit length to within 0.0004, and each pair is orthogonal to within 0.001. So the nine printed numbers really are three perpendicular directions of unit length — the object an eigenanalysis of a symmetric matrix returns — and nothing has been rounded past the point of being one.

That matters more here than in most tables, because the essay’s claim is that the axes fall out rather than being chosen. A set of weights that failed orthogonality would be a set somebody had adjusted.

Two point four orders, not three

Three orders of magnitude below blue-yellow is the essay’s summary of the red-green channel, and the three variances give a different gap.

97.1, 2.85 and 0.012 per cent put red-green 238 times below blue-yellow — 2.4 orders — and 8,092 times below achromatic, which is 3.9. So the three-orders figure sits between the two comparisons and belongs to neither.

The correction runs the same way as the finding. Red-green is nearly four orders below the achromatic channel, which is a larger gap than the essay claims for it against blue-yellow, and the surprising comparison is the one the essay does not make: the channel everybody means by colour carries one eight-thousandth of what the channel nobody thinks of as colour carries.

And the pair together carries 2.862 per cent, which is the about 3 per cent quoted later — so the two chromatic channels are 3 per cent and the red-green one is 0.4 per cent of that 3 per cent. Colour is the residual, and red-green is a residual of the residual.

The correlation is quoted a decimal short

The red-green channel is small because L and M correlate, and the two numbers do not quite fit.

For a pair correlating at r, the smaller of the pair’s own two principal components carries (1 − r)/2 of their joint variance. At the quoted 0.999 that is 0.05 per cent — four times the red-green channel’s measured 0.012 per cent of the whole.

Since the L and M cones carry very nearly the whole variance between them, the two should agree, and they do at a correlation of 0.9998 rather than 0.999. Quoting the correlation to three decimals understates the redundancy by a factor of four, and the fourth decimal is where the argument lives: 0.999 and 0.9998 are the same number to a reader and a factor of four to the channel built from their difference.

That is worth stating because the correlation is the essay’s mechanism and the variance is its finding, and a reader checking one against the other with the printed figures will not reproduce it. A quantity whose consequence is a difference needs one more digit than a quantity whose consequence is a sum, and this is the cleanest instance of that rule in the collection.

Sixteen degrees, not twenty-odd

The CIELAB section quotes the four unique hues at 25°, 88°, 161° and 250° against axes at 0, 90, 180 and 270, and calls the discrepancy twenty-odd degrees.

The four deviations are +25, −2, −19 and −20, so the mean absolute displacement is 16.5 degrees and one of the four — yellow, at 88 against 90 — is essentially on its axis.

So the failure is not uniform and it is worth knowing which axis works. Yellow sits where b*'s positive end is; blue sits twenty degrees off b*'s negative end; and the two ends of a* are off by 25 and 19 in opposite senses, so a* is not merely rotated — its two ends disagree about which way.

That is a more specific complaint than the axes are not the opponent channels. One half-axis of the four is right, and the a* axis is not a straight line through the unique hues at all — a rotation could fix a common offset and cannot fix two ends that need moving in opposite directions.

Where the model stops

This is an argument about why the recoding exists, and it is not a model of the retinal circuitry that implements it.

The real circuitry is not a clean matrix multiplication. Horizontal and amacrine cells build centre-surround receptive fields that mix spatial and chromatic opponency together, so the “red-green” cells in the primate retina are simultaneously edge detectors — a single cell’s response confounds a chromatic difference with a spatial one, and separating them is an unsolved problem in visual neuroscience.

The decorrelation argument also predicts the axes and not their weights in perception. The eigenvectors say which combinations are independent; they do not say how the visual system scales them afterwards, and it does scale them, substantially.

And the derivation depends on the input statistics. Different reflectance families give slightly different axes. What is robust is the structure — one achromatic channel dominating, an S-opposed channel second, an L−M channel far behind — and not the third decimal place of any weight.

What the pictures cannot show

The scatter plots show cone responses, which nobody experiences. There is no figure here of what an opponent channel looks like, because an opponent channel does not look like anything: it is a signal on a nerve, and what it contributes to is an experience that also involves the other two.

The bar chart is on a log scale, and that is a real distortion in service of legibility. On a linear scale the achromatic bar would fill the figure and the other two would be invisible — which is a truer picture of the proportions and a useless one. Any reader taking a visual impression of relative magnitude from that figure is reading a logarithm, and the numbers are printed for exactly that reason.

One palette under normal vision and two dichromacies. The same 7 colours simulated by the Brettel–Viénot–Mollon construction at severity 1.0. protanopia and deuteranopia collapse the red-green distinctions. This shows which discriminations survive, not what anybody sees.
Fig. 4 The second stage with one channel removed. Losing the L or M pigment collapses the difference the red-green channel is built from — the channel already carrying least, which is why its loss is survivable and why it is the common one. The handle runs the severity, so the collapse from three cones to two is a slope rather than a switch.
One palette under normal vision and tritanopia. The same 7 colours simulated by the Brettel–Viénot–Mollon construction at severity 1.0. Tritanopia leaves the red-green distinctions and collapses blue against yellow instead. This shows which discriminations survive, not what anybody sees.
Fig. 5 And with the other channel removed instead. Losing the short-wave pigment collapses the blue-yellow axis and leaves the red-green one, which is the same statement about the same recoding from the other side.

The severity of the loss is a dial rather than a switch, and most people who have one of these are somewhere along it.

One palette under normal vision and three dichromacies. The same 7 colours simulated by the Brettel–Viénot–Mollon construction at severity 0.5. protanopia and deuteranopia collapse the red-green distinctions, and tritanopia leaves them and collapses blue against yellow instead. This shows which discriminations survive, not what anybody sees.
Fig. 6 All three at half severity, which is the anomalous rather than the dichromatic case and is far commoner. An axis compressed rather than removed is what a third pigment shifted a few nanometres produces.
One palette under normal vision and deuteranopia. The same 7 colours simulated by the Brettel–Viénot–Mollon construction at severity 1.0. deuteranopia collapses the red-green distinctions. This shows which discriminations survive, not what anybody sees.
Fig. 7 The commonest form on its own, for scale: about five per cent of men against under one for the other two together. Two axes is a statement about the population as well as about the arithmetic.

All three collapses together on the palette this site checks its own figures against are what say the two axes are two rather than one or three.

One palette under normal vision and three dichromacies. The same 7 colours simulated by the Brettel–Viénot–Mollon construction at severity 1.0. protanopia and deuteranopia collapse the red-green distinctions, and tritanopia leaves them and collapses blue against yellow instead. This shows which discriminations survive, not what anybody sees.
Fig. 8 All three collapses at full severity on the palette this site checks its own figures against. Three cones make two chromatic axes; two cones make one, and which one depends on which pigment is missing.

Who found it, and when

Hering published his opponent account in the 1870s, from phenomenology — the impossible colours, the elementary hues, the afterimages — and was largely dismissed by the physiologically-minded majority.

Hurvich and Jameson vindicated him in the 1950s with hue cancellation: to measure how much redness a stimulus contains, let an observer add green until the redness is exactly cancelled, and the amount added is the measure. It converts an introspective report into a null measurement, which is the same trick colour matching uses and is why both are quantitative.

Svaetichin recorded from horizontal cells in fish retina in 1956 and found cells responding with opposite sign to different wavelengths — the first direct physiological evidence, and found in a fish because primate retina is considerably harder to record from.

Buchsbaum and Gottschalk gave the information-theoretic derivation in 1983. Ruderman, Cronin and Chiao extended it to natural images in 1998 and established the variance ordering — including the finding that the L−M axis carries least, which remains the least-quoted robust result in the field.

Where this goes next

The four hues the two chromatic channels produce, and the awkward question of where exactly they sit, is why there are four unique hues. The receptors feeding all of this are three numbers. And the dimension count that this stage preserves while redistributing what each dimension carries is why colour is exactly three-dimensional.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 16 that link here.

The objects this essay names

Each one links to every other essay that touches it.

CIELABCone correlationDecorrelationEfficient codingHering's opponent theoryOpponent processingReflectanceSecond stageUnique hues