The mosaic is not the observer
Assumes Three numbers and Nothing is in focus at both ends.
Two people with entirely normal colour vision can have retinas whose long-wavelength and medium-wavelength cone populations differ by a factor of sixteen. One has roughly equal numbers; the other has sixteen L cones for every M. Both pass every clinical test. Both agree about which shirt matches which tie.
That is a startling amount of variation in the composition of an organ, and it is startling in a specific way: it should, on any naive account, wreck colour matching entirely.
The three arguments of that picture are the ratio, the patch of retina it is drawn from, and the observer the fundamentals come from, and each moves the answer differently.
A larger patch of retina is what makes any of this reproducible, and it is worth seeing how quickly the arrangement stops mattering.
Two more draws say that the arrangement washes out with area and that the range of ratios is wider than the range anybody’s colorimetry differs over.
Why it should wreck matching, and does not
A colour match is the claim that two spectra are indistinguishable. Physiologically, that means they produce equal excitation in each of the three cone classes: the L excitation from one equals the L excitation from the other, and likewise for M and S.
Now change the number of L cones. The total L signal leaving the retina scales — twice as many L cones, twice the signal, other things being equal. Surely that changes what matches what.
It does not, and the reason is a single line of algebra. Scaling both sides of an equality by the same positive constant leaves the equality intact. If
then for any positive ,
The matches cannot move. Not approximately, not to within experimental error — exactly, for any composition whatever, including compositions no human has.
This is checked here rather than asserted. Two spectra are constructed and scaled until they agree in cone space, and then the L:M ratio is swept over 1, 2, 4, 8 and 16 to one. The relation between the two spectra in each cone class agrees across the whole sweep to 10⁻¹², which is floating-point noise. The assertion demands exactness rather than a tolerance, deliberately: a loose tolerance would hide an implementation in which the weights had accidentally been made to interact.
The quantity that does move
An invariance claim is worth nothing without a neighbouring quantity that is not invariant. If everything survived the sweep, the computation would be testing arithmetic rather than vision.
Luminance is the quantity that does not survive. It is not an equality between channels; it is a weighted sum of them — very nearly , with the weights set by how many of each there are. Change the ratio and the sum changes.
Measured on a red spectrum, where the L and M contributions differ most: luminance is 1.33 times greater at L:M 16:1 than at 1:1. A third more, from nothing but a difference in retinal composition, on a stimulus both observers would agree matches the same reference.
That asymmetry is the whole content of this essay, and it explains a fact usually reported as an annoyance. V(λ) has a much larger between-observer variance than the colour-matching functions do. Every practitioner knows this; photometry is understood to be a coarser business than colorimetry, and heterochromatic brightness matching is understood to be unreliable in a way that colour matching is not. The reason is not that brightness judgements are harder. It is that luminance is a sum and matching is an equality, and only one of those is invariant to what the retina is made of.
That overlap is worth dwelling on, because it is what makes the ratio variation tolerable in practice as well as in principle. If L and M sampled disjoint parts of the spectrum, changing their relative numbers would change which spectral regions the retina was well and badly served in, and although matching would still be invariant by the argument above, almost everything else about vision would differ between the two observers. Because they overlap almost completely, trading one for the other barely changes what the retina samples; it changes only how the two nearly-identical samples are weighted against each other.
The system that cares about that weighting is the red-green opponent channel, which takes the difference. And the difference of two heavily overlapping signals is a small quantity computed from two large ones — which is why the red-green channel carries about one part in ten thousand of the variance in natural scenes, and why it is so sensitive to exactly the composition this essay says matching ignores.
What a standard observer can and cannot standardise
The 1931 system is a set of three functions and a claim that they describe everybody well enough. The result above says precisely which part of that claim is cheap and which is expensive.
The shapes of the cone fundamentals — where each class is sensitive, as a function of wavelength — are what colour matching depends on. Those are set by the photopigments, and photopigments are gene products with very little room to vary: an L pigment is an L pigment, give or take the well-documented serine/alanine polymorphism that shifts its peak by a few nanometres. So standardising the shapes is standardising something genuinely stable.
The amounts — how many of each class, and therefore the weights in every sum — vary enormously and are standardised anyway, because is a specific weighted combination and the CIE had to choose one. That choice is a population average of something with a sixteen-fold spread.
So the standard observer is on firm ground for matching and on much softer ground for anything photometric, and the two are usually quoted with the same air of authority. This is the same distinction, arriving from the retina, that separates what an observer disagreement costs a match from what it costs a measurement.
What varies, and what does not, in one table
Setting the two out side by side makes the shape of the result visible:
| Quantity | Depends on | Varies between people |
|---|---|---|
| A colour match | the shapes of the fundamentals | barely — only through pigment polymorphism and pre-receptoral filtering |
| Tristimulus values | the shapes, again | barely |
| Luminance, and V(λ) | the shapes and the weights | substantially — 1.33× across the measured L:M range |
| Photometric quantities generally | the same | the same |
| Anything opponent | the difference of two weighted signals | most of all |
The pattern is that equalities are robust and sums are not, and that the further a quantity is from being an equality between channels, the less a standard observer standardises it. That is a more useful rule than the usual framing, which treats the standard observer as uniformly authoritative and then apologises separately for each place it turns out not to be.
The mosaics are seeded, so the honest check is to draw a larger patch from a different seed and confirm that nothing about the result was an accident of the arrangement.
Where the invariance stops
Three places, and the third is the interesting one.
It stops at the pigments. The invariance is to how many cones of each class there are, not to what they are sensitive to. An observer whose L pigment peaks at 559 nm rather than 563 nm has a genuinely different fundamental and genuinely different matches, which is observer metamerism and is a real and measurable problem.
It stops at pre-receptoral filtering. Macular pigment density varies by a factor of several between people and is a filter in front of the cones, so it changes the effective fundamental rather than its weight. This is why the 2° and 10° functions differ most in the blue, and why the variation between individuals is worst there too.
It stops at anything nonlinear. The algebra above holds because the cone excitations are linear in the spectrum and the match is an equality. Everything downstream of the receptors is nonlinear — the opponent recoding, the compressive response, the whole appearance model — and none of it inherits the invariance. Two observers with different L:M ratios agree that two patches match, and there is no guarantee whatever that they agree about how bright, how colourful, or how contrasty the matched pair looks. The invariance covers exactly the layer colorimetry stops at, and not one step past it.
That is a sharper statement of this site’s standing caution than the caution usually gets. Matching is not appearance is normally argued from surround effects and adaptation. Here it arrives from the mosaic: the very fact that makes matching robust is a fact about an equality, and appearance is not an equality.
A narrower range of compositions is the fairer test of the claim, since a nine-fold difference is nearer what a population actually spans than a sixteen-fold one.
The practical shape of it
A laboratory that needs two instruments, or two people, to agree about a colour has a choice about what to ask them to agree on, and this result says which choices are cheap.
Asking whether two samples match is cheap. Any two normal observers will agree, and the agreement does not depend on their retinal composition at all — which is why visual matching remained the reference method long after instruments existed, and why a visual match is still the arbiter in some industries.
Asking which of two samples is brighter, when they differ in colour, is expensive. Heterochromatic brightness matching has famously poor inter-observer agreement, has never been made additive, and is the reason the CIE has several different photometric procedures that give different answers. None of that is a failure of experimental technique. It is a sum being asked to do the work of an equality.
Asking by how much two samples differ is somewhere between, and depends on the formula: a difference metric built on an opponent representation inherits the opponent channel’s exposure to composition, which is one of several reasons a threshold is not a unit.
What was computed, and how
Cone excitation is linear in the spectrum, and the transform from XYZ to LMS is linear in XYZ, so integrating a spectrum against the fundamentals is the same computation as taking its tristimulus values through the cone matrix. The second route is the one used, because it reuses machinery the rest of the site is already checked on rather than introducing a second copy of the fundamentals on a different wavelength grid.
The two test spectra are built rather than stored — a single broad Gaussian, and a pair of narrower ones either side of it — and then one is scaled until the two agree in the L channel. They do not agree in all three, and that is deliberate: a two-parameter family cannot generally match three excitations, and constructing a pair that did would have required machinery whose correctness is a separate question. What is asserted is not that the pair matches but that whatever the two disagree about is unchanged by the sweep, which is the invariance, stated in a form that does not depend on having built a perfect metamer.
The mosaic pictures are drawn from a seeded generator, on a hexagonal packing rather than a square lattice, because the cone mosaic is hexagonal and a square one would be a picture of a sensor. The seed is fixed so the figure does not churn the built output on every rebuild.
Who found it, and when
That the L:M ratio varies widely was established by two independent routes in the 1980s and 1990s. Electroretinography and flicker photometry gave population estimates with a large spread; adaptive-optics imaging of living retinas, from the late 1990s, made the individual mosaics directly visible and confirmed ratios from about 1:1 to about 16:1 in people with normal colour vision.
The invariance argument is older than the measurement and is usually credited to the structure of Grassmann’s laws rather than to any one person: if colour matching is additive and the receptors are linear filters, the matches depend on the filters and not on their gains. It is the kind of result that is obvious once stated and was not stated for a long time, because the natural question — “does the ratio matter?” — sounds empirical and is largely algebraic.
Hofer, Carroll and Williams’s imaging work in the 2000s added the detail that makes the story land: the observers with the most extreme ratios were not identifiable by any colour test, and did not know.
The 1931 experiment sits oddly in this history. Its seventeen observers were selected, matched and averaged with no knowledge of any of this — the photopigments would not be sequenced for another half century — and the average it produced turned out to be robust for the quantity it was averaging and much less robust for the quantity derived from it. That is a lucky outcome rather than a designed one, and it is worth noticing that the experiment’s own record does not claim otherwise: the 1931 committee was explicit that the functions represented an average and said little about how much individuals departed from it, because they had no way to find out.
What the picture cannot show
The three mosaics above are drawn at the right proportions and the wrong everything else. They are hexagonal, which is right; they are the right relative numbers of L, M and S, which is the point; and the arrangement within each is a seeded random draw, which is wrong in a specific and known way.
Real cone mosaics are not randomly arranged. The S cones sit on a semi-regular sub-lattice with a characteristic spacing, and the L and M cones show clumping — patches of same-type cones larger than chance would produce. Neither is captured here, because neither is needed for the argument and drawing them would imply a level of fidelity the figure does not have.
The colours are also a convention rather than a claim. Nothing about an L cone is red. The three fills are the site’s three cone roles, used consistently across every figure that distinguishes the classes, and a reader who reads them as the colours those cones “see” has read something the figure did not say — the whole difficulty of drawing colour on the apparatus it is about arriving in miniature.
How far the luminance figure could have gone
The 1.33 measured on a red spectrum is the essay’s one non-invariant number, and it is worth asking what its range is, because a single value from a single test spectrum says nothing about whether the effect has been seen near its worst.
Writing the essay’s own account of luminance as arithmetic — a sum of the two long-wave signals weighted by how many of each cone there are, normalised so the total cone count is fixed — gives, for a ratio r of L to M,
and the quantity the sweep reports is V(16)/V(1). That depends on one property of the test spectrum only: the ratio of the L excitation it produces to the M excitation.
| L/M excitation of the stimulus | V(16) / V(1) |
|---|---|
| 0.25 | 0.47 |
| 0.50 | 0.71 |
| 1.00 | 1.000 |
| 2.20 | 1.33 — the published figure |
| 4.00 | 1.53 |
| 16.0 | 1.78 |
| → ∞ | 1.882 |
Three readings come out of that column.
The published 1.33 corresponds to a stimulus whose L excitation is 2.2 times its M excitation, which is a moderately red light rather than an extreme one. The red spectrum was chosen as the one where the two contributions differ most among the spectra tested, and it is not near the limit of what a spectrum can do.
The ceiling is 1.882. A stimulus exciting only the long-wave cones — a deep red monochromatic light, near enough — moves in luminance by 88 per cent between the two extreme observers, not 33. The published figure is 37 per cent of the way from no effect to the largest available one.
And the whole span is a factor of exactly sixteen, from 0.118 at the other extreme to 1.882, which it must be: at the limits only one cone class is being counted, so the luminance ratio becomes the cone ratio. That the closed form reproduces the input range is the check that the model behind the table is the essay’s own rather than a different one.
The invariance is a continuum, not a wall
The third row of that table is the one worth carrying, and the essay’s framing does not have a place for it.
At L/M = 1 the luminance ratio is exactly 1.000. A stimulus that excites the two long-wave cone classes equally has a luminance that does not depend on the mosaic at all — as invariant as a colour match is, and for a reason as exact.
So the division the essay draws — matches are equalities and therefore robust, luminance is a sum and therefore not — is right about the mechanism and describes the extremes of a continuum rather than two categories. A weighted sum is invariant to its weights whenever the things being summed are equal, and the exposure grows smoothly with the imbalance: 29 per cent at L/M = 2, 53 per cent at 4, 78 per cent at 16.
The practical form is more useful than the categorical one. A photometric measurement’s exposure to who is making it is not a fixed penalty attached to photometry; it is proportional to how far the stimulus is from balancing the two long-wave cones. A near-neutral or slightly greenish source is measured almost as reliably as a match is made. A saturated red or a narrow long-wavelength emitter is where the standard observer’s weights are doing the most work and are least entitled to.
That also sharpens the closing advice. Asking which of two samples is brighter, when they differ in colour, is expensive is true; how expensive depends on which colours, and the cost runs from nothing to a factor of nearly two across the range one population of ordinary eyes supplies.
What the opponent numbers add
The variance figures quoted from the opponent decomposition — 97.1 per cent achromatic, 2.85 per cent blue–yellow, 0.012 per cent red–green — give the same story a third scale.
The red–green channel carries one part in 8,300 of the variance, and the achromatic channel carries 8,100 times as much as it does. The channel most exposed to the mosaic is the one with the least signal in it, by nearly four orders of magnitude, which is why it is both the most sensitive and the least statistically robust thing the retina computes.
That is the same arithmetic as the table above, read from the other end. A difference of two nearly identical signals is small and is dominated by whatever makes them differ, and the number of cones of each class is one of the things that does.
Where this goes next
Upward, this ladder reaches the question the whole site keeps returning to — what the observer choice actually costs, now with an answer that distinguishes the quantities it costs nothing from the quantities it costs a third.
Downward, the eye’s own optics explain why the third cone class can be sparse, and the collapse from a spectrum to three numbers is the layer at which all of this invariance lives and above which none of it does.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- One match names the observer colour-matching functions · cone fundamentals · individual variation · observer metamerism · standard observer · trichromacy
- The laws that make colour add up colour-matching functions · cone fundamentals · individual variation · luminance · standard observer · trichromacy
- Four primaries have a choice colour-matching functions · cone fundamentals · individual variation · observer metamerism · standard observer
- A fourth primary is a design colour-matching · individual variation · observer metamerism · standard observer
- A gamut has a population cone fundamentals · individual variation · observer metamerism · standard observer
- The matches do not name the cones colour-matching functions · cone fundamentals · standard observer · trichromacy
What links here
The 8 essays that link to this one and share the most of its objects, of 12 that link here.
The objects this essay names
Each one links to every other essay that touches it.
Colour-matchingColour-matching functionsCone fundamentalsIndividual variationLuminanceLuminous efficiencyObserver metamerismStandard observerTrichromacy