Where the model breaks

Whose eyes

The standard observer is an average over seventeen people, and no reader is it. What that costs was small when displays were broad and grows every time the primaries get narrower.

Assumes Seventeen observers in 1931 and Two spectra, one colour.

Every calculation on this site runs through a standard observer, and a standard observer is nobody. It is an average over seventeen people measured in the 1920s, and the whole apparatus of colour matching rests on the assumption that the average is close enough to everyone.

How close is computable, and the answer has been getting worse.

How far a match comes apart when the observer changes. A broad source and a three-primary source, solved at each primary width so the pair is an exact tristimulus match for the CIE 1931 observer. The pair is then handed to the 1964 observer, and the gap between them is plotted. For the observer they were built for the gap is arithmetic noise at every width. For the other it grows as the primaries narrow, reaching 0.012 at 10 nm — and displays have been getting narrower for twenty years.
Fig. 1 A broad source and a three-primary source, solved at each primary width so the pair is an exact tristimulus match for one observer, then handed to another. For the observer they were built for the gap is arithmetic noise at every width. For the other it grows steadily as the primaries narrow.

Two standard observers is the smallest population anybody has, and the same measurement can be taken from either end of it.

How far a match comes apart when the observer changes. A broad source and a three-primary source, solved at each primary width so the pair is an exact tristimulus match for the CIE 1964 observer. The pair is then handed to the 1931 observer, and the gap between them is plotted. For the observer they were built for the gap is arithmetic noise at every width. For the other it grows as the primaries narrow, reaching 0.013 at 10 nm — and displays have been getting narrower for twenty years.
Fig. 2 The same pairs scored with the ten-degree observer as the reference and the two-degree one as the deviant. Neither of them is a person, and the disagreement between them is a lower bound on what a real population does.
How far a match comes apart when the observer changes. A broad source and a three-primary source, solved at each primary width so the pair is an exact tristimulus match for the CIE 1931 observer. The pair is then handed to the 1964 observer, and the gap between them is plotted. For the observer they were built for the gap is arithmetic noise at every width. For the other it grows as the primaries narrow, reaching 0.012 at 10 nm — and displays have been getting narrower for twenty years.
Fig. 3 And across the three widths a display technology actually spans. The trend the essay is about is entirely inside the range somebody is choosing between when they specify a screen.

The generator’s own default sampling is the third reading of the same pair, and it is the one this collection quotes in its other essays.

How far a match comes apart when the observer changes. A broad source and a three-primary source, solved at each primary width so the pair is an exact tristimulus match for the CIE 1931 observer. The pair is then handed to the 1964 observer, and the gap between them is plotted. For the observer they were built for the gap is arithmetic noise at every width. For the other it grows as the primaries narrow, reaching 0.012 at 10 nm — and displays have been getting narrower for twenty years.
Fig. 4 The same pair at the default sampling this collection quotes elsewhere. Two standard observers is the smallest population anybody has, and it is already too large for one number to describe.

Why a match can fail for somebody else

Metamerism is two different spectra producing the same three numbers. The three numbers are three integrals of the spectrum against three matching functions — so a metameric match is an agreement about three particular integrals, and nothing more.

Change the matching functions and the integrals change. Two spectra that agreed under one set need not agree under another, and the amount by which they come apart depends on how much they differed spectrally in the first place.

That is the whole mechanism. A pair of spectra that are nearly identical will match for anybody; a pair that agree on three integrals while differing wildly in between will match for exactly the observer whose functions were used and progressively worse for everybody else.

The measurement

The CIE publishes two observers — the 1931 two-degree functions and the 1964 ten-degree ones — which makes the calculation direct. Build a pair that matches exactly under one and evaluate it under the other.

The pair here is a broad smooth source and a three-primary source, which is the case that matters: it is a printed page against a display, or an old display against a new one. The three primary powers are solved so the tristimulus values match the broad source exactly under the 1931 observer — an exact solve, not a fit, and the residual is at the level of floating-point noise.

Handed to the 1964 observer, the same pair comes apart. And the gap grows monotonically as the primaries narrow: 0.0028 at 60 nm primaries, rising to 0.0125 at 10 nm. A factor of 4.5 across the range, in one direction, at every step.

Why narrow primaries make it worse

The reason is geometric and worth seeing clearly.

A broad primary integrates the matching function over a wide band, so it samples an average of it. Two observers whose functions differ in shape but agree roughly in area will produce similar integrals — the differences average out across the band.

A narrow primary samples the matching function at essentially one wavelength. There is nothing to average, so any difference between two observers at that wavelength passes straight through into the tristimulus value. Narrowing the primary concentrates the observer difference instead of diluting it.

And the effect is worst where the observers differ most, which is the short-wavelength end — because the lens and the macular pigment absorb in the blue, and both vary substantially between people. A narrow blue primary samples exactly the region where two observers are least likely to agree.

The direction the technology has gone

This would be a curiosity if displays were getting broader. They are not.

A cathode-ray tube’s phosphors were broad, and its gamut was correspondingly small. Every technology since has widened the gamut, and the only way to widen a gamut is to make the primaries more saturated, and the only way to make a primary more saturated is to make it narrower — which is what a wider gamut costs.

An LED-backlit panel with a phosphor-converted white backlight has moderately broad primaries. A quantum-dot backlight has narrower ones. An OLED display with emissive materials is narrower again, and a laser projector is narrowest of all — genuinely monochromatic, which is the limiting case in both directions at once.

So the industry has spent thirty years increasing the very quantity that makes observer variation matter, in pursuit of a gamut increase that is real and is a different property. Nobody made a mistake. The two effects are simply not connected in anybody’s specification, and one of them appears on the box.

The same solve runs wider at the broad end too, out to an eighty-nanometre phosphor of the kind a cathode-ray tube actually had, which puts the whole of that thirty-year progression on one axis.

How far a match comes apart when the observer changes. A broad source and a three-primary source, solved at each primary width so the pair is an exact tristimulus match for the CIE 1931 observer. The pair is then handed to the 1964 observer, and the gap between them is plotted. For the observer they were built for the gap is arithmetic noise at every width. For the other it grows as the primaries narrow, reaching 0.013 at 8 nm — and displays have been getting narrower for twenty years.
Fig. 5 Six widths from a CRT phosphor at one end to something close to a quantum dot at the other, with the gap reaching 0.013 at 8 nm. The control line — the observer the pair was built for — sits at arithmetic noise across the whole range, which is what makes the rise a property of the second observer rather than of the solve.

Where the variation actually comes from

The intuition is that people differ in their brains. Almost none of the difference is neural, and the largest sources are in front of the retina.

The lens absorbs increasingly toward short wavelengths and yellows steadily throughout life. An eighty-year-old lens transmits roughly a third of the 440 nm light a twenty-year-old lens does. This is the single largest source of observer difference, it is entirely pre-neural, and it is continuous — every person is a slightly different observer every year.

The macular pigment is a carotenoid screen over the central retina peaking near 460 nm, with optical density varying between individuals by a factor of three, for reasons that are partly dietary.

The photopigments vary too, and the L pigment in particular comes in common variants differing by a few nanometres in peak wavelength — a genuine polymorphism, present in a substantial fraction of the population, which produces two subtly different classes of normal trichromat.

The uncomfortable summary is that a standard observer has no age, no macular density, and one version of each pigment. It describes a hypothetical person of no particular years, assembled from seventeen who each had some.

What the two CIE observers are, and are not

Using the 1931 and 1964 functions as “two observers” is the calculation this essay does, and it is worth being precise about what it establishes.

They are not two people. The 1964 functions were measured with a ten-degree field rather than a two-degree one, so the difference between them is dominated by field size rather than by individual variation: a larger field extends beyond the macula, so the macular pigment’s contribution falls, and rod intrusion becomes possible.

So the pair are two averages, differing for a known structural reason. What the calculation establishes is the magnitude of the effect when two plausible sets of matching functions differ by an amount comparable to real observer variation — which is a defensible use of them, and is not the same as measuring how far apart two individuals are.

The CIE published a model of individual observer variation in 2006, parameterised by age and field size, which is the right tool for the actual question. This site does not implement it, and the essay’s numbers should be read as an illustration of the mechanism at a realistic scale rather than as a population statistic.

Extending the sweep to a ninety-nanometre primary at one end and nine at the other spans everything the technology has ever shipped.

How far a match comes apart when the observer changes. A broad source and a three-primary source, solved at each primary width so the pair is an exact tristimulus match for the CIE 1931 observer. The pair is then handed to the 1964 observer, and the gap between them is plotted. For the observer they were built for the gap is arithmetic noise at every width. For the other it grows as the primaries narrow, reaching 0.013 at 9 nm — and displays have been getting narrower for twenty years.
Fig. 6 Six widths from a phosphor broader than any CRT’s to something close to a laser. The control line stays at arithmetic noise across the whole range and the other rises monotonically, which is the trend the technology has spent thirty years moving along.

The other thing this makes worse

Observer variation is not only a problem for matches between different technologies. It is a problem for anything that assumes two people looking at the same thing agree.

Colour matching in production assumes an operator’s judgement is transferable. Two operators with different lens ages and different macular densities, comparing a print against a display, will disagree — and the disagreement will be largest in the blues and will be attributed to the equipment.

Consumer displays with wide primaries and per-unit calibration are calibrated against an instrument, which implements the standard observer. A calibrated display is therefore correct for a hypothetical average person and slightly wrong for the owner, in a direction that depends on the owner’s age.

And simulating what somebody else sees inherits all of this. A colour vision deficiency simulation assumes a standard observer as its starting point, so it is showing what a hypothetical average person with a hypothetical average deficiency would see — two idealisations stacked, each with a spread comparable to the effect being shown.

What was computed here

The primaries are Gaussian emitters at stated peaks and widths, and their powers are solved by Cramer’s rule so the resulting tristimulus values match the broad source’s exactly under the reference observer. An exact solve rather than an optimisation, because the essay’s claim is that the pair matches exactly for one observer and comes apart for another, and an approximate match would leave a residual indistinguishable from the effect.

The residual is checked: under the reference observer the gap is under 10⁻¹⁵ at every width, which is the “matched by construction” line in the figure and is the control the whole comparison rests on.

Three things are asserted. The pair must match for the observer it was built for, at every width. It must come apart for the other by at least an order of magnitude more. And — the claim in the caption, and the one that matters for the technology argument — the mismatch must grow monotonically as the primaries narrow, by at least a factor of three across the range. That last is checked step by step rather than at the endpoints, because a curve with the right endpoints and a dip in the middle would tell a different story about the trend.

The trend has to stop, and probably already has

The six widths are a geometric ladder — 60, 40, 28, 20, 14, 10, at a ratio of about 1.43 each step — and the mismatch across them fits a power law: the gap goes as the width to the power −0.84, giving 0.0028 at sixty nanometres and 0.0125 at ten.

Extrapolating that law is what the technology argument invites and it is not available. At five nanometres it would predict 0.022 and at one nanometre 0.086, and neither can happen, because a primary of zero width is a delta function and the mismatch it produces is the difference between the two observers’ matching functions at three fixed wavelengths — a finite number.

So the curve must bend. There is a worst possible primary and it is a monochromatic one, and the worst case is bounded. That matters for the essay’s own strongest claim: a laser projector is the limiting case in the sense of being the narrowest, and it is not the limiting case in the sense of being unboundedly bad.

Where the bend is can be estimated from what is being averaged. A primary of width w averages the two observers’ difference over w, and that averaging only helps while w is comparable to the scale on which the difference varies — which is tens of nanometres, since it is dominated by the lens and the macular pigment, both broad absorbers. Below about ten or fifteen nanometres there is nothing left to average, so the ten-nanometre point is probably already most of the way to the plateau, and the step from an OLED to a laser is a much smaller step than the step from a cathode-ray tube to an OLED.

That is a more useful thing to tell a manufacturer than narrower is worse. It says the damage was done between sixty and twenty nanometres, that it is nearly complete, and that the remaining technologies differ from each other far less than any of them differs from what came before.

The lens claim, checked against this collection’s own model

An eighty-year-old lens transmits roughly a third of the 440 nm light a twenty-year-old lens does is the essay’s one quantitative claim about the mechanism, and it reproduces exactly from the ocular-media model the rest of the collection uses — a density of 0.5 rising by 0.02 a year past twenty, on an exponential in wavelength with a 68-nanometre scale.

At 440 nanometres that model gives 0.621 at twenty and 0.198 at eighty, a ratio of 0.319 — a third, to three decimals.

What the single number hides is how fast the ratio moves:

wavelength age 20 age 80 ratio
400 nm 0.424 0.054 0.128
440 nm 0.621 0.198 0.319
460 nm 0.701 0.299 0.427
500 nm 0.821 0.512 0.623

The age effect changes by a factor of 3.3 between 400 and 460 nanometres, which is the width of the region a blue primary can sit in. So the effect is worst at the short-wavelength end is true and under-specified: where a narrow blue primary sits inside the blue matters about as much as how narrow it is.

A display whose blue primary peaks at 460 gives an eighty-year-old 43 per cent of what it gives a twenty-year-old; one at 440 gives 32 per cent; one at 420 gives 22. Those are three ordinary design choices, all of them made for gamut and efficiency, and they differ by a factor of two in how much the oldest and youngest viewers of the same panel disagree.

And the two levers point the same way. A narrower blue primary is usually also a shorter one, since narrowing towards saturation moves a primary outward along the locus and the locus’s blue arm runs towards 400. So the technology trend concentrates the sample and moves it into the steepest part of the lens curve at the same time, which is why the blues are where the complaints are.

Where the model stops

The largest limitation is the one above: two CIE observers differing mainly by field size are a proxy for two individuals differing by age and pigment density. The mechanism is the same and the magnitudes are illustrative.

The pigment template can generate shifted observers — that is what it is for, and the parameter is there — but turning a shifted pigment into a full set of matching functions requires the transformation from cone fundamentals to XYZ, and doing that properly for a non-standard observer is a larger piece of work than this essay needed. What is drawn here is the pigments; what is computed is the two-observer case.

There is also no rod contribution. The 1964 ten-degree observer is measured at a field size where rods are present and can intrude at lower luminances, and the standard does not model that. The mesopic range is a different failure of the standard observer in the same direction: it describes a fully light-adapted person, permanently.

The observer this site actually assumes

Worth stating plainly, because it is a stronger assumption than any single essay makes.

Every colour on this site is computed through the CIE 1931 two-degree observer unless a figure says otherwise. That observer is fully light-adapted, of no particular age, with one version of each photopigment, viewing a two-degree field. The reader is none of those things, and the mismatch is largest in the blues, on wide-gamut displays, for older readers.

The site’s response is the one available: name the observer in the corner of every figure. That does not correct anything. What it does is make the assumption visible, so a reader who knows their own vision differs knows which way and by roughly how much.

What the pictures cannot show

The figure plots a tristimulus gap, which is not something anybody experiences.

What it would take to show the effect is two observers, which is exactly the thing a page cannot supply. Every reader sees these figures through their own lens, their own macular pigment, and their own pigment variants, and there is no way for the page to know or compensate for any of them.

There is also a self-referential difficulty that runs deeper here than elsewhere on this site. The essay’s subject is that the standard observer is not the reader — and every colour on the page was computed through the standard observer. The swatches are wrong for the reader by exactly the amount the essay is about, and there is no version of the page that would not be.

That is not a reason to draw nothing. It is a reason the figures plot numbers rather than showing two versions of a swatch and inviting a comparison the apparatus cannot support.

What could be done about it

Three responses exist, and their ordering by how often they are used is roughly the reverse of their ordering by how well they work.

Use a wider observer. The 1964 ten-degree functions are a better standard for most practical work than the 1931 two-degree ones, since most colour judgements are made on samples larger than a thumbnail at arm’s length. They are used far less, for the same reasons the 1976 chromaticity diagram is used less than the 1931 one: the default was set long ago and nothing forces a change.

Use an observer model. The CIE’s 2006 physiological observer is parameterised by age and field size and can generate matching functions for a stated individual. It is the right tool and it requires knowing the individual, which a display manufacturer does not.

Broaden the primaries. This works, it directly reduces the effect, and it costs gamut — so it is the opposite of what the market rewards, and nobody does it.

The reason none of the three is happening is that observer metamerism has no owner. A display manufacturer’s product is measured with an instrument implementing the standard observer, and by that measurement it is correct. The error appears only in a person, is different for every person, and shows up as a diffuse complaint about a display looking odd — which is attributed to the panel, the calibration, or taste.

Who found it, and when

The 1931 functions were assembled from Wright’s ten observers and Guild’s seven, using different apparatus at different institutions, and combined into a single average. The observers were young, they were male, and there were seventeen of them.

That the average would not describe everybody was understood at the time and treated as acceptable, correctly, for the purpose: the system was an industrial specification for communicating colours between manufacturers, and an average was exactly what was wanted.

The 1964 ten-degree observer was added because a two-degree field is small — about the width of a thumbnail at arm’s length — and much industrial colour matching is done on larger samples. It is a better standard for most practical work and is used far less than the 1931 one, for the same reasons the 1976 chromaticity diagram is used less than the 1931 one.

Observer metamerism as a practical problem arrived with wide-gamut displays in the 2000s and became urgent with laser projection, where the primaries are genuinely monochromatic and different viewers in the same cinema disagree about the colour of the screen. That is the clearest demonstration the effect has ever had, and it is a demonstration nobody wanted.

Where this goes next

The seventeen people the standard is made of are seventeen observers in 1931. The phenomenon this depends on is two spectra, one colour. And the other place a standard observer stops describing the reader is the eye that has no colour.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 55 that link here.

The objects this essay names

Each one links to every other essay that touches it.

GamutIndividual variationLens yellowingMacular pigmentMetamerismNarrow band displaysObserver metamerismPrimariesStandard observer