Two degrees or ten
Assumes Whose eyes and Seventeen observers in 1931.
The first figure rule on this site is name the observer. A chromaticity diagram is a picture of an integral, the integral has a kernel somebody chose, and a diagram that does not say which kernel is incomplete. Every figure here carries it in the corner, and has since the site’s foundation phase.
At the start of this phase that rule was measured rather than assumed, by rendering all 277 figure placements and reading the caption strips back out of the finished SVG. The result:
- 277 of 277 placements carried a strip. The rule, satisfied outright.
- 204 of them named the CIE 1931 2° observer.
- 3 named the 1964 10° observer, and all three were in one essay.
- 2 of 64 generators had ever been called with more than one distinct note.
The comparison has a direction and a sampling, and neither of them changes what it says.
How finely the range is sampled is the other free choice, and the part of it a display technology actually spans is narrower than the sweep above.
A rule kept, and a point unmade
Sixty-two generators printed the same string on every placement they had ever had. A label that never varies carries no information — and this one is the site’s own thesis. The observer is a choice, and the site had made it once and defaulted to it two hundred and four times.
Nothing reported this and nothing could have. Every check this collection runs on a figure asks whether a label fits, contrasts and stays inside the viewBox. The parameterisation floor asks whether a placement passes the generator any argument, and 96% of these did — they pass the wavelength, the illuminant, the tolerance, and take the observer as it comes. The one property nobody measured was whether the strip ever said anything different.
That is the shape of defect worth naming, because it is general: a caption strip is checked everywhere for presence and nowhere for content. A figure can be correct, well-labelled, contrast-checked, inside its viewBox, and quietly making no claim at all.
The repair is the observer index, which sorts every placement by what its strip names and is built by rendering each figure and reading the string back out of the emitted SVG rather than out of the source that wrote it. The two can differ, and where they do it is the picture that is right.
What the two observers are
The 1931 functions come from matching experiments on a two-degree field — about the width of a thumbnail at arm’s length, which is roughly the foveal region. The 1964 functions come from a ten-degree field, and are not the 1931 experiment repeated more carefully. They are a different experiment on a different part of the retina.
Three things differ across those two fields, and all three push in the same direction.
Macular pigment. A yellow carotenoid screen lies over the central retina, absorbing strongly in the blue. A two-degree field looks through the thick of it; a ten-degree field averages over the periphery where it thins out. So the 10° observer is substantially more sensitive in the short wavelengths.
Rod intrusion. There are no rods in the foveola and plenty at ten degrees. The matching experiments were run at levels intended to be photopic, and the 10° functions still carry a rod contribution that the 2° functions do not.
Cone composition. The relative numbers of the three classes vary with eccentricity, S cones being absent from the very centre and present outside it.
The upshot is that the disagreement is concentrated in the blue, and this site keeps discovering that from different directions. The eye’s own dispersion is worst in the blue. Individual variation is worst in the blue. The macular pigment is a blue filter. These are independent facts that happen to land in the same place, and together they make the short-wavelength end the region where the phrase “the standard observer” is doing the least work.
What it costs, quantity by quantity
The useful question is not whether the two observers differ but which numbers move and by how much, and the answers vary over four orders of magnitude.
A broadband source’s photometric value: 0.1%. D65’s luminous efficacy of radiation moves by a tenth of a per cent between the two observers. For daylight-like sources the choice is genuinely negligible.
A narrowband source’s photometric value: 104%. A monochromatic source at 470 nm is more than twice as luminous under the 10° functions. This is not academic — 470 nm is close to the pump wavelength of essentially every white LED manufactured, so the photometric rating of the most common light source on earth depends materially on which observer computed it. Specifications say which. Very little else does.
A dominant wavelength: 6 nm. A fixed chromaticity exits the locus at 529.5 nm under the 2° functions and 523.5 nm under the 10°. Six nanometres is larger than the resolution of the instruments that would report such a number, which means the observer choice sits above the noise floor of the measurement it is quoted alongside.
A colour match: nothing at all, and everything. This is the subtle one. Under a single observer, matching is remarkably robust — it survives a sixteen-fold variation in cone composition exactly, because a match is an equality and equalities survive rescaling. Between different observers with different fundamental shapes, matching is not robust at all: two spectra that match for the 2° observer can visibly differ for the 10° one, which is observer metamerism and is a real industrial problem rather than a curiosity.
That last consequence is worth pausing on. Every gamut coverage percentage, every chromaticity plotted for every display ever specified, is computed under the 2° observer, and would be a different number under the 10°. The primaries do not move — they are defined as chromaticities — but the locus they are being compared against does, so the fraction of the visible region they enclose changes. Nobody quotes an observer alongside a coverage figure, and the figure is meaningless without one in exactly the way a hex code is meaningless without a space.
The same figures, drawn again
The cheapest way to see what the choice costs is to take figures this site already had and ask for them under the other kernel. Nothing below is a new computation — each is an existing generator handed a different observer, which is the capability that had sat unused for two phases.
That second figure is the sharpest illustration available of what an observer choice is. MacAdam’s data is a set of measured discrimination thresholds — it is what it is, and no change of standard observer alters what his subject could and could not tell apart. What changes is the diagram those thresholds are drawn on, and therefore every statement of the form “the ellipses are larger in the green” — which is a statement about a mapping, not about an eye.
The last of these is the one with commercial consequences. A pair of samples formulated to match — a car body panel and its plastic bumper, a garment’s shell and its trim — is formulated against a standard observer, and the people who then look at it are not that observer. When the match fails in the showroom the failure is real and the instrument still says the match is perfect.
Even that count is observer-dependent. The proportion of colour directions with no spectral answer is a property of where the locus ends and where the white sits, and both are integrals against the colour-matching functions. It is a small difference and it is a good illustration of how far the choice propagates: a statement about the structure of colour naming turns out to have an observer buried in it.
The cost of choosing between them is largest where the stimulus is narrowest, so the comparison is worth running across the full range of widths a display can span.
Which one to use, and what actually happens
The CIE’s own recommendation has been consistent for sixty years and is widely ignored. The 10° observer is recommended for fields larger than about four degrees, which is most industrial colour work: a paint chip held at reading distance, a fabric swatch, a printed sheet, a car panel. The 2° observer is recommended for small fields, which in practice means self-luminous sources viewed at a distance and not much else.
What happens instead is that almost everything uses the 2° observer, because it came first, because instruments default to it, because it is what the software offers on the first tab, and because a value computed under one observer cannot be compared with a value computed under the other and everyone’s historical data is 2°.
That is the same institutional argument this site keeps running into. The 1931 diagram survived its own replacement by fifty years. The luminous efficiency function is known to be wrong in the blue and is kept anyway, because replacing it would invalidate every photometric measurement ever made. In each case the technically better option has been available for decades and the cost of switching is comparability, which is a real cost and is not the same as the option being wrong.
The measurement, repeated
The point of building an index rather than doing the count once is that the count is now a build-time property and cannot rot.
Every figure on this site is rendered on every build, its strip is read back out, and the result is grouped and published. If a future phase adds fifteen essays that all take the observer as it comes, the spread number falls and the gate refuses the phase. If a generator stops emitting a strip at all — the failure mode this whole family of defects presents as — the index cannot read it and the build stops.
That is a stronger arrangement than a one-off audit, and it is the shape the fleet keeps converging on: a measurement that lives in the build rather than in a report, so that the number and the site cannot disagree. The alternative is what was in place before, which was a rule in a document, kept perfectly, meaning nothing.
What was computed, and how
The index is built by rendering every placement with the same call the registry makes and matching the caption strip’s right-hand slot out of the emitted SVG — the last end-anchored emphasised label, because a figure may carry other end-anchored labels and the strip is always drawn last.
Each string is then classified into one of six bands: the two observers, a space or encoding, a named model at a stated severity, an identity asserted in code, and a stated condition. The classification has to be total, and observercheck fails on a single unrecognised note.
That check earned its keep immediately. The band for named models was written as a list of names — Brettel, Viénot, CIECAM, ΔE2000 — and the first new figure of this phase arrived carrying “Thibos 1992, the chromatic eye”, which belongs in exactly that band and matched none of the names. The gate refused the build. The rule became the shape rather than the roster: a note carrying a four-digit year is a named model. A softer gate would have let the figure through into a band called “other” and the index would have quietly begun accumulating things nobody had classified, which is how the whole defect this essay is about came to exist in the first place.
The spread number — how many generators are called with more than one distinct note — is a ratchet rather than a target. It cannot fall, on the argument the fleet’s parameterisation floor makes at length: a gate turned on already-red is a gate that gets disabled.
Where the model stops
Two standard observers is two, and the number of real observers is the number of people.
Neither the 2° nor the 10° functions describe anybody. Both are averages over small selected panels — seventeen observers for the 1931 set — and the spread between individuals is comparable to the spread between the two standards. Choosing between them is choosing between two averages of a distribution, which is a much smaller decision than the distribution itself, and this site has an essay about that distribution that this one is a companion to rather than a replacement for.
There are also more than two. The Judd–Vos corrections of 1951 and 1978 fix the 1931 functions’ known error in the blue and are used in some applications and not others. The CIE’s 2006 physiologically-based functions are parameterised by field size and age, which is the honest structure, and are not used in industry at all because nothing is calibrated against them. Naming “the observer” as though it were a binary choice is itself a simplification, and one this site makes on every figure.
The index also cannot see figures that name a condition rather than an observer. Nine placements say “at Y = 0.4” and six say “IEC 61966-2-1”, and those are correct strips for figures where an observer is not the governing choice — but the index has no way to check that judgement. It can check that every figure says something and that the something is classifiable. Whether it says the right thing remains a matter of writing the generator carefully.
What a reader should take from this
Three things, in order of how much they generalise.
Ask which observer, whenever a colorimetric number is quoted. Not because the answer is usually surprising — for daylight-like sources it is usually a tenth of a per cent — but because the cases where it matters are the cases where somebody is most likely to have not thought about it. A narrowband emitter, a saturated blue, a large sample: those are where the two averages diverge, and they are all common.
Distinguish the choice between observers from the spread among people. They are different sizes of problem and the second is larger. Choosing the 10° functions over the 2° for a large-field task is straightforwardly correct and buys perhaps a few units of colour difference; the variation between two real people doing the same task is larger, and no choice of standard addresses it.
Notice when an invariant is being kept for free. That is the transferable one. This site’s rule was kept perfectly for two phases, on every figure, verified by nothing — and the way it failed was not that a figure broke the rule but that no figure ever exercised it. Rules of that shape are common and the check for them is cheap: look at the distribution of what the rule produced, not at whether it produced anything.
Who found it, and when
The 1931 functions were adopted at the CIE’s eighth session in Cambridge, from Wright’s and Guild’s experiments on ten and seven observers respectively. The 1964 supplementary standard observer came from Stiles and Burch’s fifty-three-observer ten-degree study and Speranskaya’s work, and was adopted specifically because industrial users had been reporting that visual matches on large samples disagreed with instrumental ones computed under the 2° functions.
That is worth restating: the 10° observer exists because the 2° observer was producing wrong answers in the field, and the wrongness was noticed by people matching paint and fabric rather than by anyone in a laboratory. The correction was available within a generation of the original. It has still not become the default.
The macular-pigment explanation for the difference was worked out over the following decades and is now the standard account; the rod-intrusion component was more contentious and is part of why the 2006 functions are parameterised the way they are.
Where this goes next
Downward, the seventeen observers of 1931 is where the first of these two averages came from, and whose eyes is the distribution both of them are averages of.
Sideways, the finding that produced this essay is not about colour at all. A label that is present, correct and constant is a label carrying nothing, and no check looks for that. The observer index is one site’s answer; the general question — which of a collection’s invariants are being kept in the way that costs nothing — is open, and this is the first case anybody measured.
There is one more thread, and it runs the other way. Every number in the cost table above was computed by asking the same generator for the same figure twice, under two kernels — which is only possible because the generators take the observer as an argument rather than baking it in. That was true from the first phase and had never been exercised. The defect and its repair were both already present in the code; what was missing was any reason to look, and the reason turned out to be a listing page that had to render every figure to build itself.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- Four primaries have a choice colour-matching functions · individual variation · observer metamerism · standard observer · white point
- One person is two observers colour-matching functions · individual variation · macular pigment · observer metamerism · standard observer
- A neutral is everyone's colour individual variation · observer metamerism · standard observer · white point
- Nobody here has two eyes colour-matching functions · individual variation · observer metamerism · standard observer
- One match names the observer colour-matching functions · individual variation · observer metamerism · standard observer
- The lens is worst under tungsten individual variation · macular pigment · observer metamerism · standard observer
What links here
The 8 essays that link to this one and share the most of its objects, of 27 that link here.
The objects this essay names
Each one links to every other essay that touches it.
The CIE 1931 2° observerThe CIE 1964 10° observerColour-matching functionsIndividual variationThe Judd–Vos correctionMacular pigmentObserver metamerismStandard observerWhite pointXYZ