The observers differ by a unit's worth
Assumes Two degrees or ten, A choice with no magnitude and One match names the observer.
Every figure on this site names the observer it was drawn under, because the observer is a choice. Two of them are standard, and how far apart they are turns out to be a number this collection has never printed.
The claim
Over the collection’s own test surfaces, the two standard observers differ by 2.65 ΔE2000 on average — more than any change of light in the adaptation census — and the number has a ×2.23 dependence on which formula is used to take it.
- It is larger than most of what the collection worries about. The mean over 125 surfaces under D65 is 2.65, against a census whose harshest row is 3.37 and whose median is 1.24.
- No file here publishes it, and that is not an oversight so much as a consequence: the essays that argue about the two observers argue about spectra and individual matches, not about an average over a set.
- The menu spreads it from 1.60 to 3.55, with CIELUV lowest and Oklab highest — the two unweighted spaces at opposite ends, which no other row in the inventory does.
- The reason is where the two observers differ, which is the short-wavelength end, and that is precisely where the spaces disagree with each other most.
- And an average is the wrong summary of it. The worst surface is 5.4 ΔE2000 and the best is under 0.2.
What the two observers are
The 1931 functions were measured on seventeen people looking at a two-degree field. The 1964 functions were measured on forty-nine people looking at a ten-degree field, and the CIE issued them as a supplementary observer rather than a replacement, on the grounds that a large field is a different viewing situation rather than a better measurement of the same one.
The physical difference between them is almost entirely pre-neural and has two named causes. The macular pigment covers the central few degrees of the retina and absorbs strongly in the blue, so a two-degree field is seen through it and a ten-degree field is mostly not. And rod intrusion is possible at ten degrees and not at two, though at photopic levels its contribution is small.
The consequence is a systematic difference in the short-wave end of the matching functions, and hence in the tristimulus values of anything with structure down there. Which is nearly everything: a fluorescent lamp’s mercury lines, a brightener in paper, a blue ink, a display’s blue primary.
The number nobody printed
There is no shortage of statements about the two observers in this collection. There is a shortage of one particular kind.
What exists: essays about individual spectra whose tristimulus values disagree, about which laboratory standards require which observer, about what a match under one observer does under the other, and figures drawing the same generator under both — sixty-odd of them, which is a gate this collection enforces.
What does not exist: a single number saying how far apart they are, over a stated set, in a stated unit. Every statement above is about a spectrum or a class of spectra, and each is more informative than an average would be for the question it answers.
The audit needed one anyway, because a menu comparison needs a scalar. So it computes one — 125 test surfaces under D65, each given tristimulus values by both observers, each judged against its own observer’s white — and the row is marked as having no source rather than given one that does not exist. Inventing a source would have been the worse error, and the marking is the honest form of a gap.
The number itself is worth having, and it is larger than the essays that avoided printing it would suggest.
How large
| unit | the two observers | the census’s median row | ratio |
|---|---|---|---|
| ΔE*uv | 1.596 | 1.052 | 1.52 |
| ΔE*ab | 1.792 | 0.995 | 1.80 |
| ΔE2000 | 2.650 | 1.237 | 2.14 |
| ΔE*94 | 2.797 | 1.311 | 2.13 |
| CAM16-UCS | 3.136 | 1.751 | 1.79 |
| Oklab | 3.551 | 1.115 | 3.18 |
In the published unit, choosing the wrong standard observer costs about twice what the median change of light in the adaptation census costs — and about eight times what daylight going from D65 to D50 costs, which is a change nobody would describe as negligible.
That ordering is worth sitting with. This collection spends a field on chromatic adaptation, on the grounds that a change of light is a real and quantifiable problem for a colorimetric system. A change of observer between two standards, both in current use, both cited in standards documents, is a larger problem by a factor of two, and it has no adaptation transform attached to it because there is nothing to adapt: the two observers are two different instruments, not one instrument under two conditions.
The ratio does not cancel the unit, and that is a finding about the rule
The third column of that table is there to be reassuring. This collection’s standing advice, arrived at from three directions, is to publish the ratio rather than the level, because a change of unit is a common-mode factor and a ratio divides it out. The census’s own audit says so, the quadrature audit says so, and the worst-case audit says so.
It does not work here, and the failure is instructive.
The observer column spans a factor of 2.22 across the menu. The census column spans 1.76. And the ratio between them — the column that was supposed to be free of both — spans 2.10. Dividing one by the other has bought almost nothing: the ratio is no more stable than either quantity it was formed from, and it is less stable than the census column alone.
The reason is visible in the ordering. Across the six units the two columns correlate at a rank correlation of +0.60 — the same two formulae are not at the top of both, and Oklab, which reads the observer gap highest of anything on the menu, reads the census third lowest. Two quantities that a change of unit moves in different orders cannot have a stable ratio, and a ratio only cancels a unit when the unit acts on both numerator and denominator the same way.
Which is the condition the rule quietly assumes and which this row does not meet. A change of unit is a common-mode factor when the two quantities being compared are the same kind of displacement over the same objects — two census rows against each other, two spaces scored on one set of ellipses, a worst case against a mean of the same 125 surfaces. All three of the cases the rule was derived from are of that kind. Here the numerator is a displacement concentrated in the blue and the denominator is a displacement spread over the whole hue circle, and a formula that weights the blues differently from everything else moves the two by different factors.
So the rule needs a condition attached to it, and the condition is testable before the ratio is quoted: compute both quantities under two units and see whether their ordering survives. Where it does, the ratio is a comparative number and carries almost none of the unit’s uncertainty; where it does not, the ratio is a third quantity with an uncertainty of its own that may exceed either parent’s.
That is a small correction to a rule this collection uses everywhere and it is worth making precisely because the rule is otherwise good. A comparative number is not automatically robust — it is robust when the comparison is between like things, and the whole point of this row is that the observer gap is not like a change of light. It has no adaptation transform, it is a change of instrument rather than of condition, and it lives in one part of the spectrum. Every one of those is a reason the ratio to a census row would not be expected to behave, and every one of them was already stated in this essay before the ratio was printed.
The practical form for a reader is the one sentence the table cannot carry. The observer gap is about twice the census’s median row in the published unit, and between one and a half and three times it depending on which unit is asked — and the honest way to quote it is with both the unit and the spread, in exactly the way the level is quoted, because the ratio has bought no protection from either.
Why this row spreads the way it does
The ×2.23 spread across the menu is the second largest in the inventory, and its shape is unlike every other row.
Every other quantity in the inventory has CAM16-UCS at or near the top and one of the unweighted formulae at the bottom. This one has Oklab at the top and CIELUV at the bottom, with the appearance unit in between and the two weighted CIELAB formulae close together in the middle. The pattern that governs the rest of the audit — whether the formula divides by chroma — does not govern this row.
What governs it is where in the spectrum the difference lives. The two observers differ in the blue, so the surfaces they disagree about are moved along the blue–yellow direction at moderate lightness, and the three unweighted spaces treat that direction quite differently: CIELUV’s chromaticity coordinates are a projective transform of the diagram, CIELAB’s are cube-root differences of tristimulus values, and Oklab’s are a fitted cube-root of a cone-like intermediate. A difference concentrated in one direction is measured by whichever axis that direction happens to fall on, and the three spaces put it in three different places.
The same mechanism produced the other outlier in the census table: an older lens costs 1.02 in CIELAB and 0.63 in CIELUV, and a yellowing lens absorbs in the blue for the same reason the macular pigment does. The two rows that are about a filter in front of the retina are the two rows the choice of space decides.
An average is the wrong summary, and here is what it hides
Over the 125 surfaces the gap runs from under 0.2 to 5.44 ΔE2000, and the distribution is not symmetric: most surfaces are moved a little and a minority are moved a great deal.
The minority is identifiable. A surface with most of its reflectance in the long wavelengths is barely touched, because the two observers agree closely above 550 nm. A surface with structure in the blue is moved hard. And a surface whose short-wave content is narrow — a saturated blue, an inked patch, a phosphor — is moved hardest of all, because the two observers’ short-wave functions differ in shape as well as in amplitude.
This is the pattern that makes a mean a poor description of a set, and it repeats the collection’s own finding from the adaptation side: the ratio of the worst case to the mean is about two on a broad change and about four on a sharp spectral one. Here the ratio is 2.05, which places the observer difference among the broad changes despite being caused by a filter — because a macular pigment’s absorption band is 100 nm wide, which is broad on the scale a cone can resolve.
The two observers are not two conditions
There is a reading of the number that makes it look smaller than it is, and it should be blocked.
The reading goes: the 1964 observer describes a ten-degree field and the 1931 observer a two-degree one, so the difference between them is a viewing condition — like a surround, or an adapting luminance — and a colour appearance model would absorb it. Then 2.65 ΔE2000 is not an uncertainty at all but a prediction, and the right response is to use the observer that matches the field.
That is correct as far as it goes and it does not go far. The field size in a real viewing situation is not two degrees or ten; it is whatever the object subtends, and the CIE provides functions at exactly two values with no interpolation sanctioned between them. A colour patch on a page at reading distance is about three degrees. A wall is fifty. Neither has a standard observer, and the practice everywhere is to pick the nearer of the two.
So the number is best read as the size of the quantisation error in a system that offers two settings for a continuous variable. A quantity varying continuously between two tabulated points, sampled at whichever point is nearer, has an error of up to about half the gap in the middle and the whole gap at the ends. The same shape appears in the three tabulated surrounds an appearance model offers, and it has the same repair: a continuum, with the tabulated values as points on it.
Whether a continuum is available here is a real question and the answer is partly yes. The physical cause is a filter of stated density in front of the receptors, and a filter has a density that can be varied continuously — which is exactly what this collection’s own population does, where macular density is one of five measured variates and its range covers the two-to-ten-degree difference and more. The continuum exists in the machinery here and is not in the standard, which is the same shape of gap as the surround.
What a reader should do with the number
Three things follow, and the third is the one this collection has to apply to itself.
Name the observer, which was already the rule. Every figure here does, and this number is the size of what the rule protects against. A chromaticity diagram with no observer named is a diagram whose points are uncertain by about two and a half units in a unit the reader also has not been told.
Do not adapt between them. A chromatic adaptation transform maps between white points, and the two observers do not have a white-point difference in that sense — they have different matching functions, so the same spectrum has different tristimulus values under each and no 3×3 fixes it for every spectrum at once. A matrix fitted to do it is fitted on a set and fails off it, which is the same structure as every other fitted transform here.
And treat the choice as an input with an elasticity. This collection’s own median observer is built from a pigment template rather than from either standard, and sits about 0.95 ΔE2000 from the 1931 functions. That is a third of the gap between the two standards, which is a useful calibration: the collection’s own construction is closer to the 1931 observer than the 1964 observer is.
The unit the difference is quoted in is itself a choice, and one ellipse is enough to show how much that choice is worth.
Where the model stops
The 125 surfaces are constructed, smooth and single-lobed, and real reflectances have structure the construction does not. A set with narrow-band members — inks, phosphors, interference colours — would give a larger gap, and by the argument above, considerably larger. This number is a lower bound on what the observer choice costs in practice, and it is already twice the census’s median.
It is also computed under one illuminant. Under a lamp with emission lines the two observers diverge much further, because their disagreement is a difference of two functions and a line spectrum samples that difference at a point rather than integrating over it. The essay on what a discharge lamp does to an adapted observer has the same structure and the same cause.
And it is a comparison of two standards, not of two people. The spread across a population of real observers is a different quantity, is larger, and is not made smaller by choosing the right standard.
Who found it, and when
The two observers and their difference are the CIE’s, 1931 and 1964, and the macular explanation is Stiles and Burch’s, from the measurements the 1964 functions were built on. That the difference matters for narrow-band sources is the standard warning in every colorimetry text and is the reason a display industry measuring blue primaries argues about it.
What is not standard is putting one number on it. The reason is good: the number depends entirely on the set of spectra, and a text that printed one would be inviting exactly the misreading this collection spends a round warning against. The defensible version is a number with its set named beside it, which is the convention the audit of averages arrived at and which this row now follows: 2.65 ΔE2000 over 125 constructed reflectances under D65, worst case 5.44, and a factor of 2.2 of unit-choice on top.
Where the ladder goes next
Of the six formulae, the one this collection publishes in is a patch on the space its own uniformity instrument ranks last of the three it can rank. Whether the patch works is not a matter of opinion, because the ellipses are a measurement of people — and turning that instrument on the units rather than on the spaces gives an answer with two components that move in opposite directions.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- Two yellow filters cancel on a slope colour difference · macular pigment · standard observer · test set · uncertainty
- One person is two observers colour-matching functions · macular pigment · metamerism · standard observer
- The coincidence was a mechanism calibration · colour difference · metamerism · test set
- The disagreement is at the near end calibration · colour difference · test set · uncertainty
- The objective nobody chose calibration · colour difference · test set · tristimulus
- Three choices reached calibration · colour difference · test set · uncertainty
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CalibrationColour differenceColour-matching functionsField sizeMacular pigmentMetamerismStandard observerTest setTristimulusUncertainty