Matching and measuring

The observers differ by a unit's worth

The gap between the 1931 and 1964 standard observers is the one quantity in this collection's audit with no published number under it — nothing here reports it as a single figure over a stated set. It also has the second-largest dependence on which colour-difference formula is used, running from 1.60 to 3.55 across the menu.

Assumes Two degrees or ten, A choice with no magnitude and One match names the observer.

Every figure on this site names the observer it was drawn under, because the observer is a choice. Two of them are standard, and how far apart they are turns out to be a number this collection has never printed.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 1 Six published quantities under six colour-difference formulae. The third row is the gap between the two standard observers, and it is the only row with no number in the “published” column — nothing in this collection reports it as a single figure over a stated set.

The claim

Over the collection’s own test surfaces, the two standard observers differ by 2.65 ΔE2000 on average — more than any change of light in the adaptation census — and the number has a ×2.23 dependence on which formula is used to take it.

  • It is larger than most of what the collection worries about. The mean over 125 surfaces under D65 is 2.65, against a census whose harshest row is 3.37 and whose median is 1.24.
  • No file here publishes it, and that is not an oversight so much as a consequence: the essays that argue about the two observers argue about spectra and individual matches, not about an average over a set.
  • The menu spreads it from 1.60 to 3.55, with CIELUV lowest and Oklab highest — the two unweighted spaces at opposite ends, which no other row in the inventory does.
  • The reason is where the two observers differ, which is the short-wavelength end, and that is precisely where the spaces disagree with each other most.
  • And an average is the wrong summary of it. The worst surface is 5.4 ΔE2000 and the best is under 0.2.

What the two observers are

The 1931 functions were measured on seventeen people looking at a two-degree field. The 1964 functions were measured on forty-nine people looking at a ten-degree field, and the CIE issued them as a supplementary observer rather than a replacement, on the grounds that a large field is a different viewing situation rather than a better measurement of the same one.

The physical difference between them is almost entirely pre-neural and has two named causes. The macular pigment covers the central few degrees of the retina and absorbs strongly in the blue, so a two-degree field is seen through it and a ten-degree field is mostly not. And rod intrusion is possible at ten degrees and not at two, though at photopic levels its contribution is small.

The consequence is a systematic difference in the short-wave end of the matching functions, and hence in the tristimulus values of anything with structure down there. Which is nearly everything: a fluorescent lamp’s mercury lines, a brightener in paper, a blue ink, a display’s blue primary.

The number nobody printed

There is no shortage of statements about the two observers in this collection. There is a shortage of one particular kind.

What exists: essays about individual spectra whose tristimulus values disagree, about which laboratory standards require which observer, about what a match under one observer does under the other, and figures drawing the same generator under both — sixty-odd of them, which is a gate this collection enforces.

What does not exist: a single number saying how far apart they are, over a stated set, in a stated unit. Every statement above is about a spectrum or a class of spectra, and each is more informative than an average would be for the question it answers.

The audit needed one anyway, because a menu comparison needs a scalar. So it computes one — 125 test surfaces under D65, each given tristimulus values by both observers, each judged against its own observer’s white — and the row is marked as having no source rather than given one that does not exist. Inventing a source would have been the worse error, and the marking is the honest form of a gap.

The number itself is worth having, and it is larger than the essays that avoided printing it would suggest.

How large

unit the two observers the census’s median row ratio
ΔE*uv 1.596 1.052 1.52
ΔE*ab 1.792 0.995 1.80
ΔE2000 2.650 1.237 2.14
ΔE*94 2.797 1.311 2.13
CAM16-UCS 3.136 1.751 1.79
Oklab 3.551 1.115 3.18

In the published unit, choosing the wrong standard observer costs about twice what the median change of light in the adaptation census costs — and about eight times what daylight going from D65 to D50 costs, which is a change nobody would describe as negligible.

That ordering is worth sitting with. This collection spends a field on chromatic adaptation, on the grounds that a change of light is a real and quantifiable problem for a colorimetric system. A change of observer between two standards, both in current use, both cited in standards documents, is a larger problem by a factor of two, and it has no adaptation transform attached to it because there is nothing to adapt: the two observers are two different instruments, not one instrument under two conditions.

The ratio does not cancel the unit, and that is a finding about the rule

The third column of that table is there to be reassuring. This collection’s standing advice, arrived at from three directions, is to publish the ratio rather than the level, because a change of unit is a common-mode factor and a ratio divides it out. The census’s own audit says so, the quadrature audit says so, and the worst-case audit says so.

It does not work here, and the failure is instructive.

The observer column spans a factor of 2.22 across the menu. The census column spans 1.76. And the ratio between them — the column that was supposed to be free of both — spans 2.10. Dividing one by the other has bought almost nothing: the ratio is no more stable than either quantity it was formed from, and it is less stable than the census column alone.

The reason is visible in the ordering. Across the six units the two columns correlate at a rank correlation of +0.60 — the same two formulae are not at the top of both, and Oklab, which reads the observer gap highest of anything on the menu, reads the census third lowest. Two quantities that a change of unit moves in different orders cannot have a stable ratio, and a ratio only cancels a unit when the unit acts on both numerator and denominator the same way.

Which is the condition the rule quietly assumes and which this row does not meet. A change of unit is a common-mode factor when the two quantities being compared are the same kind of displacement over the same objects — two census rows against each other, two spaces scored on one set of ellipses, a worst case against a mean of the same 125 surfaces. All three of the cases the rule was derived from are of that kind. Here the numerator is a displacement concentrated in the blue and the denominator is a displacement spread over the whole hue circle, and a formula that weights the blues differently from everything else moves the two by different factors.

So the rule needs a condition attached to it, and the condition is testable before the ratio is quoted: compute both quantities under two units and see whether their ordering survives. Where it does, the ratio is a comparative number and carries almost none of the unit’s uncertainty; where it does not, the ratio is a third quantity with an uncertainty of its own that may exceed either parent’s.

That is a small correction to a rule this collection uses everywhere and it is worth making precisely because the rule is otherwise good. A comparative number is not automatically robust — it is robust when the comparison is between like things, and the whole point of this row is that the observer gap is not like a change of light. It has no adaptation transform, it is a change of instrument rather than of condition, and it lives in one part of the spectrum. Every one of those is a reason the ratio to a census row would not be expected to behave, and every one of them was already stated in this essay before the ratio was printed.

The practical form for a reader is the one sentence the table cannot carry. The observer gap is about twice the census’s median row in the published unit, and between one and a half and three times it depending on which unit is asked — and the honest way to quote it is with both the unit and the spread, in exactly the way the level is quoted, because the ratio has bought no protection from either.

The CIE 1964 colour-matching functionsThe three functions that turn a spectrum into three numbers. They are all positive, which is why XYZ exists — the RGB functions they were derived from are not. ȳ is by construction the luminous efficiency function, which is why luminance comes out of Y.400450500550600650700wavelength / nmȳȳ is the luminous efficiency functionCIE 1964 10° observer
Fig. 2 The two sets of colour-matching functions on one axis. The separation is concentrated below about 500 nm, which is where the macular pigment absorbs and where a two-degree field is looking through it.

Why this row spreads the way it does

The ×2.23 spread across the menu is the second largest in the inventory, and its shape is unlike every other row.

Every other quantity in the inventory has CAM16-UCS at or near the top and one of the unweighted formulae at the bottom. This one has Oklab at the top and CIELUV at the bottom, with the appearance unit in between and the two weighted CIELAB formulae close together in the middle. The pattern that governs the rest of the audit — whether the formula divides by chroma — does not govern this row.

What governs it is where in the spectrum the difference lives. The two observers differ in the blue, so the surfaces they disagree about are moved along the blue–yellow direction at moderate lightness, and the three unweighted spaces treat that direction quite differently: CIELUV’s chromaticity coordinates are a projective transform of the diagram, CIELAB’s are cube-root differences of tristimulus values, and Oklab’s are a fitted cube-root of a cone-like intermediate. A difference concentrated in one direction is measured by whichever axis that direction happens to fall on, and the three spaces put it in three different places.

The same mechanism produced the other outlier in the census table: an older lens costs 1.02 in CIELAB and 0.63 in CIELUV, and a yellowing lens absorbs in the blue for the same reason the macular pigment does. The two rows that are about a filter in front of the retina are the two rows the choice of space decides.

ΔEok against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔEok, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 208 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 34.7 per cent of the mean and the rank correlation is 0.881. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 3 Oklab against the published unit on the reference pairs. Its scatter is the widest on the menu, and its disagreement is concentrated in the pairs that differ in the blues — which is exactly where the two standard observers differ from each other.
ΔE*uv against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE*uv, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 153 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 32.3 per cent of the mean and the rank correlation is 0.923. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 4 CIELUV on the same axes, at the other end of this row. The two spaces the 1976 committee could not choose between disagree about the observer question by a factor of 2.2.

An average is the wrong summary, and here is what it hides

Over the 125 surfaces the gap runs from under 0.2 to 5.44 ΔE2000, and the distribution is not symmetric: most surfaces are moved a little and a minority are moved a great deal.

The minority is identifiable. A surface with most of its reflectance in the long wavelengths is barely touched, because the two observers agree closely above 550 nm. A surface with structure in the blue is moved hard. And a surface whose short-wave content is narrow — a saturated blue, an inked patch, a phosphor — is moved hardest of all, because the two observers’ short-wave functions differ in shape as well as in amplitude.

This is the pattern that makes a mean a poor description of a set, and it repeats the collection’s own finding from the adaptation side: the ratio of the worst case to the mean is about two on a broad change and about four on a sharp spectral one. Here the ratio is 2.05, which places the observer difference among the broad changes despite being caused by a filter — because a macular pigment’s absorption band is 100 nm wide, which is broad on the scale a cone can resolve.

The two observers are not two conditions

There is a reading of the number that makes it look smaller than it is, and it should be blocked.

The reading goes: the 1964 observer describes a ten-degree field and the 1931 observer a two-degree one, so the difference between them is a viewing condition — like a surround, or an adapting luminance — and a colour appearance model would absorb it. Then 2.65 ΔE2000 is not an uncertainty at all but a prediction, and the right response is to use the observer that matches the field.

That is correct as far as it goes and it does not go far. The field size in a real viewing situation is not two degrees or ten; it is whatever the object subtends, and the CIE provides functions at exactly two values with no interpolation sanctioned between them. A colour patch on a page at reading distance is about three degrees. A wall is fifty. Neither has a standard observer, and the practice everywhere is to pick the nearer of the two.

So the number is best read as the size of the quantisation error in a system that offers two settings for a continuous variable. A quantity varying continuously between two tabulated points, sampled at whichever point is nearer, has an error of up to about half the gap in the middle and the whole gap at the ends. The same shape appears in the three tabulated surrounds an appearance model offers, and it has the same repair: a continuum, with the tabulated values as points on it.

Whether a continuum is available here is a real question and the answer is partly yes. The physical cause is a filter of stated density in front of the receptors, and a filter has a density that can be varied continuously — which is exactly what this collection’s own population does, where macular density is one of five measured variates and its range covers the two-to-ten-degree difference and more. The continuum exists in the machinery here and is not in the standard, which is the same shape of gap as the surround.

What a reader should do with the number

Three things follow, and the third is the one this collection has to apply to itself.

Name the observer, which was already the rule. Every figure here does, and this number is the size of what the rule protects against. A chromaticity diagram with no observer named is a diagram whose points are uncertain by about two and a half units in a unit the reader also has not been told.

Do not adapt between them. A chromatic adaptation transform maps between white points, and the two observers do not have a white-point difference in that sense — they have different matching functions, so the same spectrum has different tristimulus values under each and no 3×3 fixes it for every spectrum at once. A matrix fitted to do it is fitted on a set and fails off it, which is the same structure as every other fitted transform here.

And treat the choice as an input with an elasticity. This collection’s own median observer is built from a pigment template rather than from either standard, and sits about 0.95 ΔE2000 from the 1931 functions. That is a third of the gap between the two standards, which is a useful calibration: the collection’s own construction is closer to the 1931 observer than the 1964 observer is.

Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu.
Fig. 5 Where on the scale the units disagree. The observer gap sits at about 2.6 units, in the band where the relative disagreement between formulae is still a quarter to a half of the value — which is where its ×2.23 spread comes from.

The unit the difference is quoted in is itself a choice, and one ellipse is enough to show how much that choice is worth.

One of MacAdam's ellipses as each unit sees it, at x 0.365, y 0.153. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 1.37; the least round is ΔE*ab at 3.73. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 6 A single measured ellipse drawn in each of the six units, each scaled to its own mean radius. A difference of one unit means a different shape in each of them, which is the arithmetic behind quoting an observer difference in any of them.

Where the model stops

The 125 surfaces are constructed, smooth and single-lobed, and real reflectances have structure the construction does not. A set with narrow-band members — inks, phosphors, interference colours — would give a larger gap, and by the argument above, considerably larger. This number is a lower bound on what the observer choice costs in practice, and it is already twice the census’s median.

It is also computed under one illuminant. Under a lamp with emission lines the two observers diverge much further, because their disagreement is a difference of two functions and a line spectrum samples that difference at a point rather than integrating over it. The essay on what a discharge lamp does to an adapted observer has the same structure and the same cause.

And it is a comparison of two standards, not of two people. The spread across a population of real observers is a different quantity, is larger, and is not made smaller by choosing the right standard.

How far a match comes apart when the observer changes. A broad source and a three-primary source, solved at each primary width so the pair is an exact tristimulus match for the CIE 1931 observer. The pair is then handed to the 1964 observer, and the gap between them is plotted. For the observer they were built for the gap is arithmetic noise at every width. For the other it grows as the primaries narrow, reaching 0.012 at 10 nm — and displays have been getting narrower for twenty years.
Fig. 7 The two standard observers disagreeing about specific spectra, which is how this collection has always argued the point. It is a more informative picture than an average and it cannot be put in a table beside a census row, which is why the average was computed as well.
MacAdam's twenty-five ellipses, measured in each unit. The uniformity instrument used here, applied to units rather than to spaces. The upper bar is anisotropy — the mean over the twenty-five of the largest radius divided by the smallest, where 1 would be a circle. The lower is spread — the largest mean radius divided by the smallest across all twenty-five, which asks whether a step of the same size means the same thing in different parts of the diagram. Reading down the three CIELAB-based formulae in the order they were published, the anisotropy falls 3.42 → 2.89 → 2.74 and the spread rises 3.24 → 3.59 → 4.12: the weighting divides a difference by the chroma it was measured at, which equalises directions at a point and unequalises magnitudes between points. Neither number is scaled, so no calibration is applied here. CAM16-UCS is ahead on both.
Fig. 8 The six units on the ellipse instrument. The observer gap’s readings are ordered quite differently from this, which is what says the row is measuring a direction in the space rather than a general property of the formulae.

Who found it, and when

The two observers and their difference are the CIE’s, 1931 and 1964, and the macular explanation is Stiles and Burch’s, from the measurements the 1964 functions were built on. That the difference matters for narrow-band sources is the standard warning in every colorimetry text and is the reason a display industry measuring blue primaries argues about it.

What is not standard is putting one number on it. The reason is good: the number depends entirely on the set of spectra, and a text that printed one would be inviting exactly the misreading this collection spends a round warning against. The defensible version is a number with its set named beside it, which is the convention the audit of averages arrived at and which this row now follows: 2.65 ΔE2000 over 125 constructed reflectances under D65, worst case 5.44, and a factor of 2.2 of unit-choice on top.

Where the ladder goes next

Of the six formulae, the one this collection publishes in is a patch on the space its own uniformity instrument ranks last of the three it can rank. Whether the patch works is not a matter of opinion, because the ellipses are a measurement of people — and turning that instrument on the units rather than on the spaces gives an answer with two components that move in opposite directions.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationColour differenceColour-matching functionsField sizeMacular pigmentMetamerismStandard observerTest setTristimulusUncertainty