The census in six units
Assumes A choice with no magnitude, What no adaptation can remove and A mean has a set under it.
The adaptation census is fourteen changes of light with a number beside each, and the number is a mean over a constructed set of surfaces of what an adapted observer is left with. Every one of those numbers is in one unit.
The claim
With the change of scale taken out, the census’s levels move by between a factor of 1.5 and a factor of 3.7 depending on the row — and the largest movement is on the rows with the smallest answers.
- The two ends of the ranking are immovable. A blackbody at daylight’s own temperature is the mildest change in the census under all six units, and a green wall bounced twice is the harshest under all six.
- The middle is not. Between two and ten of the ninety-one pairs of rows change places, depending on which unit is asked.
- The spread is largest where the answer is smallest. The blackbody row spans ×3.71 across the menu and the green wall bounced twice spans ×1.51; the rank correlation between a row’s residual and its spread is −0.947. That trend is real and it is almost entirely one unit’s doing, which is the section below that had to be written after the table rather than before it.
- The appearance unit is the outlier upwards, uniformly. CAM16-UCS reads every one of the fourteen rows as the largest or nearly the largest, which is not a disagreement about ordering at all.
- And one row is a different shape from the others. An older lens costs 1.02 in ΔE*ab and 0.63 in ΔE*uv, a ratio of 1.6 between two formulae published in the same year by the same committee.
What the census is, before it is a table
The census exists because a change of light is not a single thing, and it is worth restating what the fourteen rows are before asking what a unit does to them.
Each row is a change of context: a light arriving at a surface, and a second light arriving at the same surface. Daylight to a warmer daylight, daylight to tungsten, daylight to a discharge lamp with emission lines in it, daylight to daylight after it has bounced off a painted wall, and — two rows that are inside the observer rather than in the room — daylight seen through the macular pigment and daylight seen through an older lens.
For each, an adapted observer applies the gain their own machinery gives them: the ratio of the two whites, in a fixed basis. What is left over is the row’s number. It is not what a better transform would fix, because the matrix that would fix it is a different matrix for every row and the observer has one mechanism.
The set the mean is taken over is 125 constructed surfaces, and that set has been audited from four directions. The unit is the thing under all of it that had not been.
Taking the scale out first
The six formulae are not on one scale, so the first column of any comparison is bookkeeping.
Uncalibrated, the census means run 0.53, 1.16, 1.31, 0.42, 0.62 and 1.58 — which reads as chaos and is mostly one fact: a Euclidean distance in CIELAB is about 1.8 times CIEDE2000 on any pair anybody measures, because CIEDE2000’s weighting functions divide by chroma and lightness. So each formula is first multiplied by the scale that best carries it onto ΔE2000 over a common sample of 374 pairs, built at the separations these arguments actually live at.
After that the means are 0.958, 1.129, 1.312, 0.962, 0.980 and 1.664. Still not equal — a calibration on one sample does not make two formulae agree on a different one, and if it did there would be nothing to measure — but now every remaining difference is a real disagreement about these surfaces under these lights.
The table
| change of light | ΔE*ab | ΔE*94 | ΔE2000 | ΔE*uv | Oklab | CAM16-UCS | spread |
|---|---|---|---|---|---|---|---|
| a blackbody at daylight’s temperature | 0.227 | 0.235 | 0.263 | 0.213 | 0.170 | 0.631 | ×3.71 |
| the macular pigment | 0.256 | 0.379 | 0.368 | 0.222 | 0.450 | 0.793 | ×3.57 |
| daylight to D50 | 0.307 | 0.402 | 0.459 | 0.327 | 0.396 | 0.910 | ×2.96 |
| daylight to D100 | 0.321 | 0.440 | 0.508 | 0.339 | 0.470 | 0.969 | ×3.02 |
| daylight to a three-primary display | 0.667 | 0.835 | 0.980 | 0.637 | 0.650 | 1.368 | ×2.15 |
| daylight to D40 | 0.672 | 0.858 | 0.980 | 0.725 | 0.816 | 1.471 | ×2.19 |
| an older lens | 1.023 | 1.046 | 1.117 | 0.630 | 0.800 | 1.586 | ×2.52 |
| a red wall | 0.995 | 1.308 | 1.357 | 1.091 | 1.354 | 1.884 | ×1.89 |
| daylight to halogen | 0.990 | 1.277 | 1.583 | 1.013 | 1.158 | 1.911 | ×1.93 |
| daylight to tungsten | 1.136 | 1.495 | 1.635 | 1.248 | 1.482 | 2.072 | ×1.82 |
| daylight to a white LED | 0.967 | 1.342 | 1.697 | 0.984 | 1.123 | 1.931 | ×2.00 |
| a green wall, one bounce | 1.333 | 1.414 | 1.722 | 1.444 | 1.165 | 2.102 | ×1.80 |
| daylight to a triphosphor tube | 1.811 | 2.020 | 2.323 | 1.651 | 1.456 | 2.455 | ×1.69 |
| a green wall, two bounces | 2.703 | 2.749 | 3.375 | 2.946 | 2.236 | 3.213 | ×1.51 |
Read down the last column and the pattern is unmistakable: the spread falls as the residual rises, almost monotonically, from ×3.71 on the mildest row to ×1.51 on the harshest.
Why the mild rows move most
That ordering is not an artefact and it is not obvious in advance, so it is worth taking apart.
A mild residual is a mean over surfaces that are nearly matched after adaptation — small differences, close to the origin of the difference formula. A harsh residual is a mean over surfaces the observer gets substantially wrong. And the six formulae were fitted to different data at different separations, so they agree best where they were all trying hardest and worst where none of them was.
Two mechanisms produce it. The weighting functions in ΔE*94 and ΔE2000 divide by chroma, which matters more when the two colours are far apart in chroma than when they are barely apart in anything. And CAM16-UCS’s compression of colourfulness is a logarithm, which is steepest near zero — so a change of light that leaves almost nothing behind still produces a report the appearance model treats as a real, if small, difference.
The consequence is practical and slightly awkward. The rows this collection is least confident about for other reasons are also the rows whose unit-dependence is largest. The blackbody row is 0.263 and is already known to be inside the standard error of its neighbours; it now also has a factor of 3.7 of unit-choice on it. A number with two independent sources of a factor of two or three under it is not a number to print to three figures.
The trend is one unit’s
The paragraph above explains the trend, and the explanation is sound as far as it goes. It is also not what is producing most of the number, which is worth catching before anybody quotes the ×3.71.
The spread is a range statistic over six samples, so it is decided by two of them and ignores the other four. And CAM16-UCS is the largest entry on thirteen of the fourteen rows — the exception being the green wall bounced twice, where ΔE2000 passes it. So the spread column is very close to a single ratio: the appearance unit over whichever matching formula happens to be lowest.
Recomputing the same column over the five matching formulae alone separates the two effects:
| over all six | over the five matching units | |
|---|---|---|
| blackbody at daylight’s temperature | ×3.71 | ×1.55 |
| the macular pigment | ×3.57 | ×2.03 |
| an older lens | ×2.52 | ×1.77 |
| daylight to a white LED | ×2.00 | ×1.75 |
| a green wall, two bounces | ×1.51 | ×1.51 |
| rank correlation with the residual | −0.947 | −0.156 |
The trend does not survive. Across the five matching formulae the spread runs from ×1.36 to ×2.03, with a mean of 1.58 and no relationship to the level at all — a rank correlation of −0.16 on fourteen points is nothing. The falling column in the table above is the appearance unit’s common-mode shift being divided by a rising denominator, which is arithmetic rather than a property of the formulae.
This site has already said why that shift exists: CAM16-UCS is measuring a slightly different quantity, and a section further down this page says so again. What had not been done is to notice that the same fact governs the spread column, and that the claim in the third bullet is therefore about one unit rather than about the menu.
And the residue is more interesting than the trend was. With the appearance unit set aside, the three rows the matching formulae disagree about most are the macular pigment at ×2.03, an older lens at ×1.77, and daylight to a white LED at ×1.75 — which is both of the two rows that are inside the observer, plus the one lamp on the list with a narrow blue emitter in it. That is not a grouping by how mild a row is. It is a grouping by where in the spectrum the change acts, and the essay already has the mechanism for it, three sections down: the formulae are built differently in the blues, and every one of those three rows moves surfaces along the blue–yellow axis.
What does not move
Three things survive every unit on the menu, and they are the three the census is actually used for.
A blackbody at daylight’s own temperature is the mildest change in the census, in all six units, by a margin of at least 1.3 to the next row. This is the census’s most-cited result and it is what a reader would hope: an observer whose visual system evolved under a thermal source discounts a thermal source almost perfectly, and no colour-difference formula disagrees.
A green wall bounced twice is the harshest, in all six, by a margin of at least 1.3. Also unsurprising and also robust: a filter applied twice is a filter squared, and nothing in a diagonal gain can undo the squaring.
And the surface changes are worse than the daylight changes, in all six. A room is harder than a lamp, which is the census’s structural finding and is not a matter of unit at all.
What moves is the middle: the discharge lamps against the surfaces, and the two ocular filters against everything. Which pairs, and how that compares against the pairs the test set itself could not resolve, is a question two independent instruments turn out to answer partly the same way.
The lens row, which is a different shape
One row disagrees in a way the others do not, and it is the one that is not about a lamp.
An older lens costs 1.023 in ΔE*ab and 0.630 in ΔE*uv — a ratio of 1.62 between two formulae published in the same year, by the same committee, as alternatives for the same job. No other row in the census has a CIELAB-to-CIELUV ratio above 1.15.
The reason is where a yellowing lens acts. It absorbs in the blue, so the surfaces it moves are moved mostly along the blue–yellow axis at low luminance, and that is exactly the region where CIELAB and CIELUV are built differently: CIELUV’s chromaticity coordinates are a projective transform of the diagram and CIELAB’s are a cube-root difference of tristimulus values, and the two treat the dark blues quite differently.
So the row that reports what happens to an observer over a lifetime is the row whose answer depends most on which of two 1976 recommendations is used. That is worth knowing about every published number describing age-related change in colour vision, most of which are quoted in one of the two without saying which.
The appearance unit is not disagreeing about order
CAM16-UCS reads every row higher than any other unit reads it, which looks like a strong disagreement and is nearly the opposite.
Its ranking is the closest of the five to the published one: Kendall’s τ of 0.956, with two reversed pairs out of ninety-one, against ΔE*ab’s eight and Oklab’s ten. What it does is shift every row upwards by a similar amount, and a common-mode shift reorders nothing.
The shift itself has a cause worth naming. CAM16-UCS is not a distance between two tristimulus values; it is a distance between two reports from a model with a surround, an adapting luminance and a degree of adaptation in it. The model’s own adaptation step is incomplete by construction — it never fully discounts a white — so a change of light that a von Kries gain has been credited with fully removing still leaves a residue in the appearance model’s report. The appearance unit is not measuring the same quantity slightly differently; it is measuring a slightly different quantity, and this collection has a separate essay about what happens when it is used as a ruler for a matching question.
Where the model stops
The census is a census of constructed changes of light acting on constructed surfaces, and neither the change nor the surface is measured. That constraint is older than this question and is stated wherever the census appears; what changes here is that a third constructed thing has been added to the list, and it is the ruler.
Nothing in this exercise establishes that any of these units is closer to what a person would report. Doing that needs a psychophysical experiment, this collection has none, and the literature’s answer depends on which threshold dataset the question is asked with. What it establishes is the size of the exposure, which is a factor of about two on a typical row and up to 3.7 on the mildest.
And it does not reach the objective the surfaces were selected under, or the basis the gain is taken in, both of which have their own audits. Stacking those with this one is not a matter of multiplying the factors together — they are not independent — but a reader adding them up in their head will not get a smaller number than either alone.
And the spread is a range, which is the weakest of the summary statistics. Six samples, of which two decide the answer, so a single unit behaving differently from the others moves it as much as a genuine disagreement among all six — which is what the section above found it doing. A standard deviation across the menu would have been the more honest column and would have hidden the appearance unit’s shift instead of exaggerating it; neither statistic is right on six points, and the reason both are quoted here is that no single number describes a disagreement between six things that were built for different purposes.
So the practical form of all this is a rule about figures rather than a correction to any row.
A census row is good to one significant figure, and its rank is good. A level that moves by ×1.5 across the matching formulae — before the appearance unit, before the test set’s own audit, and before the standard error on the mean — cannot support the third digit it is printed with. What survives every one of those exposures is the ordering, and the ordering is what every argument on this site actually uses the census for: the room is harder than the lamp, a filter applied twice is worse than the same filter once, and a thermal source under a thermal observer is nearly free.
A comparison of two rows is supported when they are more than about a factor of 1.5 apart, and not otherwise. That is the same threshold from two directions — it is the width of the unit exposure, and it is close to the margin by which the census’s two robust extremes beat their neighbours. Rows closer together than that are the rows the reversals section finds units swapping, which is the consistency a reader should expect rather than a coincidence.
Who found it, and when
The formulae are the CIE’s, over twenty-four years: CIELAB and CIELUV in 1976 as two recommendations the committee could not choose between, ΔE*94 as a graphic-arts weighting on the first of them, CIEDE2000 as a considerably more elaborate patch on the same space, and CAM16-UCS in 2017 out of an appearance model rather than out of a space. Oklab is Björn Ottosson’s, 2020, fitted to the same threshold data as its predecessors and published on a personal website rather than by a committee.
That the six disagree is thoroughly established and is why there are six. What is not usually done is to take a body of published results computed in one of them and recompute the whole body in the others, because that requires the results and the formulae in the same program — which is an advantage of arrangement rather than of insight, and is the only one claimed here.
Where the ladder goes next
One property separates the two units that agree with the published one from the three that do not: whether the formula divides a chroma difference by the chroma it was measured at. That property is not a matter of kind — an appearance model and a graphic-arts weighting do it, and three Euclidean distances do not — and it turns out to govern more than agreement. It governs the census’s own sensitivity to the set it is averaged over, which is the largest sensitivity anywhere in this collection.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A dial through a discrete menu calibration · ciede2000 · colour difference · residual · test set
- A mean is not a worst case chromatic adaptation · colour difference · mean · residual · test set
- A partial correction is worth its fraction chromatic adaptation · colour difference · residual · test set · the von kries transform
- The coincidence was a mechanism calibration · chromatic adaptation · colour difference · metamerism · test set
- The error on a gap is not the errors at its ends chromatic adaptation · colour difference · mean · residual · test set
- The row a fourth dimension improves chromatic adaptation · illuminant · metamerism · residual · test set
What links here
The 8 essays that link to this one and share the most of its objects, of 9 that link here.
- The disagreement is at the near end
- The observers differ by a unit's worth
- A model judged in another model's unit
- A gain is not an observer
- The average surface does not look average
- The identity is in the eye's own coordinates
- Which mixture bows most depends on the ruler
- Thirty unknowns instead of six
The objects this essay names
Each one links to every other essay that touches it.
CalibrationChromatic adaptationCIEDE2000Colour differenceIlluminantMeanMetamerismResidualTest setThe von Kries transform