What light is

The census in six units

Recomputing every change of light in the adaptation census under six colour-difference formulae, with the scale factor divided out, leaves a table whose levels move by up to a factor of three point seven. The rows that move most are the mild ones, which is the opposite of what a reader would guess and is a property of where each formula was fitted.

Assumes A choice with no magnitude, What no adaptation can remove and A mean has a set under it.

The adaptation census is fourteen changes of light with a number beside each, and the number is a mean over a constructed set of surfaces of what an adapted observer is left with. Every one of those numbers is in one unit.

The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.
Fig. 1 The census under six colour-difference formulae, each multiplied by the single factor that best carries it onto ΔE2000 over a common reference sample, so the vertical axis means the same thing in every column. Lines that slope are disagreements; lines that cross are reorderings.

The claim

With the change of scale taken out, the census’s levels move by between a factor of 1.5 and a factor of 3.7 depending on the row — and the largest movement is on the rows with the smallest answers.

  • The two ends of the ranking are immovable. A blackbody at daylight’s own temperature is the mildest change in the census under all six units, and a green wall bounced twice is the harshest under all six.
  • The middle is not. Between two and ten of the ninety-one pairs of rows change places, depending on which unit is asked.
  • The spread is largest where the answer is smallest. The blackbody row spans ×3.71 across the menu and the green wall bounced twice spans ×1.51; the rank correlation between a row’s residual and its spread is −0.947. That trend is real and it is almost entirely one unit’s doing, which is the section below that had to be written after the table rather than before it.
  • The appearance unit is the outlier upwards, uniformly. CAM16-UCS reads every one of the fourteen rows as the largest or nearly the largest, which is not a disagreement about ordering at all.
  • And one row is a different shape from the others. An older lens costs 1.02 in ΔE*ab and 0.63 in ΔE*uv, a ratio of 1.6 between two formulae published in the same year by the same committee.

What the census is, before it is a table

The census exists because a change of light is not a single thing, and it is worth restating what the fourteen rows are before asking what a unit does to them.

Each row is a change of context: a light arriving at a surface, and a second light arriving at the same surface. Daylight to a warmer daylight, daylight to tungsten, daylight to a discharge lamp with emission lines in it, daylight to daylight after it has bounced off a painted wall, and — two rows that are inside the observer rather than in the room — daylight seen through the macular pigment and daylight seen through an older lens.

For each, an adapted observer applies the gain their own machinery gives them: the ratio of the two whites, in a fixed basis. What is left over is the row’s number. It is not what a better transform would fix, because the matrix that would fix it is a different matrix for every row and the observer has one mechanism.

The set the mean is taken over is 125 constructed surfaces, and that set has been audited from four directions. The unit is the thing under all of it that had not been.

Taking the scale out first

The six formulae are not on one scale, so the first column of any comparison is bookkeeping.

Uncalibrated, the census means run 0.53, 1.16, 1.31, 0.42, 0.62 and 1.58 — which reads as chaos and is mostly one fact: a Euclidean distance in CIELAB is about 1.8 times CIEDE2000 on any pair anybody measures, because CIEDE2000’s weighting functions divide by chroma and lightness. So each formula is first multiplied by the scale that best carries it onto ΔE2000 over a common sample of 374 pairs, built at the separations these arguments actually live at.

After that the means are 0.958, 1.129, 1.312, 0.962, 0.980 and 1.664. Still not equal — a calibration on one sample does not make two formulae agree on a different one, and if it did there would be nothing to measure — but now every remaining difference is a real disagreement about these surfaces under these lights.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 2 What is left of each unit after its own best rescaling has been removed. The two that weight chroma sit closest to the published one; the three that do not sit furthest. That split governs almost everything in the table below.

The table

change of light ΔE*ab ΔE*94 ΔE2000 ΔE*uv Oklab CAM16-UCS spread
a blackbody at daylight’s temperature 0.227 0.235 0.263 0.213 0.170 0.631 ×3.71
the macular pigment 0.256 0.379 0.368 0.222 0.450 0.793 ×3.57
daylight to D50 0.307 0.402 0.459 0.327 0.396 0.910 ×2.96
daylight to D100 0.321 0.440 0.508 0.339 0.470 0.969 ×3.02
daylight to a three-primary display 0.667 0.835 0.980 0.637 0.650 1.368 ×2.15
daylight to D40 0.672 0.858 0.980 0.725 0.816 1.471 ×2.19
an older lens 1.023 1.046 1.117 0.630 0.800 1.586 ×2.52
a red wall 0.995 1.308 1.357 1.091 1.354 1.884 ×1.89
daylight to halogen 0.990 1.277 1.583 1.013 1.158 1.911 ×1.93
daylight to tungsten 1.136 1.495 1.635 1.248 1.482 2.072 ×1.82
daylight to a white LED 0.967 1.342 1.697 0.984 1.123 1.931 ×2.00
a green wall, one bounce 1.333 1.414 1.722 1.444 1.165 2.102 ×1.80
daylight to a triphosphor tube 1.811 2.020 2.323 1.651 1.456 2.455 ×1.69
a green wall, two bounces 2.703 2.749 3.375 2.946 2.236 3.213 ×1.51

Read down the last column and the pattern is unmistakable: the spread falls as the residual rises, almost monotonically, from ×3.71 on the mildest row to ×1.51 on the harshest.

Why the mild rows move most

That ordering is not an artefact and it is not obvious in advance, so it is worth taking apart.

A mild residual is a mean over surfaces that are nearly matched after adaptation — small differences, close to the origin of the difference formula. A harsh residual is a mean over surfaces the observer gets substantially wrong. And the six formulae were fitted to different data at different separations, so they agree best where they were all trying hardest and worst where none of them was.

Two mechanisms produce it. The weighting functions in ΔE*94 and ΔE2000 divide by chroma, which matters more when the two colours are far apart in chroma than when they are barely apart in anything. And CAM16-UCS’s compression of colourfulness is a logarithm, which is steepest near zero — so a change of light that leaves almost nothing behind still produces a report the appearance model treats as a real, if small, difference.

The consequence is practical and slightly awkward. The rows this collection is least confident about for other reasons are also the rows whose unit-dependence is largest. The blackbody row is 0.263 and is already known to be inside the standard error of its neighbours; it now also has a factor of 3.7 of unit-choice on it. A number with two independent sources of a factor of two or three under it is not a number to print to three figures.

The trend is one unit’s

The paragraph above explains the trend, and the explanation is sound as far as it goes. It is also not what is producing most of the number, which is worth catching before anybody quotes the ×3.71.

The spread is a range statistic over six samples, so it is decided by two of them and ignores the other four. And CAM16-UCS is the largest entry on thirteen of the fourteen rows — the exception being the green wall bounced twice, where ΔE2000 passes it. So the spread column is very close to a single ratio: the appearance unit over whichever matching formula happens to be lowest.

Recomputing the same column over the five matching formulae alone separates the two effects:

over all six over the five matching units
blackbody at daylight’s temperature ×3.71 ×1.55
the macular pigment ×3.57 ×2.03
an older lens ×2.52 ×1.77
daylight to a white LED ×2.00 ×1.75
a green wall, two bounces ×1.51 ×1.51
rank correlation with the residual −0.947 −0.156

The trend does not survive. Across the five matching formulae the spread runs from ×1.36 to ×2.03, with a mean of 1.58 and no relationship to the level at all — a rank correlation of −0.16 on fourteen points is nothing. The falling column in the table above is the appearance unit’s common-mode shift being divided by a rising denominator, which is arithmetic rather than a property of the formulae.

This site has already said why that shift exists: CAM16-UCS is measuring a slightly different quantity, and a section further down this page says so again. What had not been done is to notice that the same fact governs the spread column, and that the claim in the third bullet is therefore about one unit rather than about the menu.

And the residue is more interesting than the trend was. With the appearance unit set aside, the three rows the matching formulae disagree about most are the macular pigment at ×2.03, an older lens at ×1.77, and daylight to a white LED at ×1.75 — which is both of the two rows that are inside the observer, plus the one lamp on the list with a narrow blue emitter in it. That is not a grouping by how mild a row is. It is a grouping by where in the spectrum the change acts, and the essay already has the mechanism for it, three sections down: the formulae are built differently in the blues, and every one of those three rows moves surfaces along the blue–yellow axis.

How much the census depends on its test set, under each unit. Each column is a unit and each dot is one of the fourteen census rows: the elasticity of that row's residual to how saturated the test surfaces are, over the same ±25 per cent span every sensitivity in this collection uses. An elasticity of one means a set half again as saturated gives an answer half again as large. In ΔE2000 the fourteen run 0.49 to 0.91 about a mean of 0.69 — already the largest sensitivity measured anywhere in this collection. In ΔE76, which does not weight chroma at all, every one of the fourteen rises, the mean goes past one to 1.11, and the spread narrows from 0.42 to 0.25. The bar across each column is its mean.
Fig. 3 Each column is a unit and each dot is one of the fourteen rows: how much that row’s residual depends on how saturated the test set is. The unit changes the sensitivity as well as the level, and by more than the level changes.
What one change of light costs, surface by surface — daylight to a blackbodyA rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a blackbody is 0.263 ΔE₀₀. The curve runs from 0.0e+0 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 0.457, which is 1.74 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.0.00.20.30.5the published mean, 0.263worst in the set, 0.4575 surfaces at exactly zerothe surfaces, sorted by what this change of light costs themΔE₀₀daylight to a blackbodyCIE 1931 2° observer · the set, varied
Fig. 4 The mildest row in the census, surface by surface, in the published unit. Most of its answer comes from surfaces the adapted observer gets very nearly right, which is exactly the region where the six formulae were fitted to different data and agree least.

What does not move

Three things survive every unit on the menu, and they are the three the census is actually used for.

A blackbody at daylight’s own temperature is the mildest change in the census, in all six units, by a margin of at least 1.3 to the next row. This is the census’s most-cited result and it is what a reader would hope: an observer whose visual system evolved under a thermal source discounts a thermal source almost perfectly, and no colour-difference formula disagrees.

A green wall bounced twice is the harshest, in all six, by a margin of at least 1.3. Also unsurprising and also robust: a filter applied twice is a filter squared, and nothing in a diagonal gain can undo the squaring.

And the surface changes are worse than the daylight changes, in all six. A room is harder than a lamp, which is the census’s structural finding and is not a matter of unit at all.

What moves is the middle: the discharge lamps against the surfaces, and the two ocular filters against everything. Which pairs, and how that compares against the pairs the test set itself could not resolve, is a question two independent instruments turn out to answer partly the same way.

How much of the census ranking each unit keeps. Kendall's τ between each unit's ordering of the fourteen census rows and the published one; 1 would be perfect agreement. The number beside each bar is what τ is computed from — how many of the 91 pairs of rows the unit puts the other way round. The best agreement is ΔE′, the one unit here that is not a distance between two triples at all, at τ 0.956 and 2 discordant pairs. The worst is ΔEok at 0.780. A ranking that survives every unit is a ranking a reader can rely on; the ones that do not are listed in the figure that follows.
Fig. 5 Kendall’s τ between each unit’s ordering and the published one, with the number of reversed pairs beside it. Between two and ten of the ninety-one pairs change places.
Which steps of the census ranking a change of unit reverses. Every adjacent pair in the published census ranking that at least one unit puts the other way round. The bar counts how many of the five other units reverse it. The marker on the left says whether the test set had already declared the pair unresolved — a gap smaller than twice its own paired standard error, which is a statement about sampling over 125 surfaces and shares no arithmetic with a change of ruler. The two pairs every unit reverses are both flagged, which is the agreement. The pair at the bottom is the disagreement: the test set resolves it at 9.1 standard errors and four of the five units reverse it anyway, because a sampling error cannot see a change of ruler and a change of ruler cannot see a sampling error.
Fig. 6 The adjacencies at least one unit reverses, marked by whether the test set had already declared them unresolved. Two pairs are reversed by every unit and both were flagged; one pair the test set resolves at nine standard errors is reversed by four of the five.

The lens row, which is a different shape

One row disagrees in a way the others do not, and it is the one that is not about a lamp.

An older lens costs 1.023 in ΔE*ab and 0.630 in ΔE*uv — a ratio of 1.62 between two formulae published in the same year, by the same committee, as alternatives for the same job. No other row in the census has a CIELAB-to-CIELUV ratio above 1.15.

The reason is where a yellowing lens acts. It absorbs in the blue, so the surfaces it moves are moved mostly along the blue–yellow axis at low luminance, and that is exactly the region where CIELAB and CIELUV are built differently: CIELUV’s chromaticity coordinates are a projective transform of the diagram and CIELAB’s are a cube-root difference of tristimulus values, and the two treat the dark blues quite differently.

So the row that reports what happens to an observer over a lifetime is the row whose answer depends most on which of two 1976 recommendations is used. That is worth knowing about every published number describing age-related change in colour vision, most of which are quoted in one of the two without saying which.

ΔE*uv against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE*uv, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 153 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 32.3 per cent of the mean and the rank correlation is 0.923. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 7 CIELUV against ΔE2000 on the reference pairs. The scatter is the third widest on the menu and its shape is distinctive: the disagreement grows fastest on the pairs that differ in the blues, which is where the lens row lives.

The appearance unit is not disagreeing about order

CAM16-UCS reads every row higher than any other unit reads it, which looks like a strong disagreement and is nearly the opposite.

Its ranking is the closest of the five to the published one: Kendall’s τ of 0.956, with two reversed pairs out of ninety-one, against ΔE*ab’s eight and Oklab’s ten. What it does is shift every row upwards by a similar amount, and a common-mode shift reorders nothing.

The shift itself has a cause worth naming. CAM16-UCS is not a distance between two tristimulus values; it is a distance between two reports from a model with a surround, an adapting luminance and a degree of adaptation in it. The model’s own adaptation step is incomplete by construction — it never fully discounts a white — so a change of light that a von Kries gain has been credited with fully removing still leaves a residue in the appearance model’s report. The appearance unit is not measuring the same quantity slightly differently; it is measuring a slightly different quantity, and this collection has a separate essay about what happens when it is used as a ruler for a matching question.

ΔE′ against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE′, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 281 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 24.4 per cent of the mean and the rank correlation is 0.972. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 8 The appearance unit against the published one on the same pairs. The scatter is second smallest of the five, which is why its ranking of the census is the closest — and it is systematically above the line at the small end, which is why its levels are all high.

Where the model stops

The census is a census of constructed changes of light acting on constructed surfaces, and neither the change nor the surface is measured. That constraint is older than this question and is stated wherever the census appears; what changes here is that a third constructed thing has been added to the list, and it is the ruler.

Nothing in this exercise establishes that any of these units is closer to what a person would report. Doing that needs a psychophysical experiment, this collection has none, and the literature’s answer depends on which threshold dataset the question is asked with. What it establishes is the size of the exposure, which is a factor of about two on a typical row and up to 3.7 on the mildest.

And it does not reach the objective the surfaces were selected under, or the basis the gain is taken in, both of which have their own audits. Stacking those with this one is not a matter of multiplying the factors together — they are not independent — but a reader adding them up in their head will not get a smaller number than either alone.

And the spread is a range, which is the weakest of the summary statistics. Six samples, of which two decide the answer, so a single unit behaving differently from the others moves it as much as a genuine disagreement among all six — which is what the section above found it doing. A standard deviation across the menu would have been the more honest column and would have hidden the appearance unit’s shift instead of exaggerating it; neither statistic is right on six points, and the reason both are quoted here is that no single number describes a disagreement between six things that were built for different purposes.

So the practical form of all this is a rule about figures rather than a correction to any row.

A census row is good to one significant figure, and its rank is good. A level that moves by ×1.5 across the matching formulae — before the appearance unit, before the test set’s own audit, and before the standard error on the mean — cannot support the third digit it is printed with. What survives every one of those exposures is the ordering, and the ordering is what every argument on this site actually uses the census for: the room is harder than the lamp, a filter applied twice is worse than the same filter once, and a thermal source under a thermal observer is nearly free.

A comparison of two rows is supported when they are more than about a factor of 1.5 apart, and not otherwise. That is the same threshold from two directions — it is the width of the unit exposure, and it is close to the margin by which the census’s two robust extremes beat their neighbours. Rows closer together than that are the rows the reversals section finds units swapping, which is the consistency a reader should expect rather than a coincidence.

Who found it, and when

The formulae are the CIE’s, over twenty-four years: CIELAB and CIELUV in 1976 as two recommendations the committee could not choose between, ΔE*94 as a graphic-arts weighting on the first of them, CIEDE2000 as a considerably more elaborate patch on the same space, and CAM16-UCS in 2017 out of an appearance model rather than out of a space. Oklab is Björn Ottosson’s, 2020, fitted to the same threshold data as its predecessors and published on a personal website rather than by a committee.

That the six disagree is thoroughly established and is why there are six. What is not usually done is to take a body of published results computed in one of them and recompute the whole body in the others, because that requires the results and the formulae in the same program — which is an advantage of arrangement rather than of insight, and is the only one claimed here.

Where the ladder goes next

One property separates the two units that agree with the published one from the three that do not: whether the formula divides a chroma difference by the chroma it was measured at. That property is not a matter of kind — an appearance model and a graphic-arts weighting do it, and three Euclidean distances do not — and it turns out to govern more than agreement. It governs the census’s own sensitivity to the set it is averaged over, which is the largest sensitivity anywhere in this collection.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 9 that link here.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationChromatic adaptationCIEDE2000Colour differenceIlluminantMeanMetamerismResidualTest setThe von Kries transform