Where the model breaks

A choice with no magnitude

An audit can multiply a width by 1.25 and report an elasticity. It cannot multiply CIEDE2000 by anything. Auditing a structural choice needs a different instrument, and building one shows that six published numbers in this collection each carry a factor of about two of unit-choice — after the change of scale has been taken out.

Assumes A mean has a set under it, The input nobody declared and How far apart are two colours.

Almost every number this collection prints about a difference between two colours is printed in the same unit, and the unit was chosen once, in the first weeks, by importing a function.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 1 Six quantities this collection publishes, each recomputed under six colour-difference formulae. Every formula is first multiplied by the single factor that best carries it onto ΔE2000 over a common reference sample, so the bar is not a change of units in the ordinary sense. None of the six is below a factor of 1.7.

The claim

A structural choice has no magnitude, so it cannot be audited the way a width can — and when an instrument is built that can audit it, this collection’s published numbers turn out to carry about a factor of two.

  • A width has an elasticity because a width can be multiplied. A formula cannot; the menu is a handful of published alternatives with nothing continuous between them.
  • The alternatives are not even on the same scale. ΔE*76 reads about 1.8 times ΔE2000 on ordinary pairs, so a table of six columns is six different rulers before any disagreement is visible.
  • Calibration separates the two questions. Fit each formula’s best single scale onto ΔE2000 over a common reference sample; what is left is the part that can reorder something.
  • What is left is large. Six published quantities from six files with no shared machinery move by factors of 1.71, 1.82, 1.96, 2.13, 2.23 and 2.30 across the menu.
  • And the menu does not split the way it looks as though it should. The disagreement is governed by whether a formula divides a chroma difference by the chroma it was measured at, not by whether it is a matching difference or an appearance one.

Three kinds of thing an audit can be pointed at

This collection has now audited three kinds of input, and each needed a different instrument, which is worth stating in order because the third is not obviously auditable at all.

A declared scalar is the easy case. A cone’s optical density is 0.4; a lens absorbs a stated amount at a stated age; a test set’s surfaces reach a stated saturation. Multiply the number by 1.25 and by 0.8, take the ratio of the answers, divide by the ratio of the inputs, and the result is an elasticity: a single number saying how much of the answer is that input’s doing. Every declared scalar in this collection has one.

A set is harder, because a set has no magnitude either. The instrument there was four instruments — membership, measure, shape and dimension — and each of them asks a genuinely different question of the same object. Three of the four reduce to scalars in the end: a region has three declared numbers and those are the ones that dominate.

A formula has neither a magnitude nor parts that can be varied independently. CIEDE2000 is not a number with a range; it is a page of arithmetic with a rotation term in it, and the only thing that can be done to it is to replace it. So the menu is discrete, the members were published in different decades by different committees for different purposes, and the instrument has to work with a handful of points rather than a derivative.

What has to be removed before anything can be compared

The obstacle that makes a menu look useless is that its members are not on one scale.

Run the collection’s adaptation census under each of six formulae and the mean over its fourteen rows comes out at 0.53, 1.16, 1.31, 0.42, 0.62 and 1.58. Read as a table, that is a wild disagreement. Read carefully, most of it is nothing at all: ΔE*76 is roughly 1.8 times ΔE2000 on any pair anybody would measure, and CIELUV’s distances are larger again, and a factor applied to every row of a table changes no conclusion in it.

So the first move is to take that out. Each formula is multiplied by the single scale that best carries it onto ΔE2000, by least squares through the origin, over 374 pairs of surfaces built for the purpose: single-lobed reflectances at seven centres and four depths, each paired with a small perturbation of itself, at separations from a twentieth of a unit to about ten.

The construction of that sample is not a detail, and the first version of it was wrong in a way worth recording. Pairing twenty-five different surfaces with each other gives three hundred pairs whose median separation is 33 ΔE2000 and whose smallest is 2.8. That is not the band anything here lives in — a census residual is 0.26 to 1.7, a delivery tolerance is 1, a camera profile error is a few — and every one of these formulae was fitted to threshold data. Calibrating them a third of the way across the space measures the wrong end of each of them.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 2 Each unit’s remaining disagreement with ΔE2000 after its own best rescaling has been removed, over the reference pairs. The scale factor and the rank correlation are printed beside each bar. ΔE2000’s own row is zero, which is the check that the table is computed the right way round rather than a finding.

What is left after the scale comes out

The residual scatter is the quantity the whole audit turns on, and it is not small.

unit scale onto ΔE2000 scatter rank correlation
ΔE*94 ×0.97 15.0% 0.986
CAM16-UCS ×1.05 24.4% 0.972
ΔE*ab ×0.56 28.5% 0.940
ΔE*uv ×0.44 32.3% 0.923
Oklab ×1.59 34.7% 0.882

A scatter of zero would mean the unit is ΔE2000 in different money: every printed number in this collection would change and not one sentence would. Fifteen per cent is not that, and thirty-five per cent is nowhere near it. A rank correlation of 0.88 means the two formulae disagree about the ordering of a great many pairs, and an ordering of pairs is what every comparative claim here is.

Every change of light this site models, and how much of it a gain removes. Each row is a change of illumination. The pale bar is how far it moves an ordinary surface for an observer who does not adapt; the solid bar at its left end is what is left after the observer has applied the one gain adaptation gives them, which is the ratio of the two whites in the CAT16 basis and is not fitted to anything. Sorted by the fraction left rather than by the size of the change, because the two orderings are different: the largest change here is removed almost entirely and the worst row is a change less than a third its size.
Fig. 3 The adaptation census as this collection publishes it. Every level in the table is a mean in ΔE2000 over a constructed set, and the audit here asks what happens to the whole table when the unit under it is replaced.

The split, and the prediction it refused

The obvious way to read the menu is by kind, and it is wrong.

Five of the six are matching differences: they take a colour to be three numbers and ask how far apart two of them are. CAM16-UCS is an appearance difference: it takes a colour to be what an observer in a stated room would report, computed through a model with a surround and a degree of adaptation in it. The distinction is real and this collection spends a field on it, so the prediction written down before the arithmetic ran was that the appearance formula would be the outlier.

It is not. CAM16-UCS sits at 24 per cent and Oklab at 35, with plain CIELAB and CIELUV in between. An appearance model brought a room, a surround exponent and an incomplete adaptation to the question and still landed nearer the published unit than a Euclidean distance in a space fitted to the same threshold data twenty years later.

What the numbers do split on is one thing: whether the formula divides a chroma difference by the chroma at which it was measured. ΔE*94 does, by 1 + 0.045 C. CAM16-UCS does, by a logarithmic compression of colourfulness inside the model. Those two are the two closest. ΔE*76, ΔE*uv and Oklab do not, and they are the three furthest, in that order.

That is the single dominant fact about the menu, and it has an essay to itself, because the same weighting turns out to decide how sensitive the collection’s largest published sensitivity is.

ΔEok against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔEok, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 208 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 34.7 per cent of the mean and the rank correlation is 0.881. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 4 The furthest member of the menu against the published one, pair by pair, after calibration. A pure rescaling would put every point on the diagonal. The departure grows with the difference, and every pair of points on opposite sides of the line is a comparison the two formulae would decide differently.
ΔE*94 against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE*94, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 215 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 15.0 per cent of the mean and the rank correlation is 0.986. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 5 The nearest member, drawn on the same axes. The scatter is half as wide and the shape is the same: the disagreement is concentrated in the pairs that are far apart and saturated, which is where a chroma weighting does its work.

The inventory, and why it is six files

One quantity moving is a fact about one calculation. The question is whether the choice reaches the collection, and the way to find out is to pick published numbers that share nothing.

Six were chosen on that rule: a change of light after an observer has adapted, from the adaptation machinery; a camera profile’s error on saturated surfaces, from the imaging machinery; the gap between the two standard observers on one set of surfaces; a metameric pair under the lamp that breaks it; the same image on two papers; and an observer two seconds into a new room. Different files, different physics, different questions.

Five of the six are printed in ΔE2000 by their own machinery. The sixth is printed in CAM16-UCS, because the model it comes out of defines that unit and reporting an appearance shift in a matching difference would be a category error. Keeping it in the inventory is deliberate: an audit containing only quantities native to the unit under test has no way of showing what that unit costs when it is the wrong one.

Every one of the six is checked against its own source. Recomputed here in the unit its own file publishes in, each reproduces the number that file prints to the last bit — the largest disagreement anywhere is seven parts in ten thousand million million, which is floating-point noise. Without that check the table would be six re-implementations agreeing with each other about nothing in particular.

quantity published spread across the menu smallest largest
a camera profile’s error 1.184 ×2.30 Oklab 0.95 CAM16-UCS 2.18
the two standard observers ×2.23 ΔE*uv 1.60 Oklab 3.55
the same image on two papers 1.061 ×2.13 ΔE*ab 0.72 CAM16-UCS 1.54
a metameric pair under illuminant A 1.796 ×1.96 Oklab 1.04 CAM16-UCS 2.04
a change of light, after adapting 1.635 ×1.82 ΔE*ab 1.14 CAM16-UCS 2.07
an observer two seconds in 2.651 ×1.71 Oklab 1.88 ΔE2000 3.21

One row has no published source, and it is marked so rather than given one. Nothing in this collection prints the gap between the 1931 and 1964 observers as a single number over a stated set — the essays that argue about it argue about spectra and about individual matches — and inventing a source for it would be worse than admitting the gap.

What a factor of two does and does not touch

The honest reading is that the levels are soft and most of the conclusions are not, and the two need separating carefully.

A level is soft. The macular pigment costs an adapted observer 0.368 ΔE2000 is a sentence with a factor of two of unit-choice under it, on top of the few per cent it already carries from the quadrature rule and the several per cent from the region’s declared numbers. Quoted to three figures it is over-precise by about two significant digits.

A ratio is much harder. A green wall bounced twice costs nine times what a blackbody at the same temperature costs survives every unit on the menu, because a common-mode factor cancels and most of what the unit does is common-mode. This is the same conclusion the quadrature audit reached from the other end, and it is the practical advice both of them produce: publish the ratio wherever the argument allows it.

An ordering is in between, and that is the surprise. A ranking survives a change of test set almost perfectly and survives a change of unit much less well: the census’s fourteen rows come out in a different order under every one of the five alternatives, with between two and ten of the ninety-one pairs reversed. Which pairs, and how that compares with what the test set could already not resolve, is the subject of the essay that puts the two instruments on one axis.

The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.
Fig. 6 The census in six units, calibrated onto one scale. The levels move by up to a factor of two and the lines cross. A crossing is a reordering, and a reordering is a conclusion changing rather than a number changing.
How much of the census ranking each unit keeps. Kendall's τ between each unit's ordering of the fourteen census rows and the published one; 1 would be perfect agreement. The number beside each bar is what τ is computed from — how many of the 91 pairs of rows the unit puts the other way round. The best agreement is ΔE′, the one unit here that is not a distance between two triples at all, at τ 0.956 and 2 discordant pairs. The worst is ΔEok at 0.780. A ranking that survives every unit is a ranking a reader can rely on; the ones that do not are listed in the figure that follows.
Fig. 7 How much of the published census ranking each unit keeps, as Kendall’s τ, with the count it is computed from beside it. The best agreement is with the one unit that is not a distance between three numbers at all.

What the instrument cannot do

It cannot say which unit is right, and it is worth being explicit that this is not modesty.

There is no experiment in this collection that could settle it. The formulae were fitted to different threshold datasets, by committees with different remits, and the literature’s answer to which is best depends on which dataset is used to ask — a fact this collection has already measured from the ellipse side. A collection that picked a winner here would be making exactly the move it spends a field arguing against: taking a choice somebody made once and treating the answer as a property of the world.

What the instrument produces instead is a spread and a list of what the spread moves. That is what a reader needs in order to know how much weight a printed number will carry, and it is strictly more useful than a ranking of formulae would be.

It also cannot reach a seventh unit that is not on the menu at all and that turns up in the middle of the imaging machinery: a least-squares objective in XYZ, which is what every camera profile here and, as far as this collection can tell, everywhere, is fitted with. That one is not a colour-difference formula, nobody chose it as one, and refitting under each of the six that were chosen is where the audit finds its sharpest single result.

Pairs built to sit exactly on a ΔE2000 tolerance, read in every other unit. Twenty-four pairs of surfaces, each constructed by walking one member along a fixed direction until the difference is exactly 1.0 ΔE2000 under D65. The bar spans what those same pairs read in each unit, after calibration, with the tick at the mean. ΔE2000's own row is a point at 1.0 by construction. Every other unit spreads them: ΔE*uv reads them from 0.69 to 1.75, so a contract written at "one unit" accepts and rejects a different set of deliveries depending on which unit it means. CAM16-UCS rejects all twenty-four: it reads the closest of them at 1.43.
Fig. 8 Twenty-four pairs built to sit exactly on a ΔE2000 tolerance of 1.0, read in every other unit. A contract quoted at “one unit” does not become slightly wrong in another unit; it accepts and rejects a different set of deliveries.

Who found it, and when

The distinction between a scale and a ranking is Stevens’s, and the reason it belongs at the front of this is that a colour-difference formula is claimed as an interval scale — a difference of two is claimed to be twice a difference of one — while almost everything it is used for needs only an ordinal one. Most of the disagreement measured here lives in the interval claim.

The specific fact that the formulae reorder pairs rather than merely rescaling them is old and has been published repeatedly since ΔE*94 appeared, usually as a table of worked examples. What is not usually done is to point it at a whole body of published results at once and ask how many of them move, which needs the results and the formulae in the same program. That is the only advantage this collection has here, and it is entirely an advantage of arrangement rather than of insight.

The habit it comes out of is older and belongs to this collection: every claim gets a test it could fail, and a claim that cannot be varied is a claim nobody has tested. Three such claims were named two rounds ago and this is the first of the three to be reached.

The two groups, on one axis. Every comparison with noise on one side of it moves by at least 36 times when the detector changes; no comparison between two structured fields moves by more than 3.4. The gap between the groups is a factor of 11 and nothing sits in it. That is the last phase's rule holding in one direction: noise always separates the two readings. What it does not do is hold in the other — two of the structured comparisons move as well, which is why the rule as stated was too narrow.
Fig. 9 The collection’s earlier self-audit, which asked whether a filtered claim read at a point survives being read as components. The instrument here is the same shape one level out: a claim is read under one convention and then under another, and what matters is which claims change.

The split is ordinal, and the essay states it as a gap

The claim that the menu divides on whether a formula weights a chroma difference by the chroma it was measured at is described above as the single dominant fact about the menu. The scatter column supports the ordering and does not support the phrasing, and the difference is worth pinning down because it decides how much weight the split can carry.

Sorted, the five scatters are 15.0, 24.4, 28.5, 32.3 and 34.7 per cent. The two weighted units are indeed the two smallest, so the ordinal claim holds exactly. But the gaps between consecutive values run 9.4, 4.1, 3.8, 2.4 — and the largest gap in the list is inside the weighted pair, not between the two groups. A clustering that knew nothing about chroma weighting and simply cut at the widest gap would separate ΔE*94 from the other four, splitting the pair the hypothesis predicts.

The honest statement is therefore weaker and still useful: the two weighted units occupy the two closest places, and nothing about the spacing distinguishes the groups. Two units out of five landing in the top two by chance is one in ten, so the ordering is evidence; it is not the clean separation a phrase like the dominant fact implies, and a sixth formula landing between 24 and 28 per cent would be unclassifiable by this instrument.

Two columns that measure one thing

The scatter and the rank correlation move in perfect lockstep across all five rows — as scatter rises, correlation falls, with no exception and no near-tie. That is a check rather than a finding, and it is a useful one: two quantities computed by different routes over the same 374 pairs agreeing on the ordering means neither is picking up an artefact the other misses.

It also means only one of them is carrying information, and the rank correlation is the one worth quoting. A scatter is a statement about magnitudes and inherits the interval claim the closing section says most of the disagreement lives in. A rank correlation is a statement about orderings, which is what a comparative claim in this collection actually rests on.

The menu has a direction, not only a width

The inventory is presented as six spreads, and read that way it says the choice costs about a factor of two. Read as twelve cells rather than six ranges, it says something the spreads hide: CAM16-UCS is the largest reading in four of the six rows and the smallest in none of them, and Oklab is the smallest in three.

The exception is instructive rather than awkward. The one row where CAM16-UCS is not the largest is the row it is the native unit of — the observer two seconds into a new room — and there ΔE2000 takes the top place instead. So the pattern is not that an appearance unit flatters its own quantities; it is that after calibration onto a common scale, the appearance unit reads high on this collection’s material and the newest Euclidean space reads low, consistently, across six files that share no machinery.

That is a systematic effect sitting underneath what the table presents as a spread, and the two have different consequences. A spread says a number is soft. A direction says that a reader who prefers a different unit will find this collection’s quantities biased in a knowable direction, and by roughly a factor of two between the ends.

Where the published numbers sit inside their own spreads

The last thing the inventory can be asked is whether the collection’s own unit is at an extreme, and it is not. Expressing each published value as a fraction of the way from its row’s smallest reading to its largest — 0 at the bottom of the range, 1 at the top — gives 0.19, 0.42, 0.76 and 0.53 on the four ΔE2000-native rows with a published figure, a mean of 0.47; adding the CAM16-UCS-native row at 0.58 gives 0.50 over five.

The published unit is a near-median choice. The factor of two is a spread around this collection’s numbers rather than an offset away from them, which is the more comfortable of the two findings available and is not the one that was assumed: a unit chosen in the first weeks by importing a function had no particular reason to land in the middle of a menu nobody had assembled yet.

Where the ladder goes next

The menu has one member that is continuous with another. ΔE*94 is ΔE*76 with two weighting constants in it, and at zero those constants make every weight exactly one — so the two ends of the oldest disagreement in this subject are joined by a line rather than separated by a choice. Walking that line gives a derivative where the menu gives only a spread, and the derivative says something the six points cannot: whether the published choice sits on a flat part of the curve or on a steep one.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationCIEDE2000Colour differenceElasticityMeanRank correlationResidualSensitivityStructural choiceTest set