The disagreement is at the near end
Assumes The weighting is the disagreement, The census in six units and A threshold is not a unit.
Every formula on this menu was fitted to the same kind of data: pairs of colours an observer can only just tell apart. The obvious consequence is that they should agree there and part company further out.
The claim
Proportionally, the six formulae disagree most about pairs that are nearly identical, and least about pairs that are far apart — and in absolute terms they do the reverse.
- Every one of the five alternatives falls as the pairs get further apart: relative disagreement of 32, 16, 38, 59 and 91 per cent in the 0.5–1 band against 22, 12, 24, 24 and 17 per cent in the 8–12 band.
- The appearance unit is the extreme case, from 91 per cent to 15 and back to 17: a factor of six across the range, where no other unit changes by more than 1.4.
- The cause is one exponent. CAM16-UCS’s distance is
1.41 d⁰·⁶³, and a power below one stretches small differences against large ones. - In absolute terms every curve rises, from 0.14–0.68 ΔE2000-equivalent on the narrowest band to 1.16–2.37 on the widest.
- Which reading is right depends on the published quantity. A level wants the absolute reading; a ratio, a threshold and a tolerance want the relative one.
What being fitted to threshold data does and does not constrain
The expectation this essay contradicts is worth spelling out, because it is a reasonable one.
Every formula here descends from an attempt to make a just noticeable difference the same size everywhere. MacAdam’s ellipses are twenty-five such measurements and are the data CIELAB was checked against and CIEDE2000 was fitted to; Oklab was fitted to the same ellipses plus a hue-uniformity dataset; CAM16-UCS’s compression comes out of a model fitted to appearance scaling experiments. So all six have, in some sense, been pointed at the near end of the scale.
The trap is in what the fitting constrains. A formula fitted to thresholds is constrained to give the same value — one unit, or whatever the convention is — at every threshold pair. It is not constrained to agree with any other formula about how far apart two colours that are visibly different are, because no threshold experiment produces such a pair.
Two formulae can both put every threshold pair at exactly 1.0 and still disagree by a factor of two at 10, and the disagreement between them would then be zero at the near end and large at the far end. That is the prediction, and it fails.
Where they actually part
| band, ΔE2000 | pairs | ΔE*ab | ΔE*94 | ΔE*uv | Oklab | CAM16-UCS |
|---|---|---|---|---|---|---|
| 0.5–1 | 40 | 31.7% | 16.3% | 38.2% | 58.6% | 90.6% |
| 1–2 | 70 | 29.5% | 14.8% | 35.9% | 52.2% | 57.2% |
| 2–4 | 117 | 27.2% | 14.0% | 31.5% | 43.8% | 30.7% |
| 4–8 | 100 | 23.7% | 11.2% | 27.4% | 26.1% | 14.5% |
| 8–12 | 44 | 22.4% | 11.9% | 24.4% | 23.8% | 17.3% |
Four of the five fall gently — a quarter to a third of their value across a factor of sixteen in separation. The fifth falls off a cliff.
Read the four gentle curves first, because they are the ordinary case and they explain the census. A formula that disagrees proportionally more about small differences will disagree most about a published average that is small, and the census’s mildest rows are exactly the rows that spread furthest across the menu — ×3.71 on the blackbody row against ×1.51 on the harshest. The two observations are the same observation seen from either end.
Why the disagreement is larger at the near end for four Euclidean-ish formulae is not deep. Near threshold, a difference is dominated by whichever axis the two colours happen to differ along, and the four spaces scale their axes differently. Further out a difference has components along all three axes and the scaling differences partly average out. Averaging over directions is what makes large differences agree.
The exponent, and why the appearance unit is different
CAM16-UCS’s behaviour is not the same effect at a larger size, and the mechanism is worth naming because it is a single line of arithmetic.
Its distance is not the Euclidean distance in its own space. It is that distance raised to a power: ΔE′ = 1.41 d⁰·⁶³, where d is the Euclidean distance in the uniform coordinates. The exponent was fitted so that the formula’s numbers line up with observers’ judgements across a wide range of magnitudes.
A power below one is concave. Differentiate it and the local scale factor is proportional to d^(−0.37), which grows without bound as d falls. So the appearance unit stretches the near end of the scale relative to every Euclidean formula on the menu, and the stretching is unbounded. The 91 per cent in the top-left cell of the table is that, and the calibration cannot remove it, because a single multiplicative scale cannot undo a power law.
That has a specific consequence for how the appearance unit should be read against the others, and it is not the one a reader would guess. Its ranking of the census is the closest of the five to the published one, and its disagreement about individual near-threshold pairs is by far the largest. Those two facts are compatible because the census’s rows are means over 125 surfaces, and a stretch applied to every row alike moves the levels together and reorders nothing.
The two readings point opposite ways
The same table computed in absolute rather than relative terms reverses:
| band | ΔE*ab | ΔE*94 | ΔE*uv | Oklab | CAM16-UCS |
|---|---|---|---|---|---|
| 0.5–1 | 0.231 | 0.137 | 0.289 | 0.476 | 0.678 |
| 8–12 | 2.143 | 1.160 | 2.368 | 2.237 | 1.784 |
Every entry grows by a factor of between four and ten. So there are two true sentences pointing opposite ways — the units agree best about large differences and the units disagree most about large differences — and picking between them is not a matter of taste.
It depends on what the published quantity is used for.
A level — this change of light costs 0.263 ΔE2000 — is read absolutely. Its exposure is the absolute disagreement at its own size, which for a small level is small.
A ratio — this change of light is nine times that one — is read relatively, and the relative disagreements at both ends compound. This is the reading the quadrature audit recommended precisely because a ratio cancels a common-mode factor, and this essay is the qualification on that advice: a ratio between a large quantity and a small one does not cancel, because the disagreement is not common-mode across the scale.
A threshold — this delivery is within one unit — is read relatively and sits at the worst place on the curve, which is why pairs built to sit exactly on a ΔE2000 tolerance read from 0.69 to 1.93 in other units.
Two ranges quoted too narrowly
The appearance unit’s behaviour is set against the other four twice, and both comparisons understate what the other four do.
“No other unit changes by more than 1.4.” Taking each unit’s largest band value over its smallest:
| unit | largest | smallest | ratio |
|---|---|---|---|
| ΔE*ab | 31.7 | 22.4 | 1.42 |
| ΔE*94 | 16.3 | 11.2 | 1.46 |
| ΔE*uv | 38.2 | 24.4 | 1.57 |
| Oklab | 58.6 | 23.8 | 2.46 |
| CAM16-UCS | 90.6 | 14.5 | 6.25 |
All four exceed 1.4, and Oklab exceeds it by three quarters. The sentence the table supports is different and still decisive: the appearance unit changes by 6.25 and the next largest change is Oklab’s 2.46, a factor of 2.5 between them. That is the gap the argument needs, and it does not require the other four to be flat.
“Every entry grows by a factor of between four and ten.” In the absolute table the five ratios are 9.28, 8.47, 8.19, 4.70 and 2.63 — and the last of those is CAM16-UCS, below the stated floor. Which is the same finding wearing its other face: the appearance unit is the one whose absolute disagreement grows least across the scale, precisely because its relative disagreement collapses fastest. Quoting the range as four to ten hides the exception, and the exception is the unit the section is about.
And two of the five curves are not monotone. ΔE*94 runs 16.3, 14.8, 14.0, 11.2, 11.9 and CAM16-UCS runs 90.6, 57.2, 30.7, 14.5, 17.3 — both turn up in the widest band. Every one of the five alternatives falls is true endpoint to endpoint and not true step by step, and the two turn-ups are the two units the essay otherwise treats as opposite cases. Forty-four pairs is a thin band and the rises are small, so the honest reading is that both curves have flattened by four units and what happens beyond that is not resolved by this sample.
The two tables check each other
There is a consistency test available between the relative and absolute tables that neither section runs, and it comes out well.
An absolute disagreement divided by a relative one recovers the mean separation of the pairs in that band. Doing it for all five units gives 0.73, 0.84, 0.76, 0.81 and 0.75 in the 0.5–1 band, and 9.57, 9.75, 9.70, 9.40 and 10.31 in the 8–12 band. Five independent quantities, computed separately, agreeing on where each band’s pairs sit — inside a band of width 0.5 and a band of width 4 — which is the check that the two tables describe one sample rather than two.
The residual disagreement between the five is itself informative. In the widest band CAM16-UCS’s implied mean is 10.31 against the others’ 9.4 to 9.8, the largest in the set; in the narrowest it is 0.75, in the middle of the set. That is what a power law does inside a bin: the appearance unit’s ratio to ΔE2000 keeps falling as pairs get larger, so within a band its relative figure is weighted towards the band’s smaller members and its implied mean comes out high. The exponent is visible even in the residuals of a consistency check, which is about as much corroboration as five numbers can give.
A threshold is a place, not a size
The finding has an interpretation that goes beyond bookkeeping, and it is worth separating from the arithmetic.
A colour-difference formula is asked to do two jobs that are not the same job. Near threshold it is asked to be a detector: to say whether two colours can be told apart, which is a yes-or-no question with a boundary drawn through the space of pairs. Away from threshold it is asked to be a ruler: to say how much more different one pair is than another, which is a question about ratios of magnitudes.
A threshold is not a unit is a sentence this collection has already argued from the psychophysical side, and this is the same sentence arriving from the numerical one. A formula that is excellent as a detector can be poor as a ruler, and vice versa; nothing in the fitting requires both, and the six formulae here differ most exactly where the two jobs pull apart.
The practical consequence is a rule of thumb worth stating plainly. A published difference below about one unit is a detection claim and should be read as one: it says two things are nearly the same, and the number attached is worth about a factor of two. A published difference above about four units is a magnitude claim, the formulae agree about it to within a quarter, and the number is worth quoting. Between the two there is a band where neither reading is comfortable, and most of this collection’s adaptation residuals are in it.
What this does to the smallest published numbers
Applied to the collection, the finding says the exposure is concentrated where the numbers are smallest, and that is where several arguments live.
The blackbody row. 0.263 ΔE2000, the census’s mildest, and the row whose argument is that an observer built under a thermal source discounts a thermal source almost perfectly. It carries a ×3.71 unit spread. The ordering survives every unit, so the argument survives; the number does not deserve three figures.
The macular pigment row. 0.368, and restated in nanometres a round ago partly because it was known to be exposed. A ×3.57 spread.
Every “these two are indistinguishable” claim. A residual quoted as being under some fraction of a unit is a threshold claim at the near end of the scale, which is the worst place on this curve. This collection makes several, and the correct form of each is a comparison against a stated unit rather than an absolute bound.
And a t statistic is unaffected, which is the one piece of good news. A ratio of a gap to its own standard error, both computed in the same unit, is dimensionless and mostly survives — which is what makes the sampling instrument and the unit instrument independent rather than one being a rescaling of the other.
The last figure is worth pausing on, because it looks as though it contradicts the rest. On the ellipses — pure threshold data, nothing but near-end pairs — the appearance unit is the most uniform of the six by a wide margin, at an anisotropy of 1.57 against the published unit’s 2.74. And on the reference sample’s narrowest band it is the least like the published unit, at 91 per cent.
Both are true and they are not in tension. Being far from ΔE2000 at the near end and being good at the near end are different properties, and a reader who takes agreement with a published convention as evidence of quality would get this exactly backwards. The unit that disagrees most about small differences is the unit that handles them best, on the only near-end measurement of people this collection holds.
One ellipse at a time is the finest grain the comparison has, and it is where a near-threshold disagreement is visible as a shape rather than as a number.
What a specification writer should take from this
The finding has one immediate practical use and it is worth separating from the analysis.
A specification’s number and its formula have to be chosen together, and which matters more depends on where the number is. A tolerance at four units or above sits in the band where the formulae agree to within about a quarter; naming the wrong formula there is a modest error and rescaling fixes most of it. A tolerance at one unit or below sits where they disagree by between a sixth and nine tenths, and no rescaling fixes it because the disagreement is about which pairs are close rather than about how the scale runs.
So the tighter the specification, the more the formula matters — which is the opposite of the intuition that a tight tolerance is a demanding requirement independent of how it is measured. A supplier held to 0.5 units is being held to a formula’s opinion about near-threshold pairs; a supplier held to 5 is being held to something six formulae broadly agree about.
The corollary is a rule for reading somebody else’s number. A published difference under about one unit should be read as a detection claim — these two are nearly the same — with the number itself worth about a factor of two. Above about four it is a magnitude claim and the number is worth quoting. Between the two, ask which formula, and expect the answer to matter.
A second centre is what makes the previous figure a measurement rather than an anecdote about one ellipse.
Where the model stops
The bands are computed on the reference sample, which is 374 pairs of constructed reflectances under D65. Their distribution of separations is what the construction produced rather than what anybody would design, and the narrowest band has forty members. A sample built to populate the bands evenly would give tighter numbers; the shape is robust to every reweighting tried.
The relative measure divides by ΔE2000, which privileges the published unit. Dividing by the mean of the six instead moves every curve a little and changes no ordering, and the choice is stated rather than defended: this is an audit of a collection published in ΔE2000, so ΔE2000 is the reference.
And the calibration is fitted once over the whole sample. Refitting per band would flatten the relative curves substantially, and doing so would be measuring the wrong thing — a per-band calibration removes exactly the scale-dependence the essay is about. Which is worth noticing as a general trap: fitting inside the strata of a stratified comparison removes the effect being compared.
Who found it, and when
That a power-law compression cannot be undone by a scale factor is elementary and is why the exponent is in CAM16-UCS at all: Stevens’s work on magnitude estimation put sensory response as a power of stimulus intensity, and the 0.63 is a descendant of that literature rather than of colorimetry.
That the formulae disagree more about small differences than large ones in proportional terms is not, as far as this collection can find, a stated result anywhere, and the reason is probably that the comparison is not usually made this way. The standard comparison plots one formula against another over a scatter of pairs and reports a correlation, which is dominated by the large differences because they have the largest leverage. Binning by size and reporting a relative departure is a different question, and it is the question a reader with a small published number needs answered.
Where the ladder goes next
Two observers differ from each other too, and the gap between the 1931 and 1964 standard observers is the one quantity in this audit’s inventory that no file here publishes as a number. It also has the second-largest spread across the menu, which is a poor combination: a quantity nobody has pinned down, measured with a ruler nobody chose.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A dial through a discrete menu calibration · chroma · ciede2000 · colour difference · residual · test set
- A unit rests on a space that was ranked appearance model · calibration · chroma · ciede2000 · colour difference · just-noticeable difference
- A choice with no magnitude calibration · ciede2000 · colour difference · residual · test set
- A difference has no rate chroma · ciede2000 · colour difference · just-noticeable difference · threshold
- A departure is straight in the excitations chroma · ciede2000 · colour difference · residual
- A tolerance has no light level appearance model · chroma · ciede2000 · colour difference
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Appearance modelCalibrationChromaCIEDE2000Colour differenceJust-noticeable differenceResidualTest setThresholdUncertainty