Difference and uniformity

The disagreement is at the near end

Every colour-difference formula on the menu was fitted to threshold data, so the expectation is that they agree about pairs an observer can only just tell apart and diverge on large differences. They do the opposite. Proportionally the disagreement is largest at the near end, by a factor of six for the appearance unit, and the cause is an exponent of 0.63.

Assumes The weighting is the disagreement, The census in six units and A threshold is not a unit.

Every formula on this menu was fitted to the same kind of data: pairs of colours an observer can only just tell apart. The obvious consequence is that they should agree there and part company further out.

Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu.
Fig. 1 The reference pairs split into bands by how far apart they are, with each unit’s proportional departure from the published one plotted per band. Every unit is calibrated once over the whole sample rather than per band, so the shape is the effect and not an artefact of fitting. All five curves fall.

The claim

Proportionally, the six formulae disagree most about pairs that are nearly identical, and least about pairs that are far apart — and in absolute terms they do the reverse.

  • Every one of the five alternatives falls as the pairs get further apart: relative disagreement of 32, 16, 38, 59 and 91 per cent in the 0.5–1 band against 22, 12, 24, 24 and 17 per cent in the 8–12 band.
  • The appearance unit is the extreme case, from 91 per cent to 15 and back to 17: a factor of six across the range, where no other unit changes by more than 1.4.
  • The cause is one exponent. CAM16-UCS’s distance is 1.41 d⁰·⁶³, and a power below one stretches small differences against large ones.
  • In absolute terms every curve rises, from 0.14–0.68 ΔE2000-equivalent on the narrowest band to 1.16–2.37 on the widest.
  • Which reading is right depends on the published quantity. A level wants the absolute reading; a ratio, a threshold and a tolerance want the relative one.

What being fitted to threshold data does and does not constrain

The expectation this essay contradicts is worth spelling out, because it is a reasonable one.

Every formula here descends from an attempt to make a just noticeable difference the same size everywhere. MacAdam’s ellipses are twenty-five such measurements and are the data CIELAB was checked against and CIEDE2000 was fitted to; Oklab was fitted to the same ellipses plus a hue-uniformity dataset; CAM16-UCS’s compression comes out of a model fitted to appearance scaling experiments. So all six have, in some sense, been pointed at the near end of the scale.

The trap is in what the fitting constrains. A formula fitted to thresholds is constrained to give the same value — one unit, or whatever the convention is — at every threshold pair. It is not constrained to agree with any other formula about how far apart two colours that are visibly different are, because no threshold experiment produces such a pair.

Two formulae can both put every threshold pair at exactly 1.0 and still disagree by a factor of two at 10, and the disagreement between them would then be zero at the near end and large at the far end. That is the prediction, and it fails.

Where they actually part

band, ΔE2000 pairs ΔE*ab ΔE*94 ΔE*uv Oklab CAM16-UCS
0.5–1 40 31.7% 16.3% 38.2% 58.6% 90.6%
1–2 70 29.5% 14.8% 35.9% 52.2% 57.2%
2–4 117 27.2% 14.0% 31.5% 43.8% 30.7%
4–8 100 23.7% 11.2% 27.4% 26.1% 14.5%
8–12 44 22.4% 11.9% 24.4% 23.8% 17.3%

Four of the five fall gently — a quarter to a third of their value across a factor of sixteen in separation. The fifth falls off a cliff.

Read the four gentle curves first, because they are the ordinary case and they explain the census. A formula that disagrees proportionally more about small differences will disagree most about a published average that is small, and the census’s mildest rows are exactly the rows that spread furthest across the menu — ×3.71 on the blackbody row against ×1.51 on the harshest. The two observations are the same observation seen from either end.

Why the disagreement is larger at the near end for four Euclidean-ish formulae is not deep. Near threshold, a difference is dominated by whichever axis the two colours happen to differ along, and the four spaces scale their axes differently. Further out a difference has components along all three axes and the scaling differences partly average out. Averaging over directions is what makes large differences agree.

The exponent, and why the appearance unit is different

CAM16-UCS’s behaviour is not the same effect at a larger size, and the mechanism is worth naming because it is a single line of arithmetic.

Its distance is not the Euclidean distance in its own space. It is that distance raised to a power: ΔE′ = 1.41 d⁰·⁶³, where d is the Euclidean distance in the uniform coordinates. The exponent was fitted so that the formula’s numbers line up with observers’ judgements across a wide range of magnitudes.

A power below one is concave. Differentiate it and the local scale factor is proportional to d^(−0.37), which grows without bound as d falls. So the appearance unit stretches the near end of the scale relative to every Euclidean formula on the menu, and the stretching is unbounded. The 91 per cent in the top-left cell of the table is that, and the calibration cannot remove it, because a single multiplicative scale cannot undo a power law.

That has a specific consequence for how the appearance unit should be read against the others, and it is not the one a reader would guess. Its ranking of the census is the closest of the five to the published one, and its disagreement about individual near-threshold pairs is by far the largest. Those two facts are compatible because the census’s rows are means over 125 surfaces, and a stretch applied to every row alike moves the levels together and reorders nothing.

ΔE′ against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE′, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 281 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 24.4 per cent of the mean and the rank correlation is 0.972. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 2 The appearance unit against the published one, pair by pair. The upward curl at the origin is the exponent: below about one unit the power law lifts every pair above the diagonal, and no single scale factor can put it back.
ΔE*94 against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE*94, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 215 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 15.0 per cent of the mean and the rank correlation is 0.986. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 3 The most closely agreeing unit on the same axes, for contrast. Its departure grows in proportion to the difference, which is what a straight line through the origin with scatter about it looks like, and is why its band curve is nearly flat.

The two readings point opposite ways

The same table computed in absolute rather than relative terms reverses:

band ΔE*ab ΔE*94 ΔE*uv Oklab CAM16-UCS
0.5–1 0.231 0.137 0.289 0.476 0.678
8–12 2.143 1.160 2.368 2.237 1.784

Every entry grows by a factor of between four and ten. So there are two true sentences pointing opposite ways — the units agree best about large differences and the units disagree most about large differences — and picking between them is not a matter of taste.

It depends on what the published quantity is used for.

A levelthis change of light costs 0.263 ΔE2000 — is read absolutely. Its exposure is the absolute disagreement at its own size, which for a small level is small.

A ratiothis change of light is nine times that one — is read relatively, and the relative disagreements at both ends compound. This is the reading the quadrature audit recommended precisely because a ratio cancels a common-mode factor, and this essay is the qualification on that advice: a ratio between a large quantity and a small one does not cancel, because the disagreement is not common-mode across the scale.

A thresholdthis delivery is within one unit — is read relatively and sits at the worst place on the curve, which is why pairs built to sit exactly on a ΔE2000 tolerance read from 0.69 to 1.93 in other units.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 4 Six of this collection’s published quantities across the whole menu. The two smallest of them are the two with the largest spread, which is this essay’s finding applied to the collection’s own output.

Two ranges quoted too narrowly

The appearance unit’s behaviour is set against the other four twice, and both comparisons understate what the other four do.

“No other unit changes by more than 1.4.” Taking each unit’s largest band value over its smallest:

unit largest smallest ratio
ΔE*ab 31.7 22.4 1.42
ΔE*94 16.3 11.2 1.46
ΔE*uv 38.2 24.4 1.57
Oklab 58.6 23.8 2.46
CAM16-UCS 90.6 14.5 6.25

All four exceed 1.4, and Oklab exceeds it by three quarters. The sentence the table supports is different and still decisive: the appearance unit changes by 6.25 and the next largest change is Oklab’s 2.46, a factor of 2.5 between them. That is the gap the argument needs, and it does not require the other four to be flat.

“Every entry grows by a factor of between four and ten.” In the absolute table the five ratios are 9.28, 8.47, 8.19, 4.70 and 2.63 — and the last of those is CAM16-UCS, below the stated floor. Which is the same finding wearing its other face: the appearance unit is the one whose absolute disagreement grows least across the scale, precisely because its relative disagreement collapses fastest. Quoting the range as four to ten hides the exception, and the exception is the unit the section is about.

And two of the five curves are not monotone. ΔE*94 runs 16.3, 14.8, 14.0, 11.2, 11.9 and CAM16-UCS runs 90.6, 57.2, 30.7, 14.5, 17.3 — both turn up in the widest band. Every one of the five alternatives falls is true endpoint to endpoint and not true step by step, and the two turn-ups are the two units the essay otherwise treats as opposite cases. Forty-four pairs is a thin band and the rises are small, so the honest reading is that both curves have flattened by four units and what happens beyond that is not resolved by this sample.

The two tables check each other

There is a consistency test available between the relative and absolute tables that neither section runs, and it comes out well.

An absolute disagreement divided by a relative one recovers the mean separation of the pairs in that band. Doing it for all five units gives 0.73, 0.84, 0.76, 0.81 and 0.75 in the 0.5–1 band, and 9.57, 9.75, 9.70, 9.40 and 10.31 in the 8–12 band. Five independent quantities, computed separately, agreeing on where each band’s pairs sit — inside a band of width 0.5 and a band of width 4 — which is the check that the two tables describe one sample rather than two.

The residual disagreement between the five is itself informative. In the widest band CAM16-UCS’s implied mean is 10.31 against the others’ 9.4 to 9.8, the largest in the set; in the narrowest it is 0.75, in the middle of the set. That is what a power law does inside a bin: the appearance unit’s ratio to ΔE2000 keeps falling as pairs get larger, so within a band its relative figure is weighted towards the band’s smaller members and its implied mean comes out high. The exponent is visible even in the residuals of a consistency check, which is about as much corroboration as five numbers can give.

A threshold is a place, not a size

The finding has an interpretation that goes beyond bookkeeping, and it is worth separating from the arithmetic.

A colour-difference formula is asked to do two jobs that are not the same job. Near threshold it is asked to be a detector: to say whether two colours can be told apart, which is a yes-or-no question with a boundary drawn through the space of pairs. Away from threshold it is asked to be a ruler: to say how much more different one pair is than another, which is a question about ratios of magnitudes.

A threshold is not a unit is a sentence this collection has already argued from the psychophysical side, and this is the same sentence arriving from the numerical one. A formula that is excellent as a detector can be poor as a ruler, and vice versa; nothing in the fitting requires both, and the six formulae here differ most exactly where the two jobs pull apart.

The practical consequence is a rule of thumb worth stating plainly. A published difference below about one unit is a detection claim and should be read as one: it says two things are nearly the same, and the number attached is worth about a factor of two. A published difference above about four units is a magnitude claim, the formulae agree about it to within a quarter, and the number is worth quoting. Between the two there is a band where neither reading is comfortable, and most of this collection’s adaptation residuals are in it.

What this does to the smallest published numbers

Applied to the collection, the finding says the exposure is concentrated where the numbers are smallest, and that is where several arguments live.

The blackbody row. 0.263 ΔE2000, the census’s mildest, and the row whose argument is that an observer built under a thermal source discounts a thermal source almost perfectly. It carries a ×3.71 unit spread. The ordering survives every unit, so the argument survives; the number does not deserve three figures.

The macular pigment row. 0.368, and restated in nanometres a round ago partly because it was known to be exposed. A ×3.57 spread.

Every “these two are indistinguishable” claim. A residual quoted as being under some fraction of a unit is a threshold claim at the near end of the scale, which is the worst place on this curve. This collection makes several, and the correct form of each is a comparison against a stated unit rather than an absolute bound.

And a t statistic is unaffected, which is the one piece of good news. A ratio of a gap to its own standard error, both computed in the same unit, is dimensionless and mostly survives — which is what makes the sampling instrument and the unit instrument independent rather than one being a rescaling of the other.

What one change of light costs, surface by surface — daylight to a blackbodyA rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a blackbody is 0.263 ΔE₀₀. The curve runs from 0.0e+0 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 0.457, which is 1.74 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.0.00.20.30.5the published mean, 0.263worst in the set, 0.4575 surfaces at exactly zerothe surfaces, sorted by what this change of light costs themΔE₀₀daylight to a blackbodyCIE 1931 2° observer · the set, varied
Fig. 5 The mildest row in the census, surface by surface. Most of its answer comes from surfaces at a fraction of a unit, which is the band where the formulae disagree proportionally by a third to nine tenths.
MacAdam's twenty-five ellipses, measured in each unit. The uniformity instrument used here, applied to units rather than to spaces. The upper bar is anisotropy — the mean over the twenty-five of the largest radius divided by the smallest, where 1 would be a circle. The lower is spread — the largest mean radius divided by the smallest across all twenty-five, which asks whether a step of the same size means the same thing in different parts of the diagram. Reading down the three CIELAB-based formulae in the order they were published, the anisotropy falls 3.42 → 2.89 → 2.74 and the spread rises 3.24 → 3.59 → 4.12: the weighting divides a difference by the chroma it was measured at, which equalises directions at a point and unequalises magnitudes between points. Neither number is scaled, so no calibration is applied here. CAM16-UCS is ahead on both.
Fig. 6 The same six units on MacAdam’s ellipses, which is the near end of the scale and nothing else: every point on every ellipse is one threshold from its centre. The appearance unit wins there by a distance, which is the other half of what its exponent does.

The last figure is worth pausing on, because it looks as though it contradicts the rest. On the ellipses — pure threshold data, nothing but near-end pairs — the appearance unit is the most uniform of the six by a wide margin, at an anisotropy of 1.57 against the published unit’s 2.74. And on the reference sample’s narrowest band it is the least like the published unit, at 91 per cent.

Both are true and they are not in tension. Being far from ΔE2000 at the near end and being good at the near end are different properties, and a reader who takes agreement with a published convention as evidence of quality would get this exactly backwards. The unit that disagrees most about small differences is the unit that handles them best, on the only near-end measurement of people this collection holds.

One ellipse at a time is the finest grain the comparison has, and it is where a near-threshold disagreement is visible as a shape rather than as a number.

One of MacAdam's ellipses as each unit sees it, at x 0.160, y 0.200. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔEok at an anisotropy of 2.58; the least round is ΔE*94 at 7.78. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 7 A single measured ellipse drawn as its perimeter distance in each of the six units, each outline scaled to its own mean radius. The differences between the outlines are the differences between the units at this one place in the diagram.

What a specification writer should take from this

The finding has one immediate practical use and it is worth separating from the analysis.

A specification’s number and its formula have to be chosen together, and which matters more depends on where the number is. A tolerance at four units or above sits in the band where the formulae agree to within about a quarter; naming the wrong formula there is a modest error and rescaling fixes most of it. A tolerance at one unit or below sits where they disagree by between a sixth and nine tenths, and no rescaling fixes it because the disagreement is about which pairs are close rather than about how the scale runs.

So the tighter the specification, the more the formula matters — which is the opposite of the intuition that a tight tolerance is a demanding requirement independent of how it is measured. A supplier held to 0.5 units is being held to a formula’s opinion about near-threshold pairs; a supplier held to 5 is being held to something six formulae broadly agree about.

The corollary is a rule for reading somebody else’s number. A published difference under about one unit should be read as a detection claim — these two are nearly the same — with the number itself worth about a factor of two. Above about four it is a magnitude claim and the number is worth quoting. Between the two, ask which formula, and expect the answer to matter.

A second centre is what makes the previous figure a measurement rather than an anecdote about one ellipse.

One of MacAdam's ellipses as each unit sees it, at x 0.472, y 0.399. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE*uv at an anisotropy of 1.13; the least round is ΔE00 at 3.47. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 8 Another of the twenty-five, in the same six units and drawn the same way. The units do not keep their order between the two figures, which is the whole reason a ranking has to be quoted with the set it was computed over.

Where the model stops

The bands are computed on the reference sample, which is 374 pairs of constructed reflectances under D65. Their distribution of separations is what the construction produced rather than what anybody would design, and the narrowest band has forty members. A sample built to populate the bands evenly would give tighter numbers; the shape is robust to every reweighting tried.

The relative measure divides by ΔE2000, which privileges the published unit. Dividing by the mean of the six instead moves every curve a little and changes no ordering, and the choice is stated rather than defended: this is an audit of a collection published in ΔE2000, so ΔE2000 is the reference.

And the calibration is fitted once over the whole sample. Refitting per band would flatten the relative curves substantially, and doing so would be measuring the wrong thing — a per-band calibration removes exactly the scale-dependence the essay is about. Which is worth noticing as a general trap: fitting inside the strata of a stratified comparison removes the effect being compared.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 9 Each unit’s disagreement summarised as one number over the whole sample, which is the figure the rest of this audit uses. It is dominated by the middle bands, where most of the pairs are, and it understates both ends.

Who found it, and when

That a power-law compression cannot be undone by a scale factor is elementary and is why the exponent is in CAM16-UCS at all: Stevens’s work on magnitude estimation put sensory response as a power of stimulus intensity, and the 0.63 is a descendant of that literature rather than of colorimetry.

That the formulae disagree more about small differences than large ones in proportional terms is not, as far as this collection can find, a stated result anywhere, and the reason is probably that the comparison is not usually made this way. The standard comparison plots one formula against another over a scatter of pairs and reports a correlation, which is dominated by the large differences because they have the largest leverage. Binning by size and reporting a relative departure is a different question, and it is the question a reader with a small published number needs answered.

Where the ladder goes next

Two observers differ from each other too, and the gap between the 1931 and 1964 standard observers is the one quantity in this audit’s inventory that no file here publishes as a number. It also has the second-largest spread across the menu, which is a poor combination: a quantity nobody has pinned down, measured with a ruler nobody chose.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Appearance modelCalibrationChromaCIEDE2000Colour differenceJust-noticeable differenceResidualTest setThresholdUncertainty