Difference and uniformity

The weighting is the disagreement

Five colour-difference formulae, three decades and two committees, and the single property that predicts which of them agree is whether a chroma difference gets divided by the chroma it was measured at. It sorts the menu exactly, it cuts across the distinction between a matching difference and an appearance one, and it halves the census's largest sensitivity.

Assumes A choice with no magnitude, How far apart are two colours and MacAdam measured it.

Five alternatives to the unit this collection publishes in, and one property that sorts them exactly.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 1 Each formula’s remaining disagreement with ΔE2000 after its own best rescaling has been divided out, over 374 pairs of surfaces at the separations these arguments live at. The two that weight chroma are the two nearest. The three that do not are the three furthest.

The claim

Whether a formula divides a chroma difference by the chroma at which it was measured predicts which formulae agree with each other, better than anything else about them does.

  • The split is exact on this menu. The two weighted formulae sit at 15.0 and 24.4 per cent scatter about ΔE2000; the three unweighted ones sit at 28.5, 32.3 and 34.7. There is no overlap.
  • It cuts across the obvious classification. One of the two weighted formulae is a graphic-arts recommendation and the other is an appearance model. Two of the three unweighted ones were published by the same committee in the same year as the space the weighted ones are built on.
  • It is not a small effect on the levels. Turning the weighting off raises the census’s sensitivity to its own test set from 0.69 to 1.11 — past one, which means the answer would then move more than proportionally with the input.
  • And it is what CIEDE2000 mostly is. The rotation term everybody remembers matters in one corner of the space; the division by chroma matters everywhere.

What the weighting is

Two formulae, side by side, is the shortest way to see it.

ΔE*ab is the Euclidean distance in CIELAB: the square root of the sum of squares of the lightness, a and b differences. There is nothing else in it. It was recommended in 1976 with the space, and its whole content is the claim that CIELAB is uniform enough that a straight distance in it means something.

ΔE*94 is the same three differences, resolved into lightness, chroma and hue rather than lightness, a and b, and then each of the last two divided by something: the chroma difference by 1 + 0.045 C, the hue difference by 1 + 0.015 C, where C is the chroma of the reference colour. That is the whole of the change.

The consequence is that a difference of a given size between two saturated colours counts for less than the same difference between two pale ones. At a chroma of 60 — an ordinary saturated red — a chroma difference is divided by 3.7. The formula is saying that CIELAB’s chroma axis is stretched by nearly a factor of four out there, which is exactly what MacAdam’s ellipses had already said and which the space was built without.

CIEDE2000 does the same and more: the same two divisions with a more elaborate chroma dependence, a lightness weighting, a term correcting the a axis near the grey axis, and a rotation term that only bites in the blues. The rotation term is what people remember about the formula. It is not what the formula mostly does.

The menu, sorted

formula year weights chroma? scatter rank correlation
ΔE*94 1994 yes, by 1 + 0.045 C 15.0% 0.986
CAM16-UCS 2017 yes, inside the model 24.4% 0.972
ΔE*ab 1976 no 28.5% 0.940
ΔE*uv 1976 no 32.3% 0.923
Oklab 2020 no 34.7% 0.882

Scatter is the root-mean-square departure from that formula’s own best rescaling of ΔE2000, as a fraction of the mean. A formula that were a pure rescaling would sit at zero and swapping to it would change every printed number and no conclusion. None of the five is close to that.

The sort is exact, and the two things it does not sort by are worth saying out loud. Not by age: the closest is from 1994 and the furthest from 2020, with 1976 in between twice. Not by kind: the second-closest is the only entry that is not a distance between three numbers at all.

CAM16-UCS weights chroma too, and that is why it is where it is

Calling an appearance model “a formula that weights chroma” needs justifying, since it does nothing that looks like 1 + 0.045 C.

What it does is compress. The model reports a colourfulness M, and the uniform space built on top of it uses not M but (1/0.0228) log(1 + 0.0228 M) — a logarithm, which is very nearly linear near zero and increasingly flat as M grows. Differentiate it and the local scale factor on a colourfulness difference is 1/(1 + 0.0228 M), which is the same shape as ΔE*94’s 1/(1 + 0.045 C) with a different constant.

So the two arrive at the same operation from opposite directions: one by fitting a correction to a space known to be wrong, the other by modelling the response that made it wrong in the first place. That they land 9 percentage points apart on this instrument, with three unweighted formulae 4 to 10 points further out, is the strongest evidence available here that the operation rather than the derivation is what matters.

ΔE*94 against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE*94, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 215 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 15.0 per cent of the mean and the rank correlation is 0.986. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 2 The most closely agreeing member of the menu against ΔE2000, pair by pair. The remaining scatter is real but narrow, and it grows with the difference rather than being a constant fog.
ΔE′ against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE′, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 281 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 24.4 per cent of the mean and the rank correlation is 0.972. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 3 The appearance unit on the same axes. Its scatter is wider than ΔE*94’s and narrower than any unweighted formula’s, and it sits systematically above the line at the small end — which is the model’s incomplete adaptation showing through.

What the weighting does to a sensitivity

A published level moving is one thing. A published sensitivity moving is worse, and this is where the weighting turns out to matter most.

Every residual in the adaptation census is a mean over a set of constructed surfaces, and that set’s saturation is the most elastic input in this collection: make the test surfaces 25 per cent more saturated and the answers rise by between 0.49 and 0.91 times as much, proportionally. That measurement is in ΔE2000.

Recompute it under each unit and the elasticity moves further than the levels do:

unit lowest row highest row mean
CAM16-UCS 0.469 0.631 0.532
ΔE2000 0.494 0.911 0.688
ΔE*94 0.673 1.071 0.835
ΔE*uv 0.975 1.186 1.024
Oklab 1.033 1.112 1.071
ΔE*ab 1.066 1.313 1.105

The ordering is the weighting again, and now it is a continuum rather than a split: the more a formula divides by chroma, the less its answer depends on how saturated the test set is. That is not a coincidence, it is a tautology being made visible. A more saturated test set produces larger chroma differences; a formula that divides chroma differences by chroma discounts exactly the part that grew.

The size is what matters. Under an unweighted formula the elasticity is above one: a test set half again as saturated gives an answer more than half again as large, so the published residual would be more a statement about the surfaces than about the change of light. Under the published unit it is 0.69, and under the appearance unit 0.53. The choice of unit is worth about a factor of two on how much of the collection’s adaptation numbers belong to a set nobody measured.

How much the census depends on its test set, under each unit. Each column is a unit and each dot is one of the fourteen census rows: the elasticity of that row's residual to how saturated the test surfaces are, over the same ±25 per cent span every sensitivity in this collection uses. An elasticity of one means a set half again as saturated gives an answer half again as large. In ΔE2000 the fourteen run 0.49 to 0.91 about a mean of 0.69 — already the largest sensitivity measured anywhere in this collection. In ΔE76, which does not weight chroma at all, every one of the fourteen rises, the mean goes past one to 1.11, and the spread narrows from 0.42 to 0.25. The bar across each column is its mean.
Fig. 4 Fourteen census rows against six units, with each column’s mean drawn across it. The horizontal rule is an elasticity of one, above which an answer moves more than proportionally with the set it was averaged over. Three of the six units are entirely above it.

The two instruments do not rank the same

The essay now has two orderings of the same five formulae — how far each sits from ΔE2000 pair by pair, and how elastic each makes the census — and they agree about the split and disagree about almost everything inside it.

rank by scatter rank by elasticity
ΔE*94 1 2
CAM16-UCS 2 1
ΔE*ab 3 5
ΔE*uv 4 3
Oklab 5 4

Rank correlation 0.60 — the same two at the top, the same three at the bottom, and no agreement finer than that.

The instructive row is ΔE*ab. It is the closest of the three unweighted formulae to ΔE2000 on individual pairs, at 28.5 per cent against Oklab’s 34.7 — and it is the worst of all six for the elasticity, at 1.105, further from the published unit than any other entry on the menu. A formula that tracks ΔE2000 reasonably well one pair at a time is the single worst choice for a result that is a mean over a set of pairs.

That is not a paradox and it is worth having the mechanism. Scatter is a symmetric measure — it counts a formula reading high and reading low as the same size of disagreement — and an average over a set does not: reading high on the saturated members and low on the pale ones leaves the mean nearly right and makes the sensitivity to which members are in the set completely wrong. ΔE*ab’s errors happen to cancel in the middle and to be systematically ordered by chroma, which is the one arrangement that is invisible to a scatter and fatal to an elasticity.

So the practical warning is a sharp one. Agreement with the published unit is not transitive across the kind of quantity being computed. A reader checking whether a body of results survives a change of unit cannot check it on pairs and conclude anything about means, and the collection has now measured both because the second was not predictable from the first.

And the published unit has the least stable sensitivity

One more column falls out of the same table, and it is a caution about this collection rather than about the formulae.

Reading the elasticity table by row spread rather than by mean — how much a unit’s elasticity varies across the fourteen changes of light:

unit elasticity, lowest to highest row spread
ΔE2000 0.494 – 0.911 0.417
ΔE*94 0.673 – 1.071 0.398
ΔE*ab 1.066 – 1.313 0.247
ΔE*uv 0.975 – 1.186 0.211
CAM16-UCS 0.469 – 0.631 0.162
Oklab 1.033 – 1.112 0.079

The unit this collection publishes in has the widest spread of the six. Its elasticity is not one number but a range running from about a half to about nine tenths, so a sentence of the form the census’s answers move by seven tenths as much as its test set’s saturation is a mean of fourteen quantities that differ by nearly a factor of two — and it is least summarisable in exactly the unit every published figure uses.

Oklab is the extreme in the other direction: elasticity 1.07 on every row, to within 0.08. Its answers are almost exactly proportional to the test set’s saturation, uniformly, which is a bad property to have and a very predictable one.

That inverts the usual reading of these two columns. A low mean elasticity says a unit is discounting the input; a low spread says it is discounting it consistently, and the two are separate virtues. CAM16-UCS is the only entry that has both. ΔE2000 has the first and the worst of the second, which is a fair description of a formula assembled from several corrections that each bite in a different part of the space.

The prediction, and what happened to it

This is the point at which a prediction written down a round earlier can be scored, and it is worth doing carefully because it was half right in a specific way.

The sentence was: ΔE*76, having no chroma weighting, raises the saturation elasticity towards one and leaves the extremes alone.

Every row rises, which is the first clause and is exactly right: all fourteen, without exception, with a mean increase of 0.418.

It goes past one rather than towards it. The mean lands at 1.105 and the smallest row at 1.066, so the whole table crosses the threshold rather than approaching it. “Towards one” implied an asymptote the arithmetic does not have.

And the extremes do not stay put — they close. The mildest row’s elasticity goes from 0.494 to 1.106, a rise of 0.61; the harshest from 0.911 to 1.313, a rise of 0.40. The spread across the census narrows from 0.417 to 0.247, so the weighting is not a common-mode discount but something that acts hardest where the elasticity was lowest. The prediction had that exactly inverted.

Recording all three is worth more than recording the one that was right. The habit is the collection’s own: a prediction deleted after the fact leaves a file that looks as though it never made one.

The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.
Fig. 5 The levels the elasticities belong to. The weighted units read the census high and the unweighted ones read it low, after calibration — the same operation acting on the levels and on their sensitivity at once.

What a division by chroma is a claim about

A divisor is a modelling statement, and it is worth saying what this one asserts before asking whether it is right.

It says that the just noticeable chroma difference grows in proportion to chroma, over the range the fit covers — the Weber behaviour that turns up wherever a sensory system reports ratios rather than differences, and that this collection meets again in the lightness axis and again in the response of a cone. Where the pattern is Weber’s, the natural coordinate is a logarithm, and a logarithm is what the appearance model uses.

Where it is asserted that the pattern is Weber’s, three things follow that a reader should hold together. A difference has no absolute size — only a size relative to where it was taken — which is a sentence this collection has already argued from the ellipse side. A tolerance is therefore not a number but a number plus a place, and a specification quoting one without the other is under-determined. And an average over a set of differences depends on where the set sits, which is exactly the elasticity above and is why the two questions are the same question.

The divisor also has a range of validity that the formulae do not state. 1 + 0.045 C is fitted over chromas up to about 60; extrapolated to a chroma of 130 — reachable by a saturated ink and by nothing on a screen — it divides by nearly seven, and there is no data behind that. The same problem attends every extrapolated constant in this collection, and it is worth flagging because the most saturated members of the census’s test set are exactly where the divisor is doing the most work.

Why the rotation term is not the story

CIEDE2000’s most-discussed feature is a rotation of the chroma–hue ellipse in the blue region, and it deserves a paragraph mostly to be set aside.

It exists because the ellipses in the blues are tilted rather than aligned with the chroma and hue axes, and no amount of scaling the two axes independently can fix a tilt. Its magnitude is governed by a Gaussian centred at a hue angle of 275° with a width of 25°, so it is exactly zero over most of the hue circle and reaches its maximum in one place.

On the census, that place is barely visited. The test surfaces are single-lobed reflectances at levels and depths spanning the whole hue circle, and the rows most affected are the two involving daylight going bluer. Removing the rotation term entirely and leaving everything else moves the census’s mean by well under the width of the calibration’s own residual scatter.

So the practical summary of what CIEDE2000 does to this collection is: it divides by chroma, and then it does a number of other things that are individually visible and collectively small. The literature’s emphasis is on the other things because they are what is new; the effect on a body of results is almost entirely the division.

Where the model stops

None of this says the weighting is right. It says the weighting is what separates the formulae, and separately that a formula’s agreement with ΔE2000 is a poor guide to its merit — the two of the five that agree least about the census are the two that do best on the ellipses the whole question was originally settled by.

It also does not establish that the constant is right. 1 + 0.045 C is a fitted line through a scatter of measurements, and nothing here varies the 0.045 — which is the natural next question and is a dial rather than a menu, because the two ends of it are two of the formulae on the menu.

And the reference sample the calibration is fitted on is constructed, like everything else here. A different sample gives different scale factors and therefore different scatters. The split is robust to that — the two weighted formulae remain the two closest on any sample tried — but the numbers are not, and the honest form of the finding is the ordering rather than the percentages.

MacAdam's twenty-five ellipses, measured in each unit. The uniformity instrument used here, applied to units rather than to spaces. The upper bar is anisotropy — the mean over the twenty-five of the largest radius divided by the smallest, where 1 would be a circle. The lower is spread — the largest mean radius divided by the smallest across all twenty-five, which asks whether a step of the same size means the same thing in different parts of the diagram. Reading down the three CIELAB-based formulae in the order they were published, the anisotropy falls 3.42 → 2.89 → 2.74 and the spread rises 3.24 → 3.59 → 4.12: the weighting divides a difference by the chroma it was measured at, which equalises directions at a point and unequalises magnitudes between points. Neither number is scaled, so no calibration is applied here. CAM16-UCS is ahead on both.
Fig. 6 The same six units on the collection’s own uniformity instrument, which is where the weighting’s cost appears: it makes each of MacAdam’s ellipses rounder and makes them more unequal in size.

Who found it, and when

The observation that discrimination ellipses grow with chroma is MacAdam’s, 1942, and is the oldest quantitative fact in this part of the subject. That CIELAB does not account for it was known when CIELAB was recommended; the committee said so.

What took until 1994 to arrive was the form of the correction, and what took until 2000 was the admission that one linear divisor was not enough. Ottosson’s Oklab in 2020 goes the other way — a space designed so that a plain Euclidean distance works — and it lands on the far side of this menu from the weighted formulae, which is a result about how hard that is rather than about the space.

Two sensitivities from two libraries, under every unit. Two quantities that share no code, no test set and no physical question: how much the adaptation census's residual depends on how saturated its surfaces are, and how much a camera profile's reported error depends on how saturated its test chart is. The first is a mean over fourteen changes of light built from cosine combinations; the second is one number about one silicon sensor scored on Gaussian bumps. Under the published unit they sit at 0.687 and 0.656. Across the whole menu they move together, from about 0.5 under the appearance unit to about 1.15 under plain CIELAB, staying within 12 per cent of each other at the worst point. Two numbers agreeing once is a coincidence; two curves agreeing at six points across a factor of two and a half is a shared mechanism, and the mechanism is the compression the unit applies to a chroma difference.
Fig. 7 The weighting acting on two unrelated measurements at once: a camera profile’s dependence on its chart and the census’s dependence on its test set, both falling together as the chroma weighting strengthens.

Where the ladder goes next

The menu is discrete except in one place. ΔE*94’s two weighting constants can be scaled continuously, and at zero they make every weight exactly one, so the formula becomes ΔE*ab — not approximately, but as an identity. That line joins the two ends of the oldest disagreement in the subject and can be walked, which turns a spread over six points into a derivative.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 10 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Appearance modelCalibrationChromaCIEDE2000Colour differenceElasticityMacAdam's ellipsesRank correlationTest setUniformity