Difference and uniformity

A dial through a discrete menu

ΔE*94 is ΔE*ab with two weighting constants in it, and at zero those constants make every weight exactly one — so the two ends of the oldest disagreement in colour difference are joined by a line rather than separated by a choice. Walking it gives a derivative where a menu gives only a spread, and the derivative says the published weighting is on the far side of the interesting part.

Assumes The weighting is the disagreement, A choice with no magnitude and Saturation is nearly everything.

A menu gives a spread. A dial gives a derivative, and a derivative answers a question a spread cannot: whether the value somebody chose sits on a flat part of the curve or on a steep one.

The census along the line from ΔE76 to ΔE94, and past it. ΔE94 is ΔE76 with two weighting constants in it, and at zero those constants make every weight exactly one, so the two formulae are joined by a line rather than separated by a choice. The horizontal axis is how much of the published weighting is applied: 0 is exactly ΔE76, 1 is exactly ΔE94, and 3 is three times more weighting than anybody has proposed. The falling curve is the census's mean elasticity to how saturated its test set is, which drops from 1.11 to 0.74 — most of the fall happening before the published value is reached. The other curve is Kendall's τ against ΔE2000's ranking, and it peaks at w = 0.5, not at 1: the weighting that best reproduces the published ordering is about half the published weighting. There is no value of this dial that reaches ΔE2000, whose rotation term is not on this line at all.
Fig. 1 The adaptation census walked along the line from ΔE*ab to ΔE*94 and three times past it. The falling curve is the census’s mean sensitivity to how saturated its test set is; the other is Kendall’s τ against ΔE2000’s ordering. Both are computed at every stop with the same calibration applied.

The claim

Two of the six formulae on the menu are the two ends of one line, and walking the line puts the published weighting well past the steep part.

  • The identity is exact, not approximate. At zero weighting ΔE*94’s divisors become exactly 1 and the formula is bit-for-bit the Euclidean distance in CIELAB.
  • Most of the effect happens in the first quarter. The census’s saturation sensitivity falls from 1.105 to 0.986 between no weighting and a quarter of the published amount, and only from 0.986 to 0.835 over the remaining three quarters.
  • The agreement with ΔE2000’s ranking peaks at half the published weighting, not at the published value: Kendall’s τ of 0.934 at w = 0.5 against 0.912 at w = 1.
  • The levels turn over. The census mean rises to 1.153 at twice the published weighting and falls again by three times it, so more weighting is not monotonically anything.
  • And there is no value of the dial that reaches ΔE2000. Its rotation term is not on this line at all, and saying so is the point of drawing the line.

The line, and why it is exact

ΔE*94’s two additions to the Euclidean distance are divisors. The chroma difference is divided by 1 + K₁C and the hue difference by 1 + K₂C, with K₁ = 0.045 and K₂ = 0.015 in the graphic-arts recommendation and C the chroma of the reference colour.

Scale both constants by a single number w, and at w = 0 both divisors are 1 + 0 — exactly one, for every colour, with no rounding. The formula collapses term by term to the square root of the sum of squares of the lightness, chroma and hue differences, which is the Euclidean distance in CIELAB written in polar coordinates and is therefore ΔE*ab exactly.

That the identity is exact rather than a limit is what makes the dial a reconstruction of the two formulae rather than a curve passing near them, and it is asserted as such: the largest disagreement between the dial at zero and ΔE*ab over the reference pairs is below 10⁻¹², and the same at one against ΔE*94.

Eighteen years of committee work separate the two formulae, and the whole of the difference is one number. Which is not a slight on the committee — deciding that the divisor should exist, that it should be linear, and that the constant should be 0.045 rather than 0.02 or 0.09, is the work. But it does mean the choice has a magnitude after all, and a choice with a magnitude can be audited the way every other input in this collection is.

What the derivative says

The census’s sensitivity to how saturated its test surfaces are is the largest sensitivity anywhere in this collection, and it is the quantity worth watching along the dial.

weighting census mean saturation sensitivity τ against ΔE2000
0 — ΔE*ab exactly 0.958 1.105 0.824
0.25 1.038 0.986 0.912
0.5 1.081 0.916 0.934
0.75 1.109 0.869 0.912
1 — ΔE*94 as published 1.129 0.835 0.912
1.5 1.149 0.791 0.890
2 1.153 0.764 0.824
3 1.133 0.737 0.780

The sensitivity falls fastest at the start. A quarter of the published weighting buys 0.119 of the fall and the remaining three quarters buy 0.151 — so the first quarter is 44 per cent of everything the published weighting achieves. Past the published value it keeps falling and keeps flattening, so there is no point at which more weighting stops helping and no point at which it helps much.

Two denominators are available for that fraction and they say different things. Against the fall over the published range, 0 to 1, the first quarter buys 44 per cent. Against the fall over the whole swept range, 0 to 3, it buys 32. The first is the number a reader wants, because it is a statement about the choice somebody actually made; the second is a statement about a range nobody has proposed, and pairing a numerator from one with a denominator from the other understates the effect by a quarter.

That shape has a consequence for how the published choice should be read. If the argument for a divisor is that it reduces how much a published average depends on the set it was averaged over, the argument is nearly all made by w = 0.25. The rest of the way to 1 is bought for other reasons — agreement with threshold data, mostly — and those reasons are not visible from here.

The weighting moves the robustness twice as much as it moves the levels

The dial’s whole purpose is to put the choice of formula on the same axis as every other input this collection audits, and doing that properly means computing the elasticity of both curves rather than one.

Over the same ±25 per cent span the collection uses everywhere else, the census mean has an elasticity of +0.052 and the saturation sensitivity has one of −0.130. Two and a half times the size, and the opposite sign.

So the choice of weighting is a small decision about what the numbers are and a much larger one about how much they depend on the set they were computed over. A reader comparing two published census levels is barely exposed to it; a reader comparing two rows’ robustness — which is the quantity the saturation audit exists to report — is exposed two and a half times as much.

That inverts which of the two curves matters. The census mean is the headline quantity and it is nearly indifferent to the dial; the sensitivity is the diagnostic quantity and it is the one the dial moves. A formula choice that would barely register in the published table registers clearly in the audit of the published table, which is an awkward property for an audit to have and is worth stating rather than leaving for somebody to find.

It also puts the choice in its place among the collection’s other inputs. An elasticity of 0.05 on the levels is the smallest of anything measured here — below every declared population width and far below the test set’s own three numbers, the largest of which is near one. On the levels, which formula is used is the least consequential choice in the audit. On the sensitivity it is 0.13, which is still small but is no longer negligible beside the population widths.

The honest summary is therefore two sentences rather than one. The published numbers in this collection would look much the same under any weighting on this line. The statements the collection makes about how far those numbers can be trusted would not, and the weighting is one of the things deciding them.

The peak that is not where it should be

The second curve is the one that is genuinely surprising, and it is worth being careful about what it does and does not show.

Kendall’s τ between the dialled census’s ranking and ΔE2000’s ranking rises from 0.824 at no weighting, peaks at 0.934 at half the published weighting, and falls thereafter. The published ΔE*94 sits at 0.912, on the far side of the peak, and so does everything above it.

The naive reading is that ΔE2000 is between ΔE*ab and ΔE*94 in effective weighting. That is not right, or not only right. CIEDE2000’s chroma divisor is 1 + 0.045 C — the same constant — but it is applied to a modified chroma, after an a-axis correction that inflates chroma near the grey axis, and alongside a lightness weighting and a rotation term. The effective weighting is not a scalar multiple of ΔE*94’s, so a peak at 0.5 does not mean “ΔE2000 weights half as much”.

What it does mean is narrower and still useful: on this body of results, the ordering produced by ΔE2000 is best reproduced by a formula that is not one of the ones on the menu. A reader who wanted a cheap approximation to ΔE2000’s ranking of these fourteen rows would do better with a half-weighted ΔE*94 than with either published formula, and would have had no way of knowing that from six points.

How much of the census ranking each unit keeps. Kendall's τ between each unit's ordering of the fourteen census rows and the published one; 1 would be perfect agreement. The number beside each bar is what τ is computed from — how many of the 91 pairs of rows the unit puts the other way round. The best agreement is ΔE′, the one unit here that is not a distance between two triples at all, at τ 0.956 and 2 discordant pairs. The worst is ΔEok at 0.780. A ranking that survives every unit is a ranking a reader can rely on; the ones that do not are listed in the figure that follows.
Fig. 2 The menu’s own agreement with the published ranking, for comparison. The dial passes above every point on this chart except CAM16-UCS’s, at a value of the weighting nobody has published.

The levels turn over, and that is a warning

The census mean does not rise monotonically with the weighting. It goes 0.958, 1.038, 1.081, 1.109, 1.129, 1.149, 1.153, 1.133 — up to a maximum somewhere near twice the published weighting, and then down.

The mechanism is the calibration, and it is worth following because the same trap is available in any comparison of rescaled quantities. Each stop is multiplied by the scale that best carries it onto ΔE2000 over the reference pairs. As the weighting grows, the formula shrinks every chroma difference, so the fitted scale grows to compensate. The census’s surfaces are more saturated than the reference sample’s on average, so the shrinkage bites harder on the census than on the sample that set the scale, and the two effects race. Below w ≈ 2 the growing scale wins; above it the shrinking differences do.

A turning point in a calibrated quantity is a statement about the difference between two sets, not about the formula. Reporting it as though the formula had an optimum would be the mistake, and the honest form of the observation is that the census’s surfaces are more saturated than the pairs the calibration was fitted on — which is itself worth knowing, since it means every calibrated number in this audit is slightly conservative.

The objects the average is over, and the region they come from. Two panels. On the left, eight of the 125 reflectance spectra in the test set, drawn as reflectance against wavelength from 380 to 780 nanometres — smooth, broad curves between about 0.02 and 0.9, with at most two gentle undulations each, because each is a level times a combination of two cosines. None of them has a narrow feature, because the family has no basis function that could make one. On the right, the region those surfaces come from, drawn in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the constraint that the two depths may not exceed 0.7 in sum, and 25 lattice points inside it. Five levels of each of those pairs is the whole test set. The square's four corners — the most saturated surfaces the two cosines could make — are outside the diamond and are not in the set at all.
Fig. 3 The region the census’s test surfaces are drawn from, in its own coordinates. Its corners are the most saturated members, and they are the surfaces on which a chroma divisor does the most work — which is why the calibration and the census pull against each other.
What one change of light costs, surface by surface — daylight to a triphosphor tube. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a triphosphor tube is 2.323 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 4.786, which is 2.06 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 4 The harshest lamp in the census, surface by surface. The steep upper end is where the saturated surfaces are, and it is the part of this curve that a weighting flattens.

What a dial buys that a menu cannot

Three things, and the third generalises past this subject.

A magnitude for a choice that did not have one. The elasticity of the census mean to the weighting, over the same ±25 per cent span this collection uses for every other input, is 0.052 — the smallest in the audit, and the sensitivity’s is two and a half times it. Both are now comparable with the elasticities of the population widths and of the test set’s own three numbers on one scale. Before the dial there was no way to put the choice of formula on the same axis as anything else.

A place to ask whether the published value is special. It is not, on either curve here. That is not an argument against it — the value was fitted to threshold data and neither curve here is threshold data — but it removes one possible defence, which is that the constant sits at some natural optimum.

And a way to distinguish two kinds of disagreement. Where two formulae are joined by a line, the disagreement between them is a matter of degree and interpolating is meaningful. Where they are not — and ΔE2000 is not on this line, nor is CAM16-UCS, nor is Oklab — the disagreement is structural and interpolation would be a fiction. A menu that contains a line inside it is a menu whose members are not all the same kind of alternative, and finding out which are is worth doing before treating a spread across six of them as one number.

How much the census depends on its test set, under each unit. Each column is a unit and each dot is one of the fourteen census rows: the elasticity of that row's residual to how saturated the test surfaces are, over the same ±25 per cent span every sensitivity in this collection uses. An elasticity of one means a set half again as saturated gives an answer half again as large. In ΔE2000 the fourteen run 0.49 to 0.91 about a mean of 0.69 — already the largest sensitivity measured anywhere in this collection. In ΔE76, which does not weight chroma at all, every one of the fourteen rises, the mean goes past one to 1.11, and the spread narrows from 0.42 to 0.25. The bar across each column is its mean.
Fig. 5 Where the dial’s two ends sit among the menu’s six points. The line runs from the rightmost column to the second, and passes through nothing else on the chart.

What the dial does to the fourteen rows separately

An average along a dial is still an average, and this collection spends a round on what that hides, so the per-row curves are worth looking at before the summary is trusted.

The spread of sensitivities across the fourteen rows narrows as the weighting rises: from 1.066–1.313 at no weighting, to 0.796–1.128 at half, to 0.673–1.071 at the published value, to 0.505–1.034 at three times it. The bottom of the range falls twice as fast as the top does. So the weighting is not a uniform discount; it acts hardest on the rows whose sensitivity was already lowest, and those are the rows involving thermal sources and discharge lamps rather than walls.

The reason is the same one that makes a wall harder than a lamp in the first place. A wall multiplies the light, so a more saturated test surface meets a more saturated illuminant and the two compound; the resulting differences are large in chroma and a chroma divisor cuts them hard. A change of lamp shifts the whole set together and leaves less chroma for the divisor to act on.

The practical consequence is that the weighting compresses the census’s own ranking of sensitivities while leaving its ranking of levels largely alone. A reader comparing two rows’ robustness rather than two rows’ residuals is comparing something the unit has substantially rearranged, and that is a second thing the choice of formula quietly decides.

Where the line does not reach

Three of the five alternatives are off it, and each is off it in a different way.

CIELUV is a different space, not a different weighting: its chromaticity coordinates are a projective transform of the diagram rather than a cube-root difference of tristimulus values, so no divisor applied to CIELAB differences produces it. The 1976 committee recommended both because it could not choose, and the choice remains unmade.

Oklab is a third space, fitted to the same threshold data as the corrections were, on the argument that a space in which a plain distance works is better than a space with a correction bolted on. Its position at the far end of this audit’s agreement scale is a result about how hard that is.

CAM16-UCS is not a space at all in the same sense. It is the output of a model with a room in it, and there is no parameter of ΔE*94 that turns it into one. Its logarithmic compression is shaped like the divisor and is not reachable by scaling one.

So the dial covers one of the five and the other four are genuinely discrete. That is the honest scope, and it is why the essays either side of this one report a spread rather than a derivative.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 6 The spread the dial cannot replace: six published quantities across the whole menu, where four of the six alternatives are not on any line through the published unit.

Where the model stops

The sweep is computed on one body of results — fourteen changes of light, one test set, one basis — so the derivative is a derivative of this census and not of colour difference in general. A different body of results would give a different curve, and the one property that is likely to transfer is the shape rather than the numbers: most of the effect in the first quarter, a flat tail, no optimum.

It also holds K₁ and K₂ in a fixed ratio. The recommendation sets them at 0.045 and 0.015 — chroma weighted three times as heavily as hue — and nothing here varies the ratio. Doing so would turn a dial into a plane, and there is a real question in it, because the two divisors act on quantities the census’s surfaces populate very unevenly.

And it says nothing about whether a linear divisor is the right shape. 1 + K C is a line through a scatter of measurements; a Weber law would give a divisor proportional to C alone, and the 1 + is what keeps the formula finite at the grey axis. Whether the fitted line survives extrapolation past the data it was fitted on is a question this collection has asked about other constants and not about this one.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 7 The menu’s five alternatives by how far each is from a rescaling of the published unit. The dial’s two ends are the first and third rows; nothing between them is a member of the menu, and the best point on the dial is not a member either.

Who found it, and when

That ΔE*94 reduces to ΔE*ab at zero weighting is arithmetic and is stated in the recommendation itself, usually in a sentence about the parametric factors. What is not usually done is to treat the reduction as a road rather than a boundary condition.

The reason it is not usually done is that nobody normally has a body of results in the formula to sweep. A formula is validated against threshold data, which is a different exercise: it asks how well the formula predicts what observers report, and the answer is a correlation rather than a table of published numbers moving. Sweeping a constant through a collection’s own results is only possible for a collection that computes its results rather than quoting them, which is the habit this whole site is arranged around.

Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu.
Fig. 8 Where on the scale the fixed members of the menu disagree. The dial’s two ends are the outermost and the innermost of the CIELAB family here, and everything between them on this chart is reachable by a value of the weighting.

Where the ladder goes next

The census’s ranking changes under every unit on the menu, by between two and ten of ninety-one pairs. A separate instrument, built for a different reason a round ago, already reported which of the ranking’s steps the test set itself could not establish. The two lists have no arithmetic in common, and comparing them turns out to say something neither could say alone.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationChromaCIEDE2000Colour differenceElasticityInterpolationRank correlationResidualSensitivityTest set