A dial through a discrete menu
Assumes The weighting is the disagreement, A choice with no magnitude and Saturation is nearly everything.
A menu gives a spread. A dial gives a derivative, and a derivative answers a question a spread cannot: whether the value somebody chose sits on a flat part of the curve or on a steep one.
The claim
Two of the six formulae on the menu are the two ends of one line, and walking the line puts the published weighting well past the steep part.
- The identity is exact, not approximate. At zero weighting ΔE*94’s divisors become exactly 1 and the formula is bit-for-bit the Euclidean distance in CIELAB.
- Most of the effect happens in the first quarter. The census’s saturation sensitivity falls from 1.105 to 0.986 between no weighting and a quarter of the published amount, and only from 0.986 to 0.835 over the remaining three quarters.
- The agreement with ΔE2000’s ranking peaks at half the published weighting, not at the published value: Kendall’s τ of 0.934 at w = 0.5 against 0.912 at w = 1.
- The levels turn over. The census mean rises to 1.153 at twice the published weighting and falls again by three times it, so more weighting is not monotonically anything.
- And there is no value of the dial that reaches ΔE2000. Its rotation term is not on this line at all, and saying so is the point of drawing the line.
The line, and why it is exact
ΔE*94’s two additions to the Euclidean distance are divisors. The chroma difference is divided by 1 + K₁C and the hue difference by 1 + K₂C, with K₁ = 0.045 and K₂ = 0.015 in the graphic-arts recommendation and C the chroma of the reference colour.
Scale both constants by a single number w, and at w = 0 both divisors are 1 + 0 — exactly one, for every colour, with no rounding. The formula collapses term by term to the square root of the sum of squares of the lightness, chroma and hue differences, which is the Euclidean distance in CIELAB written in polar coordinates and is therefore ΔE*ab exactly.
That the identity is exact rather than a limit is what makes the dial a reconstruction of the two formulae rather than a curve passing near them, and it is asserted as such: the largest disagreement between the dial at zero and ΔE*ab over the reference pairs is below 10⁻¹², and the same at one against ΔE*94.
Eighteen years of committee work separate the two formulae, and the whole of the difference is one number. Which is not a slight on the committee — deciding that the divisor should exist, that it should be linear, and that the constant should be 0.045 rather than 0.02 or 0.09, is the work. But it does mean the choice has a magnitude after all, and a choice with a magnitude can be audited the way every other input in this collection is.
What the derivative says
The census’s sensitivity to how saturated its test surfaces are is the largest sensitivity anywhere in this collection, and it is the quantity worth watching along the dial.
| weighting | census mean | saturation sensitivity | τ against ΔE2000 |
|---|---|---|---|
| 0 — ΔE*ab exactly | 0.958 | 1.105 | 0.824 |
| 0.25 | 1.038 | 0.986 | 0.912 |
| 0.5 | 1.081 | 0.916 | 0.934 |
| 0.75 | 1.109 | 0.869 | 0.912 |
| 1 — ΔE*94 as published | 1.129 | 0.835 | 0.912 |
| 1.5 | 1.149 | 0.791 | 0.890 |
| 2 | 1.153 | 0.764 | 0.824 |
| 3 | 1.133 | 0.737 | 0.780 |
The sensitivity falls fastest at the start. A quarter of the published weighting buys 0.119 of the fall and the remaining three quarters buy 0.151 — so the first quarter is 44 per cent of everything the published weighting achieves. Past the published value it keeps falling and keeps flattening, so there is no point at which more weighting stops helping and no point at which it helps much.
Two denominators are available for that fraction and they say different things. Against the fall over the published range, 0 to 1, the first quarter buys 44 per cent. Against the fall over the whole swept range, 0 to 3, it buys 32. The first is the number a reader wants, because it is a statement about the choice somebody actually made; the second is a statement about a range nobody has proposed, and pairing a numerator from one with a denominator from the other understates the effect by a quarter.
That shape has a consequence for how the published choice should be read. If the argument for a divisor is that it reduces how much a published average depends on the set it was averaged over, the argument is nearly all made by w = 0.25. The rest of the way to 1 is bought for other reasons — agreement with threshold data, mostly — and those reasons are not visible from here.
The weighting moves the robustness twice as much as it moves the levels
The dial’s whole purpose is to put the choice of formula on the same axis as every other input this collection audits, and doing that properly means computing the elasticity of both curves rather than one.
Over the same ±25 per cent span the collection uses everywhere else, the census mean has an elasticity of +0.052 and the saturation sensitivity has one of −0.130. Two and a half times the size, and the opposite sign.
So the choice of weighting is a small decision about what the numbers are and a much larger one about how much they depend on the set they were computed over. A reader comparing two published census levels is barely exposed to it; a reader comparing two rows’ robustness — which is the quantity the saturation audit exists to report — is exposed two and a half times as much.
That inverts which of the two curves matters. The census mean is the headline quantity and it is nearly indifferent to the dial; the sensitivity is the diagnostic quantity and it is the one the dial moves. A formula choice that would barely register in the published table registers clearly in the audit of the published table, which is an awkward property for an audit to have and is worth stating rather than leaving for somebody to find.
It also puts the choice in its place among the collection’s other inputs. An elasticity of 0.05 on the levels is the smallest of anything measured here — below every declared population width and far below the test set’s own three numbers, the largest of which is near one. On the levels, which formula is used is the least consequential choice in the audit. On the sensitivity it is 0.13, which is still small but is no longer negligible beside the population widths.
The honest summary is therefore two sentences rather than one. The published numbers in this collection would look much the same under any weighting on this line. The statements the collection makes about how far those numbers can be trusted would not, and the weighting is one of the things deciding them.
The peak that is not where it should be
The second curve is the one that is genuinely surprising, and it is worth being careful about what it does and does not show.
Kendall’s τ between the dialled census’s ranking and ΔE2000’s ranking rises from 0.824 at no weighting, peaks at 0.934 at half the published weighting, and falls thereafter. The published ΔE*94 sits at 0.912, on the far side of the peak, and so does everything above it.
The naive reading is that ΔE2000 is between ΔE*ab and ΔE*94 in effective weighting. That is not right, or not only right. CIEDE2000’s chroma divisor is 1 + 0.045 C — the same constant — but it is applied to a modified chroma, after an a-axis correction that inflates chroma near the grey axis, and alongside a lightness weighting and a rotation term. The effective weighting is not a scalar multiple of ΔE*94’s, so a peak at 0.5 does not mean “ΔE2000 weights half as much”.
What it does mean is narrower and still useful: on this body of results, the ordering produced by ΔE2000 is best reproduced by a formula that is not one of the ones on the menu. A reader who wanted a cheap approximation to ΔE2000’s ranking of these fourteen rows would do better with a half-weighted ΔE*94 than with either published formula, and would have had no way of knowing that from six points.
The levels turn over, and that is a warning
The census mean does not rise monotonically with the weighting. It goes 0.958, 1.038, 1.081, 1.109, 1.129, 1.149, 1.153, 1.133 — up to a maximum somewhere near twice the published weighting, and then down.
The mechanism is the calibration, and it is worth following because the same trap is available in any comparison of rescaled quantities. Each stop is multiplied by the scale that best carries it onto ΔE2000 over the reference pairs. As the weighting grows, the formula shrinks every chroma difference, so the fitted scale grows to compensate. The census’s surfaces are more saturated than the reference sample’s on average, so the shrinkage bites harder on the census than on the sample that set the scale, and the two effects race. Below w ≈ 2 the growing scale wins; above it the shrinking differences do.
A turning point in a calibrated quantity is a statement about the difference between two sets, not about the formula. Reporting it as though the formula had an optimum would be the mistake, and the honest form of the observation is that the census’s surfaces are more saturated than the pairs the calibration was fitted on — which is itself worth knowing, since it means every calibrated number in this audit is slightly conservative.
What a dial buys that a menu cannot
Three things, and the third generalises past this subject.
A magnitude for a choice that did not have one. The elasticity of the census mean to the weighting, over the same ±25 per cent span this collection uses for every other input, is 0.052 — the smallest in the audit, and the sensitivity’s is two and a half times it. Both are now comparable with the elasticities of the population widths and of the test set’s own three numbers on one scale. Before the dial there was no way to put the choice of formula on the same axis as anything else.
A place to ask whether the published value is special. It is not, on either curve here. That is not an argument against it — the value was fitted to threshold data and neither curve here is threshold data — but it removes one possible defence, which is that the constant sits at some natural optimum.
And a way to distinguish two kinds of disagreement. Where two formulae are joined by a line, the disagreement between them is a matter of degree and interpolating is meaningful. Where they are not — and ΔE2000 is not on this line, nor is CAM16-UCS, nor is Oklab — the disagreement is structural and interpolation would be a fiction. A menu that contains a line inside it is a menu whose members are not all the same kind of alternative, and finding out which are is worth doing before treating a spread across six of them as one number.
What the dial does to the fourteen rows separately
An average along a dial is still an average, and this collection spends a round on what that hides, so the per-row curves are worth looking at before the summary is trusted.
The spread of sensitivities across the fourteen rows narrows as the weighting rises: from 1.066–1.313 at no weighting, to 0.796–1.128 at half, to 0.673–1.071 at the published value, to 0.505–1.034 at three times it. The bottom of the range falls twice as fast as the top does. So the weighting is not a uniform discount; it acts hardest on the rows whose sensitivity was already lowest, and those are the rows involving thermal sources and discharge lamps rather than walls.
The reason is the same one that makes a wall harder than a lamp in the first place. A wall multiplies the light, so a more saturated test surface meets a more saturated illuminant and the two compound; the resulting differences are large in chroma and a chroma divisor cuts them hard. A change of lamp shifts the whole set together and leaves less chroma for the divisor to act on.
The practical consequence is that the weighting compresses the census’s own ranking of sensitivities while leaving its ranking of levels largely alone. A reader comparing two rows’ robustness rather than two rows’ residuals is comparing something the unit has substantially rearranged, and that is a second thing the choice of formula quietly decides.
Where the line does not reach
Three of the five alternatives are off it, and each is off it in a different way.
CIELUV is a different space, not a different weighting: its chromaticity coordinates are a projective transform of the diagram rather than a cube-root difference of tristimulus values, so no divisor applied to CIELAB differences produces it. The 1976 committee recommended both because it could not choose, and the choice remains unmade.
Oklab is a third space, fitted to the same threshold data as the corrections were, on the argument that a space in which a plain distance works is better than a space with a correction bolted on. Its position at the far end of this audit’s agreement scale is a result about how hard that is.
CAM16-UCS is not a space at all in the same sense. It is the output of a model with a room in it, and there is no parameter of ΔE*94 that turns it into one. Its logarithmic compression is shaped like the divisor and is not reachable by scaling one.
So the dial covers one of the five and the other four are genuinely discrete. That is the honest scope, and it is why the essays either side of this one report a spread rather than a derivative.
Where the model stops
The sweep is computed on one body of results — fourteen changes of light, one test set, one basis — so the derivative is a derivative of this census and not of colour difference in general. A different body of results would give a different curve, and the one property that is likely to transfer is the shape rather than the numbers: most of the effect in the first quarter, a flat tail, no optimum.
It also holds K₁ and K₂ in a fixed ratio. The recommendation sets them at 0.045 and 0.015 — chroma weighted three times as heavily as hue — and nothing here varies the ratio. Doing so would turn a dial into a plane, and there is a real question in it, because the two divisors act on quantities the census’s surfaces populate very unevenly.
And it says nothing about whether a linear divisor is the right shape. 1 + K C is a line through a scatter of measurements; a Weber law would give a divisor proportional to C alone, and the 1 + is what keeps the formula finite at the grey axis. Whether the fitted line survives extrapolation past the data it was fitted on is a question this collection has asked about other constants and not about this one.
Who found it, and when
That ΔE*94 reduces to ΔE*ab at zero weighting is arithmetic and is stated in the recommendation itself, usually in a sentence about the parametric factors. What is not usually done is to treat the reduction as a road rather than a boundary condition.
The reason it is not usually done is that nobody normally has a body of results in the formula to sweep. A formula is validated against threshold data, which is a different exercise: it asks how well the formula predicts what observers report, and the answer is a correlation rather than a table of published numbers moving. Sweeping a constant through a collection’s own results is only possible for a collection that computes its results rather than quoting them, which is the habit this whole site is arranged around.
Where the ladder goes next
The census’s ranking changes under every unit on the menu, by between two and ten of ninety-one pairs. A separate instrument, built for a different reason a round ago, already reported which of the ranking’s steps the test set itself could not establish. The two lists have no arithmetic in common, and comparing them turns out to say something neither could say alone.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The disagreement is at the near end calibration · chroma · ciede2000 · colour difference · residual · test set
- A chart decides what a camera scores chroma · colour difference · elasticity · sensitivity · test set
- The census in six units calibration · ciede2000 · colour difference · residual · test set
- Two instruments and one ranking calibration · colour difference · rank correlation · residual · test set
- A departure is straight in the excitations chroma · ciede2000 · colour difference · residual
- A partial correction is worth its fraction colour difference · interpolation · residual · test set
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CalibrationChromaCIEDE2000Colour differenceElasticityInterpolationRank correlationResidualSensitivityTest set