Which of two is worse
Assumes A difference is not a distance and A tolerance is a shape.
Every argument this site has made about colour-difference formulae has been about their values: that ΔE76 overstates a difference in the blue, that ΔE2000’s corrections are large, that the formulae disagree by units. All of that is true and none of it is the question a specification asks.
A specification asks whether a batch passes. A quality manager asks whether this week is worse than last week, against a tolerance whose shape is already a choice. A codec comparison asks whether the new version is better, on an image rather than on two patches. Every one of those is a question about order, and a disagreement about order is a disagreement no amount of rescaling can fix.
The claim
Two colour-difference formulae disagree about which of two pairs is worse far more often than their numerical disagreement suggests, and they disagree most exactly where a tolerance is written.
Three pairs of formulae, sampled over CIELAB and again among pairs near ΔE 1:
| over the whole space | near a tolerance of 1 | |
|---|---|---|
| ΔE2000 against ΔE76 | 13.4% | 43.2% |
| ΔE2000 against ΔE94 | 7.7% | 36.7% |
| ΔE94 against ΔE76 | 13.4% | 35.0% |
A coin flip is fifty per cent. At the size a tolerance is written at, ΔE2000 and ΔE76 carry almost no information about each other’s ordering.
And it is not an artefact of the pairs being too alike to order. Restricting the count to comparisons where the first formula separates the two pairs by more than twenty per cent — a clear ordering, not a tie — the rate at the tolerance is still 34.4 per cent against 7.4 per cent over the whole space. The disagreement survives the obvious objection, and it survives it much better at the tolerance than elsewhere, which is a section of its own below.
Why an inversion is worse than a discrepancy
The distinction is worth being exact about, because it is the difference between a nuisance and a defect.
A numerical discrepancy is absorbed by the specification. If ΔE76 reads consistently 1.6 times what ΔE2000 reads, then a tolerance of 1.6 in the first is a tolerance of 1 in the second, and a document that names its formula loses nothing at all. Every quoted number moves and every decision is unchanged.
An inversion cannot be absorbed by anything. If formula A says pair X is worse than pair Y and formula B says the opposite, no monotone transformation of either — no scaling, no offset, no curve — makes them agree. One of the two orderings is going to be acted on and the other is going to be wrong, and which one is a property of the document rather than of the colours.
That is why the rate above is the interesting measurement and the ratio between the formulae’s values is not — and it is a different failure from the metric axioms breaking, which is about a single formula rather than about two. A specification inherits the ordering of the formula it names, all the way down.
The worst case, and what it looks like
The largest inversion found in the search at a tolerance of one is a pair of pairs where ΔE2000 reads 0.75 and 1.25 — a factor of 1.66 apart, unambiguously ordered — and ΔE76 reads 1.248 and 1.238, which is the opposite order and effectively a tie.
Read as an acceptance decision at a tolerance of 1: the first pair passes on both formulae, the second passes on neither, and the two formulae disagree about which of the two batches the supplier should be told about. Read as a process-improvement decision: a change that ΔE2000 calls a two-thirds improvement, ΔE76 calls no change at all.
Neither reading is exotic. Both are the ordinary use of the number.
The tolerance inversions are the real ones
The clear-ordering figures are quoted above as an answer to an objection and they say something stronger than that, which comes out of comparing how much the restriction removes at each scale rather than reading the two numbers separately.
Over the whole space the restriction takes ΔE2000 against ΔE76 from 13.4 per cent to 7.4 — so fifty-five per cent of the inversions survive it, and the other forty-five were near-ties, cases where the first formula barely separated the two pairs and the second happened to put them the other way round.
At the tolerance it takes the same comparison from 43.2 per cent to 34.4 — eighty per cent survive.
So the near-tie explanation covers nearly half of what happens over the whole space and a fifth of what happens at the tolerance. The inversions near the line are not mostly ties being resolved differently; they are clear orderings being reversed, four times out of five.
That is worth stating as a number a practitioner can hold: when ΔE2000 says one batch is at least a fifth worse than another, and both are near a tolerance of one, ΔE76 disagrees about which is worse in more than a third of cases. Not “cannot tell them apart” — disagrees, on a comparison the first formula calls unambiguous.
The amplification going the other way is worth one line too, because it picks out an unexpected pair. Dividing the tolerance rate by the whole-space rate gives 3.2 for ΔE2000 against ΔE76, 2.6 for ΔE94 against ΔE76, and 4.8 for ΔE2000 against ΔE94 — the largest of the three. The two formulae that agree best over the space as a whole, at 92.3 per cent, are the pair whose agreement degrades fastest as the difference shrinks, down to 63.3 per cent.
Which is the opposite of a reassuring result about a modern formula replacing its predecessor. ΔE94 and ΔE2000 share a construction: the same space, the same resolution into lightness, chroma and hue, the same division by chroma. Over ordinary differences that shared core dominates and they agree. Near a threshold the shared core contributes almost nothing — the raw distance is small — and what is left is the part where the 2000 formula’s extra corrections act, which is exactly the part they do not share. A family resemblance between two formulae is a resemblance at large differences and evaporates at small ones, and a specification written at one unit is written where it has evaporated.
There is a practical reading of that which is more useful than the general complaint. A shop moving from ΔE94 to ΔE2000 — a common upgrade, made because the newer formula is better fitted — will find its historical trend data mostly intact at the differences it was collected over and its acceptance decisions changing on more than a third of borderline cases. The two consequences look like one change and are not, and only the first is what anybody was told to expect.
Why the tolerance is the worst place
The rate rising near the tolerance rather than falling is the part that needs an explanation, because the naive expectation goes the other way — small differences are where the formulae were fitted hardest.
Two things are happening.
The corrections are largest at small differences. ΔE2000’s weighting functions divide by terms that depend on the chroma and the hue, and those divisions are what make it disagree with ΔE76. Near a threshold the whole of the difference is the correction, because the raw distance is small and the correction is a multiplicative factor on it. At a difference of twenty units the two formulae are both dominated by the same Euclidean core and agree about ordering much more often.
And the pairs at a tolerance are near-ties by construction. Two batches both near a tolerance of one differ by a fraction of a unit, and the fraction is comparable with the size of the corrections. That is the objection the clear-ordering measurement answers: even restricting to comparisons where one formula separates the pairs by a fifth, the rate is a third.
The consequence for practice is uncomfortable and simple: the formula named in a specification matters most in exactly the cases the specification exists to decide.
What a document would have to say
The gap here is the same shape as the one a colour difference’s missing arrangement left, and it is smaller and easier to close.
A specification that names a tolerance already names a formula, in every well-written case. What it does not name, and what the measurement above says it needs to, is what happens to the comparisons the number is used for:
- Whether the decision is a threshold or a ranking. A threshold is a single number against a single number and inherits only the formula’s value. A ranking — this batch against that one, this week against last — inherits the formula’s order, and the two are not the same requirement.
- Which formula the ranking is in. A document that specifies acceptance in ΔE2000 and reports trends from a meter set to ΔE76 has two orderings in it, disagreeing about a seventh of everything and about a third of what sits near the line.
- And whether a re-measurement counts. Two measurements of one sample differ by the instrument’s own repeatability, which is not small, and two near-tied batches re-measured can swap order without any formula changing at all.
None of these is expensive to state. The first is one sentence, the second is one word, and the third is a number the instrument’s manufacturer already publishes.
What was computed, and how
Pairs are constructed rather than sampled from a list. A point in CIELAB, a random direction, and a step of a stated length; the pair is kept if it stays inside the lightness range. When a band is named, the step is drawn from a spread around it and the pair is kept only if the first formula’s value lands within a quarter of the band — so the pairs are near the tolerance in the sense the tolerance means.
The count is over pairs of pairs, up to forty comparisons per pair to keep the total quadratic term bounded, giving between twenty thousand and ninety-five thousand comparisons per measurement.
And the control is a formula against itself, which returns exactly zero inversions to the last bit. That is what makes the numbers measurements rather than noise in the sampler.
The clear-ordering variant is the answer to the obvious objection and is computed in the same pass: a comparison counts as clear when the first formula separates the two pairs by more than a fifth in relative terms.
The control that makes these numbers measurements
Three things could produce a large inversion rate without anything being wrong with the formulae, and each is ruled out in the computation rather than argued away.
A formula against itself returns exactly zero inversions, to the last bit, over the same sampler and the same comparison count. That is the check that the search is not manufacturing disagreements out of ties or floating-point noise.
The clear-ordering restriction rules out near-ties. Counting only comparisons where the first formula separates the two pairs by more than a fifth leaves 20,180 of the 95,000 comparisons at the tolerance, and a third of those are still inverted.
And the rate falls with the size of the difference, which is the direction a real effect should go. Over the whole space — differences from one to twenty units — the rate is 13.4 per cent; at a band of four units it is lower again. A sampling artefact would not care.
What is left is a property of the formulae.
Where the model stops
A rank inversion is not by itself an error. Two formulae disagreeing means at most one of them is right, and this essay does not say which. What it establishes is that the choice between them is a decision with a measurable consequence, not a matter of house style.
The sampling is uniform in CIELAB, which is not uniform in anything a manufacturer produces. A real acceptance decision is made on pairs concentrated near a small number of standard colours, and the inversion rate there is a property of those colours rather than of the whole space.
There is no observer variation. Every pair here is evaluated for one standard observer; real observers vary, and two observers disagreeing about which of two pairs is worse is a separate and probably larger effect.
And there is no arrangement. Every pair is a pair of numbers, which is a statement about two large patches side by side and not about anything anybody is looking at.
Sampling the pairs in a wider band of difference is where a real comparison lives, since two candidates are rarely both near a threshold.
The generalisation
The sentence worth carrying: a formula’s value is a convention and its ordering is a claim.
Everything a specification does with a colour difference is ordinal — pass or fail, better or worse, improved or not. The magnitude is a scale that any document can renormalise. So the property to ask of a formula is not whether its numbers are right but whether its order is, and nothing in the way colour-difference formulae are presented, compared or standardised is stated in those terms.
The surprising connection is with how many colours there are. That essay found the count of distinguishable colours differing by a factor of several depending on which formula does the counting, and treated it as a fact about volumes. It is the same fact as this one: a count is an integral over a metric, and two metrics that order pairs differently in thirteen per cent of cases are integrating different things. The volume result and the ordering result are one measurement seen twice.
Who found it, and when
Rank correlation as the right way to compare colour-difference formulae is not new: the datasets the formulae are fitted to are sets of visual judgements, and the standard measure of fit — the STRESS statistic and its predecessors — is a measure of agreement with a set of magnitudes. That the ordering of two arbitrary pairs is a separate and harsher test appears rarely, and the rate near a tolerance appears, as far as this site has found, nowhere.
The formulae are 1976, 1994 and 2000, each a correction to the last, and each standardised while the previous one remained in specifications already written. The result is that all three are in use simultaneously, in documents that name one of them, and about parts that are supplied against another.
And the practice that makes it matter is older than any of them: a supplier and a customer agreeing a numerical tolerance rather than a physical master panel. Where a master panel is kept — automotive colour work is the standing example, and it is the practice the arrangement essay found had solved a different unstated problem the same way — the arrangement, the observer and the ordering are all fixed by example, and none of this arises.
At four units the two formulae are being asked about differences anybody can see, and the scatter is the honest picture of their disagreement.
What the pictures cannot show
They cannot show a disagreement. Every scatter on this page is a cloud of points, and the claim is about pairs of points on opposite sides of a diagonal — a relation between two dots rather than a property of one. A reader can see the cloud’s width and has to take the rate on the arithmetic.
And the worst case is four colours. Showing it as four patches would show two pairs that look similarly different, which is the whole difficulty: the numbers disagree by a factor of 1.66 in one formula and not at all in the other, and no arrangement of four swatches on a page carries that.
Where the ladder goes next
The nearest unfinished piece is CAM16-UCS, which this site has and which was deliberately left out of the comparison. It is an appearance difference rather than a matching difference and putting it in the same table would invite treating the four as interchangeable — but the inversion rate between an appearance difference and a matching one is exactly the number that would say how much of the disagreement is about uniformity and how much is about the viewing condition.
The second is the observer. Two standard observers give two sets of Lab coordinates for one pair of spectra, and the inversion rate between them is a lower bound on how much of a specification’s ordering is an artefact of the kernel. Every ingredient is on the site and nothing has run it.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- Where the formula is not smooth ciede2000 · cielab · δe · just-noticeable difference · macadam's ellipses · perceptual uniformity · quality control · specification · threshold · tolerance
- A name is not a threshold ciede2000 · δe · just-noticeable difference · perceptual uniformity · specification · threshold · tolerance
- MacAdam measured it cielab · δe · just-noticeable difference · macadam's ellipses · perceptual uniformity · threshold · tolerance
- A catalogue is not a vocabulary cielab · δe · perceptual uniformity · quality control · specification · standard observer
- A colour has a name ciede2000 · cielab · δe · perceptual uniformity · specification · standard observer
- A mean is not a difference ciede2000 · colour management · δe · quality control · specification · tolerance
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CIEDE2000CIELABColour managementΔEJust-noticeable differenceMacAdam's ellipsesPerceptual uniformityQuality controlSpecificationStandard observerThresholdTolerance