Difference and uniformity

How far apart are two colours

ΔE is meant to be a distance with the property that the same number means the same perceived difference everywhere. Three successive formulae have tried, they disagree with each other by more than a just-noticeable difference, and the disagreement decides real matching questions.

The useful thing to be able to say about two colours is how different they are. Industrial colour control needs it — is this batch within tolerance? — and so does anything that has to decide whether two things will be confused.

The quantity is called ΔE, and there are at least three of them.

Three colour-difference formulae, disagreeing. ΔE76, ΔE94 and ΔE2000 for the same 9 pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.
Fig. 1 ΔE76, ΔE94 and ΔE2000 evaluated on the same nine pairs. The largest disagreement between the oldest and newest here is more than four units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether a pair counts as matching.

How many pairs are drawn changes nothing about the disagreement, and how large they are changes a great deal.

Three colour-difference formulae, disagreeing. ΔE76, ΔE94 and ΔE2000 for the same 24 pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.
Fig. 2 The same three formulae on twenty-four pairs rather than nine. The spread between them is a property of where in the space a pair sits, so a larger sample widens the range of answers rather than settling it.

Drawn large enough to be compared by eye, the same disagreement is easier to see and harder to argue with than any table of it.

Three colour-difference formulae, disagreeing. ΔE76, ΔE94 and ΔE2000 for the same 7 pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.
Fig. 3 And seven pairs, drawn large enough to be compared by eye. Every pair here is one number to one formula and three different numbers to three, which is the whole of what a unit of colour difference is.
Three colour-difference formulae, disagreeing. ΔE76, ΔE94 and ΔE2000 for the same 16 pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.
Fig. 4 Sixteen pairs, which is enough to see that the disagreement is not concentrated in any one region of the space. Every one of these is one number to a specification and three numbers to the three formulae a specification might have named.

Two more samplings say that the disagreement between the formulae is not a property of how many pairs are drawn.

Three colour-difference formulae, disagreeing. ΔE76, ΔE94 and ΔE2000 for the same 5 pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.
Fig. 5 Five pairs, large enough that the disagreement can be judged by eye rather than read off an axis. Every one of these is one number to a specification and three to the three formulae.
Three colour-difference formulae, disagreeing. ΔE76, ΔE94 and ΔE2000 for the same 9 pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.
Fig. 6 And the generator’s own default sampling, which is the version this collection quotes in its other essays. Four samplings, one conclusion: how far apart two colours are is a question with three published answers.

The 1976 formula, and why it was not enough

CIELAB was designed so that Euclidean distance in it would be perceptually meaningful. If that had worked, ΔE would be

ΔE76=(ΔL)2+(Δa)2+(Δb)2\Delta E_{76} = \sqrt{(\Delta L^*)^2 + (\Delta a^*)^2 + (\Delta b^*)^2}

and there would be nothing more to say. The formula is simple, it is symmetric, it obeys the triangle inequality, and it is a genuine metric in the mathematical sense.

It does not work well enough. Transforming MacAdam’s ellipses into CIELAB shows why: the space equalised the size of the discrimination contours far better than it made them round, leaving a mean anisotropy of about 3.4. A step of a given size still means substantially different things depending on direction and position — particularly in the saturated regions, where CIELAB systematically overstates differences.

ΔE94: weighting by chroma

The 1994 formula keeps the space and weights the components:

ΔE94=(ΔLkLSL)2+(ΔCkCSC)2+(ΔHkHSH)2\Delta E_{94} = \sqrt{\left(\frac{\Delta L}{k_L S_L}\right)^2 + \left(\frac{\Delta C}{k_C S_C}\right)^2 + \left(\frac{\Delta H}{k_H S_H}\right)^2}

with SCS_C and SHS_H growing linearly with chroma. The effect is that differences between saturated colours are discounted relative to differences between muted ones, which is the direction the data demands.

The kk terms are application-dependent parameters — textiles use different values from graphic arts — which is the first sign of something significant. A “perceptual distance” that needs to be told what industry it is being used in is not purely perceptual.

ΔE2000: and this is where it gets ugly

CIEDE2000 adds three more corrections on top:

  • A chroma-dependent rotation term RTR_T, which handles the blue region specifically. CIELAB’s worst-known failure is that blues shift toward purple as chroma changes, so the discrimination ellipses there are rotated relative to the axes, and a term is needed to account for it.
  • A lightness weighting SLS_L that varies with lightness, since discrimination is poorer at the extremes than in the middle.
  • A correction to aa^* applied before anything else, which improves behaviour for near-neutral colours.

The formula runs to about twenty lines. It is not a metric — it violates the triangle inequality — and it has discontinuities in its derivatives. It is also, on the data it was fitted to, clearly the best of the three.

The trajectory is worth noticing. Each formula is a patch on the space rather than a fix to it, and the patches accumulate. By 2001 the community had a very accurate empirical fit and a space that was still not uniform.

How often the three disagree about an ordering

The formulae differ in value, and the sharper question is whether they differ in ranking — because a tolerance decision is a comparison rather than a measurement.

Drawing twenty thousand pairs at ordinary chroma and asking each formula how far apart each is, the rank correlations are 0.799 between ΔE76 and ΔE94, 0.730 between ΔE76 and ΔE2000, and 0.910 between ΔE94 and ΔE2000. The 1994 formula is the middle one, and it sits much closer to the newest than to the oldest — which is what a shared chroma weighting does.

The ratio between the oldest and the newest is not a constant either. Per pair it runs from 0.93 at the fifth percentile to 2.28 at the ninety-fifth, with a median of 1.42, so no scale factor converts one into the other.

The ordering statistic is the one to carry. Taking pairs of pairs and asking which of the two differs more, ΔE76 and ΔE2000 disagree on 22.6 per cent of comparisons — and on 28.3 per cent among pairs where both are under three units, which is the region a tolerance is written in.

So more than one comparison in four, near threshold, is settled by which formula happens to be in use. That is not the caption’s warning softened by a qualification; it is the same warning with a frequency attached, and the frequency is worse exactly where the warning matters.

What a ΔE value licenses

Three claims are commonly made, and only the first is safe.

“ΔE around 1 is roughly a just-noticeable difference.” Approximately true for ΔE2000 under controlled side-by-side viewing. This is the number’s design point.

“ΔE 3 is three times as different as ΔE 1.” Not reliable. The formulae are fitted to threshold and near-threshold data, and their behaviour at large differences is an extrapolation. Two pairs at ΔE 20 are not usefully described as differing by the same amount.

“ΔE under 2 means nobody will notice.” Depends entirely on the arrangement. Two patches sharing an edge are compared far more sensitively than two separated by white space, and two seen at different times are compared very poorly indeed. The published thresholds assume side-by-side comparison, and the viewing situation is not part of the formula.

That last one causes the most trouble in practice. A gradient with a ΔE 1 step between adjacent bands will band visibly, because adjacent comparison is the most sensitive case there is. The same step between two elements at opposite corners of a page is invisible.

Where the numbers are used

Manufacturing tolerance. Paint, plastics, textiles and print all specify acceptable ΔE ranges against a standard. Automotive paint is among the tightest, since adjacent body panels are the worst possible arrangement — large areas, shared edge, viewed together.

Display characterisation. Reviews quoting an average ΔE for a monitor are measuring how far its output sits from the target for a set of test colours.

Image compression and rendering. Deciding whether a difference is worth encoding is a ΔE question, though most codecs use cruder approximations for speed.

In all three, the number is being asked to stand in for a judgement about a specific viewing situation, using a formula fitted to a different one. This is not a reason to abandon it; it is a reason to treat thresholds as guidance rather than as physics.

Why not just build a uniform space

This is the obvious response, and it has been tried repeatedly.

The difficulty is that perceptual colour space appears not to be flat. The discrimination contours vary in a way that cannot be removed by any smooth coordinate change — the geometry is intrinsically curved, and a flat coordinate system on a curved space always leaves distortion somewhere. Making the ellipses circular in one region distorts them elsewhere.

Doing the job properly means treating colour space as a Riemannian manifold and defining distance as the geodesic length under a metric fitted to discrimination data. This has been done, it works better than any of the ΔE formulae, and almost nobody uses it, because computing a geodesic is far more work than evaluating twenty lines of arithmetic and the improvement is modest for most purposes.

So the field has settled on a complicated formula over a simple space rather than a simple formula over a complicated one, on grounds that are entirely practical.

The shape of the disagreement

The three formulae do not disagree randomly. They disagree systematically, in ways that follow from what each was trying to fix.

ΔE76 overstates differences between saturated colours, because CIELAB stretches the space out toward high chroma more than perception does. ΔE94 discounts them with a chroma-dependent weighting. ΔE2000 discounts them further and adds the blue-region rotation on top.

So the gap between the oldest and newest formula is largest exactly where the colours are most saturated, which is also where a great deal of practical matching happens — brand colours, product finishes, dyed textiles are rarely muted.

Where the formulae are not the right tool

Three situations where a ΔE is the wrong measurement, and something else is wanted.

Ordering a palette. ΔE answers “how different”, not “how different in what way”. Two colours can be ΔE 20 apart in lightness or in hue, and for a palette those are not interchangeable. Working directly in a perceptual space’s lightness, chroma and hue coordinates says more than a single distance does.

Banding in a gradient. The relevant quantity is the step between adjacent bands, and adjacent comparison is far more sensitive than the thresholds ΔE was fitted to. A gradient with steps well under 1 can still band visibly.

Accessibility contrast. This is a luminance-ratio question, not a colour-difference question. Two colours can be far apart in ΔE and have almost identical luminance — a saturated blue and a saturated red, for instance — which makes them a poor foreground-background pair however different they look.

Why the formulae keep the space

An obvious question is why the CIE did not simply define a better space in 2000 rather than a more complicated distance on the old one.

Part of the answer is inertia — CIELAB was embedded in standards, instruments and software by then. The more interesting part is that it may not be possible. The discrimination data suggests perceptual colour space is intrinsically curved, and a flat coordinate system on a curved space always leaves distortion somewhere. Making the contours circular in one region distorts them in another.

Treating the problem properly means defining distance as a geodesic under a metric fitted to discrimination data, which has been done and works better than any ΔE formula. Almost nobody uses it, because a geodesic is far more expensive than twenty lines of arithmetic and the gain is modest for industrial tolerance work.

So the field has a complicated formula over a simple space rather than a simple formula over a complicated one, entirely for practical reasons. The measurement showing why is eighty years old.

What was computed here

The CIEDE2000 implementation on this site is validated against the Sharma–Wu–Dalal test set, and this is one of the few places where a published test set exists precisely because the formula is so easy to get subtly wrong.

Fifteen pairs are checked, chosen by their authors to exercise the branches implementations fail on: hue angles straddling 360°, zero chroma where the hue is undefined, and the rotation term in the blues. The maximum deviation from the published values is 4.2×1054.2\times10^{-5}.

The gate also checks that passing the test set means something. A plain Euclidean distance is evaluated on the first test pair and must be far from the ΔE2000 answer — it gives 4.01 against 2.04, so the set discriminates by a factor of two rather than certifying anything that happens to be close.

The uniformity figures come from the ellipse measurement, computed at build time rather than quoted.

The uniformity numbers quoted are computed at build time from the ellipse parameters rather than transcribed, so the table and the figures cannot drift apart. The ranking assertion requires raw chromaticity to come out worse than CIELAB, which is the result MacAdam’s experiment established — a measurement failing to reproduce it would be measuring nothing.

Choosing a formula in practice

The decision is usually simpler than the formulae suggest.

Use ΔE2000 for anything where the number will be compared against a published threshold or a tolerance specification. It is the current standard, the thresholds in circulation are calibrated to it, and its complexity is somebody else’s problem once the implementation is verified.

Use ΔE76 when the metric properties matter — clustering, nearest-neighbour search, anything requiring the triangle inequality. ΔE2000 is not a metric and will produce incoherent results in an algorithm that assumes one.

Use a perceptual space’s coordinates directly when the question is about how two colours differ rather than how much. Lightness, chroma and hue separately say more than any scalar.

The implementation trap

CIEDE2000 has a reputation for being difficult to implement correctly, and the reputation is earned.

The failure modes are all in the hue arithmetic. Hue is an angle, angles wrap at 360°, and the formula requires both a hue difference and a hue mean, each with its own wrapping rule. When either chroma is zero the hue is undefined and must be special-cased rather than allowed to become a NaN that propagates silently.

Sharma, Wu and Dalal published their test set in 2005 after finding that independent implementations disagreed with one another — not by rounding, but by amounts large enough to change a pass into a fail. The set exists because the formula is hard, and any implementation not checked against it should be assumed wrong.

Why the formulae are as complicated as they are

One sentence, and it is the summary of this essay and the one before it: the space is not uniform, nobody has managed to make one that is, and every term beyond Euclidean distance is compensating for the shortfall.

A final caution about thresholds. Numbers like “ΔE 1 is a just-noticeable difference” travel widely and lose their conditions on the way. The condition is side-by-side comparison, under controlled lighting, against a neutral background, by an observer attending to the comparison. Loosen any of those and the threshold moves, usually upward and sometimes by a large factor — which is why a specification quoting a ΔE tolerance should also quote the viewing arrangement it applies to.

It is worth adding that the formulae describe thresholds, and most design decisions are made well above threshold. Choosing whether two categories in a chart are distinguishable is a suprathreshold question about whether a reader will confuse them at a glance, across a page, while attending to something else. ΔE is calibrated for a much more favourable comparison than that, so passing a threshold is a floor rather than a guarantee.

Which question the formulae were fitted to

There is an assumption underneath every formula above, and it is rarely stated: that “how different do these look” and “can these be told apart at all” are the same question at different scales.

They are not the same question, and the two bodies of data answering them were collected by different methods on different observers half a century apart. MacAdam measured thresholds — the smallest change from a reference that could be detected at all. ΔE2000 was fitted to suprathreshold judgements, where both colours are plainly visible and the observer is asked how far apart they look. The tempting assumption is that the second is the first multiplied by a constant.

The contours disagree about orientation by a substantial margin, and the factor that would turn one family into the other varies several-fold across the diagram. So the assumption is false, and a formula fitted to one kind of judgement is being quoted for the other every time somebody cites a threshold to justify a tolerance.

This matters for reading every number in this essay. A ΔE of 1 is routinely described as a just-noticeable difference. That description imports a threshold claim into a formula fitted to suprathreshold data, and the two are related but not by any constant anybody has found. A threshold is not a unit takes the argument up properly; what belongs here is the warning that the formulae above answer a narrower question than the one they are usually asked.

What the pictures cannot show

A ΔE value is a scalar summarising a three-dimensional difference, and no figure can show what was lost in the summary. Two pairs with the same ΔE can differ in completely different ways — one in lightness, one in hue — and the number does not distinguish them. In practice that distinction often matters more than the magnitude, which is why industrial specifications frequently give separate tolerances for ΔL\Delta L, ΔC\Delta C and ΔH\Delta H rather than a single ΔE.

The figures here also present colour differences on a display, under the reader’s viewing conditions, which are exactly the variables the formulae hold fixed. A pair the formula calls ΔE 2 may be clearly different or invisible depending on the room.

Who found it, and when

CIELAB and CIELUV were both adopted in 1976, as the CIE’s response to thirty years of complaints about the 1931 diagram’s non-uniformity. The intention was to standardise on one, and the failure to agree resulted in two.

ΔE94 came from the graphic arts industry’s dissatisfaction with ΔE76. CIEDE2000 was developed by a CIE technical committee through the late 1990s and published in 2001, fitted to several combined datasets. Sharma, Wu and Dalal published their implementation notes and test data in 2005 after finding that independent implementations disagreed with one another — which is a fair indication of how tractable the formula is.

Where this goes next

The measurement problem underneath all of it is MacAdam measured it. The reason a ΔE cannot answer appearance questions is matching is not appearance. And the other place where a scale fails to be linear is the midpoint is not half.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 25 that link here.

The objects this essay names

Each one links to every other essay that touches it.

CIEDE2000CIELABΔEJust-noticeable differenceLightnessMacAdam's ellipsesThresholdTolerance