Difference and uniformity

Where the formula is not smooth

A colour difference formula is a distance, and a distance ought to vary gently. CIEDE2000's does not everywhere — it carries a hue-rotation term with a hard edge in it, and the discontinuity sits where a great many industrial samples live.

Assumes How far apart are two colours and A tolerance is a shape.

A difference formula is a distance function. Two colours go in, a number comes out, and the number is supposed to say how different they look.

Distances have properties. They are non-negative, they are zero only between identical points, they are symmetric, and — the one nobody states because it seems too obvious — they vary smoothly. Move one of the two colours a little and the distance changes a little.

CIEDE2000 has the first three. The fourth is where it gets interesting.

The ΔE = 1 contour at points across the L* = 55 plane, magnified 14×Each closed curve is the set of colours exactly one unit of difference from the dot at its centre, traced by bisection along 30 directions and drawn 14 times life size. Under ΔE2000 the contours run from 0.68 to 5.07 CIELAB units across this slice — a ratio of 7.44 — and they are not circles and not aligned with each other. A formula whose contours were circles of one radius everywhere would be claiming that CIELAB is uniform, which is what the 1976 formula claims and what the measurements refuse.a*b*semi-axes 0.68–5.07 ΔE unitsslice at L* = 55, 14× magnifiedCIELAB, D65
Fig. 1 The ΔE = 1 contour at points across a lightness slice, traced by bisection and magnified fourteen times. Each closed curve is the set of colours exactly one unit from the dot at its centre. They are not circles, they are not the same size, and they are not aligned with each other.

What the contours say

Under a perfectly uniform metric every one of those contours would be a circle of the same radius. They are not, and the departures are systematic rather than noisy.

They vary in size across the slice by a substantial ratio, which says that a unit of ΔE2000 corresponds to different CIELAB distances in different places — which is the point of the formula, because CIELAB distance is known not to be perceptually uniform and ΔE2000 exists to correct it.

They vary in shape, being elongated along the chroma direction far from the neutral axis, which encodes the measured fact that a difference in chroma is less visible at high chroma than at low.

And they rotate, and the rotation is not a smooth function of position everywhere.

The term with the edge in it

CIEDE2000’s construction has a rotation term whose whole purpose is to handle the blues, where MacAdam’s ellipses and every dataset since show an orientation that the chroma and hue weightings alone cannot produce.

The term is built from a Gaussian in hue angle, centred at 275°, and it enters the formula multiplied by a chroma-dependent factor. Both of those are smooth. The non-smoothness comes from elsewhere in the construction: the formula computes a mean hue angle between the two samples, and a mean of two angles is not a continuous function of them. Two hue angles either side of the 0°/360° wrap have a mean that jumps by 180° depending on which branch the implementation takes, and the standard specifies a rule for choosing the branch that is itself a case distinction.

The consequence is a formula that is smooth almost everywhere and has seams. Near the neutral axis there is a second and separate problem: hue angle is undefined at zero chroma, so the mean hue of a pair straddling the axis is arbitrary, and the weightings that depend on it inherit the arbitrariness.

The ΔE = 1 contour at points across the L* = 55 plane, magnified 14×. Each closed curve is the set of colours exactly one unit of difference from the dot at its centre, traced by bisection along 30 directions and drawn 14 times life size. Under the 1976 formula the contours run from 1.00 to 1.00 CIELAB units across this slice — a ratio of 1.00 — and they are not circles and not aligned with each other. A formula whose contours were circles of one radius everywhere would be claiming that CIELAB is uniform, which is what the 1976 formula claims and what the measurements refuse.
Fig. 2 The same slice under the 1976 formula. Every contour is a circle of the same radius, because ΔE*ab is Euclidean distance in CIELAB and its unit ball is a sphere everywhere. This is what a formula asserting that CIELAB is uniform looks like.

One slice is one lightness, and the formula’s behaviour is not the same at all of them.

The ΔE = 1 contour at points across the L* = 35 plane, magnified 14×. Each closed curve is the set of colours exactly one unit of difference from the dot at its centre, traced by bisection along 30 directions and drawn 14 times life size. Under ΔE2000 the contours run from 0.68 to 6.43 CIELAB units across this slice — a ratio of 9.43 — and they are not circles and not aligned with each other. A formula whose contours were circles of one radius everywhere would be claiming that CIELAB is uniform, which is what the 1976 formula claims and what the measurements refuse.
Fig. 3 The same tracing twenty lightness units lower. The contours are a different set of shapes, so the formula’s departure from a circle is a function of three coordinates rather than a property of the formula.

Twenty units above the first slice is the other direction, and between the three of them the contours change size, orientation and eccentricity together.

The ΔE = 1 contour at points across the L* = 75 plane, magnified 14×. Each closed curve is the set of colours exactly one unit of difference from the dot at its centre, traced by bisection along 30 directions and drawn 14 times life size. Under ΔE2000 the contours run from 0.68 to 5.10 CIELAB units across this slice — a ratio of 7.48 — and they are not circles and not aligned with each other. A formula whose contours were circles of one radius everywhere would be claiming that CIELAB is uniform, which is what the 1976 formula claims and what the measurements refuse.
Fig. 4 And twenty above. Between the three slices the contours change size, orientation and eccentricity, and a specification quoting a single tolerance is quoting one number for all of it.

The comparison is the argument. The 1976 formula is perfectly smooth and wrong; the 2000 formula is right about far more and has seams. Those are not competing defects — one is a modelling error and the other is a mathematical property of a piecewise construction — and a specification has to live with whichever it picks.

Setting the measurement beside the model makes the shape of the problem visible. MacAdam’s ellipses vary in size by a factor of about twenty and rotate systematically; ΔE2000’s contours vary and rotate too, and reproducing that behaviour is what the formula is for. The question is never whether the corrections are justified — they plainly are — but what form they should take, and the form chosen was a product of weighting functions with a piecewise angle computation inside it.

A box tolerance and a ΔE tolerance around the same colourA slice through CIELAB at L* 55, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 2.5 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 74% — accepted by one specification and rejected by the other, on the same measurement.gold outline: ΔE2000 = 1blue box: ±1 in each componentelongation 2.51disagree on 74%of everything either acceptsΔa*at L* 55, a* 40, b* -50 — the shape changes elsewheresame measurement, two verdictsCIE 1964 10° observer
Fig. 5 The ΔE2000 = 1 contour around a blue reference, traced point by point, with a component box drawn over it. This is the region where the rotation term does its work, and the contour’s tilt is that term made visible.

Why it matters, and to whom

For most uses it does not matter at all. A pair of samples sitting a couple of units apart in the middle of the gamut is nowhere near a seam, and the formula behaves like a well-conditioned distance.

It matters in three places.

Optimisation. Anything that minimises ΔE2000 — gamut mapping, profile construction, ink formulation, a solver looking for the nearest reproducible colour — is running a numerical method over a function that is not everywhere differentiable. Gradient-based methods can stall or oscillate at a seam, and the failure looks like a convergence problem rather than like a property of the objective.

Tolerancing near neutral. A tolerance is a shape, and near the neutral axis that shape is being computed from a hue angle that is barely defined. Two nearly-neutral samples can produce a ΔE2000 that is sensitive to a rounding difference in a coordinate, which is exactly the regime where a pass-or-fail decision is being made about a grey.

Comparability between implementations. The branch rules are specified, and implementations differ anyway. The CIE published a set of test data specifically so that implementations could be checked against each other, which is an unusual thing for a standards body to have to do and is an admission of how easy the formula is to get subtly wrong.

That the orderings differ is the strongest available argument that these are three different claims rather than three precisions of one. If they were the same model at different accuracies, one would be a monotone function of another, and a tolerance written in one could be translated into the other. It cannot, which is why converting a ΔE*ab tolerance into a ΔE2000 tolerance is a negotiation rather than an arithmetic operation.

What was computed, and how

The contours are traced rather than derived. At each sample point, for each of thirty directions in the aba^*b^* plane at fixed LL^*, a bisection finds the radius at which the difference formula returns exactly 1, to twenty-four halvings. That is slower than evaluating a closed form and it has one large advantage: it uses this site’s own implementation of the formula, so the figure is a picture of the code the rest of the site computes with rather than of an independent derivation that could agree with the standard while the code did not.

The same technique gives the metric tensor used to count distinguishable colours, and it is worth noting that the counting essay’s factor of 4.68 between the two formulae is this figure, integrated.

The generator asserts a different thing depending on which formula it was handed, and the pair is worth more than either alone.

For ΔE2000 it asserts that the contours vary in size across the slice by more than a factor of 1.3 — a structural claim rather than a stored number, and the one that would fail if the weighting functions were accidentally disabled.

For ΔE*ab it asserts that they are identical, to within 2%. That is the control. The 1976 formula is Euclidean distance in CIELAB, so its unit ball is a sphere of radius 1 at every point and in every direction, and a bisection with an off-by-one, a swapped argument or a mis-scaled direction vector could not produce identical radii by accident.

The first version of this generator asserted “varies” for both. It failed on the 1976 formula, for exactly the right reason, and the tempting repair was to loosen the threshold until both passed — which would have thrown away the only check in the figure capable of catching an error in the tracing itself.

ΔE2000 itself carries a separate assertion, checked against the CIE’s published test pairs to 10⁻⁴. Those pairs were designed to exercise the seams, which is why they are the right test data and why an implementation that passes them is more trustworthy than one checked on ordinary colours.

What a well-behaved formula would look like

It is worth being concrete about the alternative, because “smoother” on its own is not a specification.

A difference formula is well behaved if it is the length of a path under a Riemannian metric — that is, if there is a smoothly varying positive-definite tensor field on colour space and the difference between two colours is the geodesic distance under it. That form has every property anyone wants: it is symmetric, it satisfies the triangle inequality, it is smooth wherever the tensor is, and it is differentiable everywhere, so optimisation over it is well conditioned.

ΔE2000 is not of that form. It is a formula for the distance between two points that depends on both of them in a way that does not decompose into an integral along a path, and one consequence is that it does not satisfy the triangle inequality. Three colours can be arranged so that going via the middle one is shorter than going direct, which for a quantity called a distance is a real defect and is not widely advertised.

Whether that matters depends entirely on the use. For pass-or-fail against a reference it does not: only pairwise differences from one point are being computed. For anything that chains differences, or that interpolates, or that searches, it does.

A seven-by-seven grid is a coarse sample of a plane, and the honest check is whether a finer one finds contours outside the range the coarse one reported.

The ΔE = 1 contour at points across the L* = 55 plane, magnified 14×. Each closed curve is the set of colours exactly one unit of difference from the dot at its centre, traced by bisection along 30 directions and drawn 14 times life size. Under ΔE2000 the contours run from 0.68 to 5.63 CIELAB units across this slice — a ratio of 8.25 — and they are not circles and not aligned with each other. A formula whose contours were circles of one radius everywhere would be claiming that CIELAB is uniform, which is what the 1976 formula claims and what the measurements refuse.
Fig. 6 The same slice at eleven centres per axis rather than seven. The contours now run from 0.68 to 5.63 CIELAB units, a ratio of 8.25 against the coarser grid’s 7.44 — so the coarse sample understated the spread rather than inventing it.

How far the contours actually vary

The gate demands a factor of 1.3 and the phenomenon is much larger than that. Across the L* = 55 slice, at the twenty-nine sample points that fall inside a radius of 110, the mean radius of the ΔE00 = 1 contour runs from 0.827 at the neutral axis to 4.277 in the blue corner at a* = −73, b* = −73. That is a ratio of 5.17 in radius and 26.8 in the area each contour encloses.

A unit of ΔE2000 is therefore a CIELAB step five times longer in a saturated blue than in a grey, and it covers twenty-seven times the ground. That is the formula doing exactly what it exists to do — the chroma weighting is there to say a difference at high chroma must be larger to be equally visible — and the size of it says how much of the number is weighting rather than distance.

The individual contours are far from round as well. Within a single contour, the ratio of the longest radius to the shortest runs from 1.49 at its most circular to 5.99 at its most elongated: even the roundest place on the slice is half again as long one way as the other.

So the gate’s 1.3 sits four times below the phenomenon, which is the right way round for a gate. It is there to catch the weighting functions being switched off, not to record the number.

The triangle inequality, with a witness

That ΔE2000 fails the triangle inequality is asserted above without an example. One is easy to find, and having it changes the claim from a caution into a measurement.

Three colours, all at chroma an ordinary paint could reach: a blue-green at L* 56, a* −52, b* −30; a near-neutral at L* 50, a* −3, b* −1; and a red at L* 45, a* 59, b* 10. Direct, the blue-green and the red are 89.75 apart. By way of the grey the journey costs 24.55 and then 17.55, which is 42.10 — so the direct distance is 2.13 times the sum of the two legs it can be walked in.

Nor is such a triple rare enough to have to be hunted for. Drawing three hundred thousand random triples with chroma up to 60 — an ordinary ink or paint gamut — 2.41 per cent violate the inequality, and restricting to chroma under 40 barely moves it, at 2.34 per cent. About one triple in forty.

Two controls make that a property of the formula rather than of the sampling. The 1976 formula never violates it — nought out of two hundred thousand, which it cannot, being a Euclidean distance in a vector space. And ΔE94 violates it at 0.65 per cent, a quarter of ΔE2000’s rate, which is what a formula that has chroma weighting but no hue rotation and no mean-hue branch ought to do.

The defect is therefore graded across the three formulae in the same order as their accuracy. That is this essay’s trade, arriving in a form that can be counted instead of described — each correction buys agreement with the data and spends one of the properties that made the previous formula a distance.

What the picture cannot show

The contours are magnified fourteen times. At life size a ΔE = 1 contour is about one part in three hundred of the slice and would be smaller than a pixel.

That magnification is honest and it distorts one thing: it makes the variation between contours look larger relative to the plane than it is. A reader should take from the figure that the contours differ in size, shape and orientation, and should not take any impression about how large a just-noticeable difference is compared with the gamut — for that, the count is the right figure.

The other thing no figure here can show is a seam. The discontinuity is in the formula’s derivative with respect to hue angle, and a contour plot at a grid of points steps straight over it. Rendering it would need a much finer sampling in one dimension and a plot of the derivative rather than the value, and it would show a spike whose height depends on the sampling — which is a picture of a numerical artefact as much as of the formula.

So the seam is described here and not drawn, which is an unsatisfying position for a figure-first site and is the correct one.

Where the model stops

The whole of this analysis is about a formula, not about vision.

ΔE2000 is a fit. Its seams are properties of the fitting form somebody chose — piecewise, with a Gaussian bolted on — and not properties of human colour discrimination, which as far as anybody knows varies perfectly smoothly. A better-conditioned formula with the same predictive accuracy is possible in principle, and several have been proposed.

CAM16-UCS is the most serious of them, and it is a different kind of object: rather than correcting distances in CIELAB with weighting functions, it builds a space out of an appearance model and takes Euclidean distance there. That gives smoothness by construction, because there is no piecewise term anywhere in it, and it gives the appearance model’s viewing-condition arguments for free.

It has not displaced ΔE2000 in industry, for the reason everything in this subject fails to displace anything: tolerances written in ΔE2000 are in contracts, instruments report it, historical data is in it, and a formula that is better conditioned but not comparable with twenty years of records is not obviously an improvement to anyone who has to sign for a batch of fabric.

Pulling the ring of centres in towards the neutral axis asks the same question of the part of the plane most production work actually sits in.

The ΔE = 1 contour at points across the L* = 55 plane, magnified 14×. Each closed curve is the set of colours exactly one unit of difference from the dot at its centre, traced by bisection along 30 directions and drawn 14 times life size. Under ΔE2000 the contours run from 0.68 to 4.97 CIELAB units across this slice — a ratio of 7.29 — and they are not circles and not aligned with each other. A formula whose contours were circles of one radius everywhere would be claiming that CIELAB is uniform, which is what the 1976 formula claims and what the measurements refuse.
Fig. 7 The same construction over a ring of radius 60 rather than 110. The spread narrows to 0.68 through 4.97, a ratio of 7.29, so moving inwards to the less saturated colours buys a seventh of the anisotropy and not an order of it.

What a practitioner should do about it

Very little, and the “very little” is specific.

For tolerancing, keep using it. The seams are not where ordinary industrial samples sit, the formula’s accuracy advantage over its predecessors is real and large, and comparability with existing contracts is worth more than mathematical tidiness.

For anything near neutral, be careful with the hue-dependent terms. A pair of nearly-grey samples is being judged partly on a hue angle that is barely defined, and two implementations can disagree. Where a grey has to be toleranced tightly, component tolerances in CIELAB are better conditioned than a ΔE2000 limit, which is one of the few cases where the box beats the ellipsoid on grounds other than convenience.

For optimisation, use something else as the objective and ΔE2000 as the check. CAM16-UCS or ΔE94 as the thing being minimised, with the final answer scored in ΔE2000, gets a well-conditioned search and a comparable number.

Check the implementation against the published test data, always. The pairs Sharma and colleagues supplied exist because the formula is easy to get subtly wrong, and an implementation that has not been checked against them is an implementation whose behaviour at the seams is unknown.

Who found it, and when

CIEDE2000 is Luo, Cui and Rigg, 2001, fitted to a combination of several datasets including the RIT-DuPont, Witt and Leeds data.

The discontinuities were reported almost immediately. Sharma, Wu and Dalal’s 2005 paper is the standard reference: it works through the formula’s implementation in detail, supplies the test data that everybody now checks against, and catalogues the places where the specification is ambiguous or the function ill-behaved. Their conclusion is worth quoting in spirit — the formula is a considerable improvement in accuracy and a considerable step backwards in mathematical hygiene, and the trade was probably worth making.

The CIE’s own technical report acknowledges the issues, which is unusually candid, and the formula remains the recommendation. That is a defensible position: the discontinuities are real, they are small, they are in regions where the fitting data was sparse anyway, and no proposed replacement has enough of an accuracy advantage to justify breaking comparability.

The general shape of the problem

There is a pattern here that recurs across this whole subject, and it is worth naming because it explains why the field looks the way it does.

A simple model is proposed. It is elegant, it has good mathematical properties, and it is wrong about the measurements. Corrections are added, each justified by data, each destroying one of the elegant properties. The corrected model predicts much better and is much uglier, and the ugliness is not decoration — it is the shape of the disagreement between the simple model and the world.

CIELAB was the simple model and ΔE2000 is the corrected one. The same story runs for the chromaticity diagram and its 1976 replacement, for colorimetry and appearance modelling, and for Planckian colour temperature and the Duv that had to be added beside it.

In every case the corrections are right, the ugliness is informative, and the simple version survives in practice for decades afterwards because it is what everybody’s data is in. That is not institutional failure. It is what happens when a model is also a unit of exchange.

Where this goes next

The counting essay is this figure integrated, and how many colours are there is where the variation shown here turns into a factor of nearly five in an answer people quote.

Downward, how far apart two colours are sets out the formulae themselves, and a tolerance is a shape is where the contours above stop being a diagram and start being a contract.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

  • Which of two is worse ciede2000 · cielab · δe · just-noticeable difference · macadam's ellipses · perceptual uniformity · quality control · specification · threshold · tolerance
  • A difference has no place ciede2000 · cielab · δe · just-noticeable difference · quality control · specification · threshold · tolerance
  • A name is not a threshold ciede2000 · δe · just-noticeable difference · perceptual uniformity · specification · threshold · tolerance
  • Three constants nobody quotes ciede2000 · δe · perceptual uniformity · quality control · specification · tolerance
  • A catalogue is not a vocabulary cielab · δe · perceptual uniformity · quality control · specification
  • A colour has a name ciede2000 · cielab · δe · perceptual uniformity · specification

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CIEDE2000CIELABΔEJust-noticeable differenceMacAdam's ellipsesPerceptual uniformityQuality controlSpecificationThresholdTolerance