Difference and uniformity

MacAdam measured it

In a perceptually uniform space the just-noticeable-difference contours would be circles of equal size. MacAdam's ellipses are neither, by a factor of eighty — and transforming them into each candidate space settles which spaces improved matters and by how much.
14 min read 6 figures Computed, not quotedSay which colour

In the late 1930s David MacAdam sat a single observer in front of a colorimeter and asked, for each of twenty-five reference colours, how far a second colour had to move before the difference became detectable. He repeated each setting enough times to get a standard deviation, and he did it in every direction.

The results are ellipses, they are on every serious discussion of colour spaces since, and they are the reason the chromaticity diagram cannot be used to measure anything.

MacAdam's discrimination ellipses, drawn 10 times actual sizeTwenty-five ellipses of colours indistinguishable from their centres. They are drawn at 10× because at true scale most are thinner than a line. Their areas vary by a factor of 74, which is the whole result: a step of the same size in xy means very different things in different places.0.00.20.40.60.80.00.20.40.60.8xyellipses at 10× scaleCIE 1931 2° observer
Fig. 1 The twenty-five ellipses, drawn ten times actual size because at true scale most are thinner than a line. Each encloses the colours indistinguishable from its centre. Their areas vary by a factor of about eighty across the diagram.

What the picture says

Two things, and both are damaging to the diagram they are drawn on.

They are not circles. Discrimination is directionally dependent — a step of a given size in one direction is detectable while the same step at right angles is not. Most of the ellipses are elongated by a factor of two to four, and some more.

They are not the same size. This is the worse problem. The largest, in the green region, has roughly eighty times the area of the smallest, in the blue. A step of 0.010.01 in xx and yy is a large, obvious change near the blue corner and an invisible one in the greens.

The consequence is that Euclidean distance on the chromaticity diagram is not a measure of anything perceptual. Two pairs of colours the same distance apart on the plot can be one obviously different and one indistinguishable. Any statement of the form “these colours are close together on the diagram” is, on its own, meaningless.

What the diagram’s shape hides

This also disposes of a common misreading of the horseshoe. The green region occupies a very large share of the diagram’s area, which invites the idea that human vision is somehow richest in green.

It is the opposite. The green region is large and its ellipses are the largest, which means it contains rather few distinguishable colours for its size. The blue corner is small and its ellipses are tiny, so it packs many distinguishable colours into little area. The apparent prominence of green is an artefact of the projection, and the projection was chosen in 1931 for computational convenience rather than for perceptual fidelity.

Fixing it: two strategies

Everything since has been an attempt to deal with this, and the attempts fall into two camps.

Fix the space. Transform the coordinates so that the ellipses become more nearly circular and more nearly equal, then use plain Euclidean distance. CIELUV (1976), CIELAB (1976) and Oklab (2020) are all of this kind.

Fix the metric. Keep the space and weight the distance according to where in the space it is being measured. ΔE94 and ΔE2000 are of this kind, and they are what most industrial practice uses.

The second approach is an admission that the first did not fully succeed. If CIELAB were uniform, ΔE76 — plain Euclidean distance in it — would have been sufficient, and the CIE would not have issued two replacements.

Measuring how far the fixes got

The claim that a space is more uniform than another is testable, and the ellipses are the instrument. Transform each ellipse’s boundary into the candidate space, measure the distance from the transformed centre to each boundary point, and two numbers come out:

Anisotropy — the mean ratio of longest to shortest radius per ellipse. A perfect space gives 1, meaning every discrimination contour is a circle.

Spread — the largest mean radius across the twenty-five, divided by the smallest. A perfect space gives 1, meaning a step of a given size means the same thing everywhere.

Both are ratios, which matters: Lab, Luv and Oklab have natural magnitudes differing by more than an order of magnitude, so any absolute distance would just be measuring the scaling.

Four colour spaces ranked by how uniform they actually areThe spread — largest ellipse over smallest, after transformation — for each space. Raw chromaticity is worst at 10.4. The three perceptual spaces are all far better and close to one another, with CIELUV at 2.21 and Oklab at 2.57 — a smaller gap than the usual advocacy suggests.CIE 1931 xy10.42aniso 2.95CIELAB3.24aniso 3.42CIELUV2.21aniso 2.37Oklab2.57aniso 2.27spread: largest JND ellipse ÷ smallest (1 is perfect)measured, not quotedat Y = 0.4
Fig. 2 The four spaces ranked by spread. Raw chromaticity is worst at 10.4. The three perceptual spaces are all far better and much closer to one another than the usual advocacy suggests.

The measured result:

space spread anisotropy
CIE 1931 xy 10.42 2.95
CIELAB 3.24 3.42
CIELUV 2.21 2.37
Oklab 2.57 2.27

Three things in that table are worth dwelling on.

The perceptual spaces are a large improvement on raw chromaticity. Spread falls from 10.4 to between 2.2 and 3.2. This is real and is the reason the 1976 spaces exist.

None of them is uniform. A perfectly uniform space scores 1 on both columns. The best result here is 2.21, meaning a step of fixed size still means twice as much in one place as another. That residual non-uniformity is exactly what ΔE94 and ΔE2000 were built to patch.

CIELAB fixes the spread far better than the shape. Its spread improves from 10.4 to 3.2, while its anisotropy actually gets worse than raw chromaticity — 3.42 against 2.95. CIELAB equalises how big the ellipses are much more successfully than it makes them round.

That last finding is not one this site expected to produce, and it explains something about the ΔE2000 formula. Its most awkward features are the chroma and hue weighting terms SCS_C and SHS_H, which scale differences by where in the space they occur. Those terms are compensating for precisely this: shape errors that CIELAB left behind.

What is being quoted and what is being computed

The ellipse parameters themselves are quoted, and this is a case where quoting is correct. They are measurements — twenty-five centres with axis lengths and orientations, from an experiment nobody is going to repeat. This site’s rule is that measurements should be quoted and derivable results should not, and MacAdam’s table is squarely the first kind.

Everything downstream is computed. The transformation into each space, the radii, the anisotropy, the spread, the ranking — all done at build time from the ellipse parameters and the conversion functions, so the table above cannot drift out of agreement with the figures.

The experiment’s limits, which are real

MacAdam’s data is foundational and it is also narrow, and the caveats are usually omitted.

One observer. The published ellipses come from a single subject, identified as PGN. Later work with more observers found the general pattern holds and the details vary substantially between people.

One luminance, one surround. The measurements were made at fixed luminance against a uniform surround. Discrimination changes with both, so the ellipses are a slice through a situation with more dimensions than the diagram shows.

Threshold, not suprathreshold. They measure the smallest detectable difference. Whether differences ten times the threshold are judged proportionally is a separate question, and the evidence is that they are not — which is why colour-difference formulae for industrial tolerance are fitted to suprathreshold judgements rather than to MacAdam’s data.

Chromaticity only. Lightness differences were not part of the experiment, and lightness is where a good deal of practical colour difference lives.

So the ellipses are the right instrument for the specific question of whether a chromaticity space is uniform, and are not a general model of colour discrimination.

Reading the numbers back into practice

The uniformity measurement has consequences that show up well outside colour science.

A gradient stepped uniformly in xyxy will not look uniform. The steps will be invisible in the greens and obvious in the blues, which is the spread ratio of ten appearing directly. Generating a gradient in a perceptual space instead reduces the ratio to about two, which is not perfect and is five times better.

A tolerance specified in xyxy means different things in different places. Industrial standards do not specify tolerances that way, and the reason is exactly this measurement.

A palette spaced evenly on the chromaticity diagram will be badly unbalanced. Some pairs will be barely distinguishable and others obviously different.

Three colour-difference formulae, disagreeingΔE76, ΔE94 and ΔE2000 for the same nine pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.ΔE76ΔE94ΔE2000same pairs, three answersCIELAB, D65
Fig. 3 The formulae that exist because of the residue. ΔE94 and ΔE2000 are corrections applied on top of CIELAB, weighting differences according to where in the space they are measured — which is precisely what the anisotropy figure above says is still needed.

What the ellipses do not measure

Four limits, all of them reasons not to over-read the numbers.

They are chromaticity only. Lightness differences were not part of the experiment, and a great deal of practical colour difference is lightness. A space could score perfectly on this measurement and still be badly non-uniform in lightness.

They are threshold measurements. Whether a difference ten times the threshold is judged ten times as large is a separate question, and the evidence says no. Colour-difference formulae for industrial tolerance are fitted to suprathreshold judgements rather than to MacAdam’s data, which is why ΔE2000 is not simply a metric derived from these ellipses.

One observer, one condition. A single subject, at one luminance, against one surround. Later work confirms the general pattern and finds substantial individual variation.

Discrimination is not appearance. These measure the smallest detectable difference, which is a different question from how different two clearly different colours look. The distinction runs through the whole subject.

The CIE 1931 chromaticity diagram with its unreachable region markedThe spectral locus encloses every chromaticity a human eye can see. Cells inside the sRGB triangle are drawn in their own colour; the 85 per cent outside it are hatched, because no value this display accepts is the colour belonging there.0.00.20.40.60.80.00.20.40.60.8xyD65460480500520540560580600620hatched: outside sRGB15% of the visible area is reachableat luminance Y = 0.55CIE 1931 2° observer
Fig. 4 The diagram the ellipses are drawn on, with its reachable region marked. Note where the two facts coincide: the greens, where discrimination is worst and the ellipses largest, are also largely unreachable — so the region that looks most generous on this diagram is the least useful in both senses.

What was computed here

The uniformity measurement transforms each ellipse boundary at sixty-four points, at a stated luminance, and records the extreme and mean radii in the target space. The two summary ratios are computed from those.

The check that keeps the measurement honest is a ranking assertion: raw chromaticity must come out less uniform than CIELAB, and its spread must exceed five. MacAdam’s entire result is that xyxy is badly non-uniform, so a measurement that failed to reproduce that would be measuring nothing, and the gate would catch it before any figure was drawn.

The ellipses are also drawn magnified, which the caption states, because at true scale they would be sub-pixel. The magnification is a stated parameter rather than a fudge, and the areas quoted are computed from the unmagnified parameters.

The magnification is a stated parameter rather than an adjustment: the areas quoted in the text are computed from the unmagnified table, and only the drawing is scaled. At true scale the whole set of twenty-five would be close to invisible, which is itself the most striking fact about human chromatic discrimination.

What the residual non-uniformity looks like

The summary numbers compress twenty-five measurements into two, so it is worth seeing the distribution.

How uniform CIELAB is, measured against MacAdam's ellipsesFor each of the 25 ellipses, the ratio of its longest to its shortest radius after transforming into CIELAB (a circle would give 1), and its mean radius. Mean anisotropy is 3.42 and the largest ellipse is 3.2 times the smallest. A perfectly uniform space would give 1 and 1.1CIELAB: anisotropy 3.42, spread 3.2the 25 ellipses, in table ordergold: max/min radius per ellipse · grey: relative size1 would be a circleat Y = 0.4
Fig. 5 The per-ellipse anisotropy in CIELAB. A perfectly uniform space would be a flat line at 1. The variation across the twenty-five is what the summary figure of 3.42 averages, and the worst cases are considerably worse than the mean.

The spread within a space matters as much as the spread between spaces. A space with a good average and a few very bad regions will behave unpredictably, and the bad regions are typically the saturated blues — which is exactly where CIELAB’s known hue-linearity failure lives and exactly what the ΔE2000 rotation term addresses.

What replaced the diagram

The direct descendant of MacAdam’s critique is the CIE 1976 uvu'v' diagram, which is a projective transformation of xyxy chosen to make the ellipses more nearly equal.

It is a genuine improvement and it is still a chromaticity diagram, so it inherits the deeper problem: it discards luminance, and discrimination depends on luminance. Any two-dimensional diagram is a slice, and the ellipses on it are the ellipses at one luminance.

That is the reason the modern answer is a three-dimensional space rather than a better diagram. CIELAB, CIELUV and Oklab all carry lightness as a coordinate, and none of them can be drawn as a plane without the same loss. The horseshoe survives in textbooks because it is a good picture of mixing, which it genuinely is, and it has not been a good picture of difference since 1942.

The through-line

MacAdam measured a defect in 1942. Every colour space and every difference formula since is a response to it, and none has removed it.

Three colour-difference formulae, disagreeingΔE76, ΔE94 and ΔE2000 for the same nine pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.ΔE76ΔE94ΔE2000same pairs, three answersCIELAB, D65
Fig. 6 The formulae the shortfall produced, disagreeing on the same nine pairs by more than a just-noticeable difference.

The experiment also stands as a model of how to settle an argument about a colour space. The claim “this space is perceptually uniform” is not a matter of taste or of design intent; it is a prediction about discrimination contours, and discrimination contours can be measured. Any new space arriving with a uniformity claim can be put through exactly this procedure, which is what makes the comparison on this page a measurement rather than a preference.

The historical shape of this is worth noticing too. A coordinate system was adopted in 1931 for computational convenience, a defect in it was measured in 1942, and the field has spent the eighty years since building corrections rather than replacing the foundation. That is not obviously the wrong choice — replacement would have invalidated the accumulated record — and it does mean the complexity of modern colour difference is inherited rather than intrinsic.

What the pictures cannot show

The ellipses are drawn on a diagram whose colours are mostly unavailable, so the region where they are largest — the greens — is largely hatched. That is unavoidable and mildly ironic: the part of the diagram whose colours are hardest to display is also the part where discrimination is worst.

The magnification also makes them look like large regions when they are the opposite. At true scale the entire set would be almost invisible, which is the fact worth carrying away: human chromatic discrimination is extremely fine, and the diagram is enormous compared to any step anyone can see.

And a transformed ellipse is not shown as a shape anywhere here. The uniformity figures report the radii statistically rather than drawing the transformed contours, because the transformed shapes live in three dimensions and projecting them back to a plane would reintroduce exactly the kind of distortion being measured.

Who found it, and when

MacAdam published the ellipses in 1942, working at Kodak. The experiment was extraordinarily laborious — thousands of individual settings — and its conclusion was immediate and unwelcome: the coordinate system the CIE had adopted eleven years earlier was not fit for measuring differences.

Judd had already proposed uniform-chromaticity transformations in the 1930s, and the CIE adopted the UCS diagram in 1960 and the uvu'v' diagram and CIELAB in 1976. ΔE94 arrived in 1994 and CIEDE2000 in 2001, each patching residual errors in the previous.

Björn Ottosson’s Oklab, published in 2020, was fitted with attention to hue linearity specifically — CIELAB’s worst-known failure is that blues shift toward purple as they darken — and the measurement above suggests it succeeded on shape while not improving on CIELUV for spread.

Where this goes next

The formulae built to patch these residuals are how far apart are two colours. The diagram the ellipses are drawn on is most of this diagram cannot be shown. And the other place non-linearity bites is the midpoint is not half.