Difference and uniformity

A threshold is not a unit

MacAdam measured the smallest difference anyone could detect. ΔE2000 was fitted to how far apart plainly different colours look. The two are quoted interchangeably, and the contours they produce are not even the same shape.

Assumes How far apart are two colours and MacAdam measured it.

“A ΔE of 1 is a just-noticeable difference” is the most repeated sentence in applied colour, and it contains a substitution nobody notices.

A just-noticeable difference is a threshold: the smallest change detectable at all. ΔE2000 was fitted to suprathreshold judgements, where both colours are plainly visible and an observer is asked how far apart they look. Those are different experiments answering different questions, and the assumption joining them is that one is the other multiplied by a constant.

Threshold and suprathreshold contours, normalised to the same size. At five of MacAdam's centres: the measured just-noticeable-difference ellipse in grey and the ΔE2000 = 1 contour in gold, each scaled to the same mean radius so that only shape and orientation are being compared. A scale change preserves orientation exactly, so any rotation between the pair settles the question. They differ by 24° on average and by 70° at worst, and the ratio between their sizes varies 4.8-fold across the diagram — so no single factor turns one into the other.
Fig. 1 At five of MacAdam’s centres: the measured threshold ellipse and the ΔE2000 unit contour, each normalised to the same mean radius so only shape and orientation are compared. A pure scale change preserves orientation exactly, so any rotation between a pair settles the question on its own.

Five centres is a sample of twenty-five, and the comparison is worth making at more of them before anything is concluded from it.

Threshold and suprathreshold contours, normalised to the same size. At five of MacAdam's centres: the measured just-noticeable-difference ellipse in grey and the ΔE2000 = 1 contour in gold, each scaled to the same mean radius so that only shape and orientation are being compared. A scale change preserves orientation exactly, so any rotation between the pair settles the question. They differ by 24° on average and by 70° at worst, and the ratio between their sizes varies 4.8-fold across the diagram — so no single factor turns one into the other.
Fig. 2 Five different centres, chosen to interleave with the five above. The mismatch between the measured contour and the formula’s unit is the same kind of mismatch and a different size at each one, which is what a systematic disagreement looks like rather than a sampling accident.

The observer the two objects are computed for is the other thing that can be changed here, and only one of them was measured on either.

Threshold and suprathreshold contours, normalised to the same size. At five of MacAdam's centres: the measured just-noticeable-difference ellipse in grey and the ΔE2000 = 1 contour in gold, each scaled to the same mean radius so that only shape and orientation are being compared. A scale change preserves orientation exactly, so any rotation between the pair settles the question. They differ by 24° on average and by 70° at worst, and the ratio between their sizes varies 4.8-fold across the diagram — so no single factor turns one into the other.
Fig. 3 The first five again under the ten-degree observer. Both objects move — the threshold data was collected on a two-degree field and the formula is computed from whatever observer it is handed — and they do not move together.

Drawn at fewer centres the shapes themselves are legible, which is the last thing worth looking at before the comparison is put into words.

Threshold and suprathreshold contours, normalised to the same size. At five of MacAdam's centres: the measured just-noticeable-difference ellipse in grey and the ΔE2000 = 1 contour in gold, each scaled to the same mean radius so that only shape and orientation are being compared. A scale change preserves orientation exactly, so any rotation between the pair settles the question. They differ by 24° on average and by 70° at worst, and the ratio between their sizes varies 4.8-fold across the diagram — so no single factor turns one into the other.
Fig. 4 And three centres drawn larger, for the shape rather than for the census. What a threshold contour and a suprathreshold unit have in common at these centres is their centre.

Two experiments, and why they differ

A threshold experiment shows an observer a reference and a comparison, makes the difference very small, and asks whether there is any difference at all. The answer is yes or no, the measurement is a probability, and the threshold is conventionally where detection reaches some criterion — 50%, or 75%, depending on the method. That criterion is a convention, chosen rather than discovered, as much so as the tolerance it eventually becomes. MacAdam’s 1942 experiment is of this kind.

A suprathreshold experiment shows an observer two colours that are obviously different and asks how different. The answer is a magnitude, and it is obtained by comparison against other pairs — is this pair further apart than that pair — because a bare magnitude judgement is unreliable. The datasets ΔE2000 was fitted to are of this kind.

There is no reason in advance to expect the second to be the first scaled up. Detection near threshold is limited by noise in the visual system; magnitude estimation well above threshold is limited by whatever the system does to turn differences into a sense of distance. Those are different mechanisms, and they could easily produce different geometries.

They do produce different geometries

Two measurements settle it, and the second is the decisive one.

The scale factor is not constant. Around each of MacAdam’s twenty-five centres, trace the locus of colours at exactly ΔE2000 = 1 and compare its size with the measured threshold ellipse. If suprathreshold difference were scaled threshold difference, the ratio would be the same at all twenty-five. It varies by a factor of 4.8.

The orientations disagree. This is the argument that does not depend on any scaling at all. A pure scale change preserves orientation exactly — stretch a shape uniformly and its long axis stays where it was. The two families of contours differ in orientation by 23.6° on average and by 70° at worst.

A single rotation settles the question. Two shapes that differ in orientation cannot be related by any scale factor whatsoever, constant or otherwise, so no amount of rescaling turns one family into the other.

What that means for the sentence

“ΔE 1 is a just-noticeable difference” is quoting a threshold to describe a suprathreshold metric. The formula answers a narrower question than the one it is asked.

What ΔE2000 = 1 actually means is: as far apart as pairs the fitting dataset placed at that distance, under the fitting dataset’s viewing conditions — side by side, mid-grey surround, specified illuminant, a trained observer looking for differences. That is a useful quantity. It is not a detection threshold, and the two are related by no constant anybody has found.

The practical consequences run in both directions, which is what makes the substitution hard to notice.

Separated in space or time, much larger differences go unnoticed. A ΔE of 3 between two panels of a car, on opposite sides, is often invisible, and the instruments would disagree by less than that. Colour memory is poor and the comparison is not simultaneous.

Side by side with a shared edge, much smaller differences are obvious. Two paints meeting at a seam show a ΔE well under 1, because a visible edge is a far more sensitive test than two separated patches. This is the reason a touch-up is visible when a repaint is not, and it is the same edge sensitivity that makes the illusion figures on this site work.

So the same number understates the tolerance in one arrangement and overstates it in another, by factors of several in each direction. The number is not wrong. The arrangement it was measured in is part of the number, and it does not travel with it.

Why threshold data was used anyway

Given all that, it is worth asking why colour difference formulae are so entangled with threshold measurements.

Because the threshold data existed. MacAdam measured in 1942, thoroughly, and produced the only quantitative map of how discrimination varies across the diagram. Suprathreshold datasets came decades later, are smaller, and disagree with each other more.

More importantly, the threshold data got the shape of the problem right even where it got the metric wrong. The great insight in MacAdam’s ellipses is not their size but that they are ellipses at all, oriented differently in different places — that discrimination is anisotropic, and that a colour space with a Euclidean metric cannot be uniform. That structural finding drove the design of every space and every formula since, and it does not depend on the threshold-versus-suprathreshold distinction at all.

MacAdam's discrimination ellipses, drawn 10 times actual size. Twenty-five ellipses of colours indistinguishable from their centres. They are drawn at 10× because at true scale most are thinner than a line. Their areas vary by a factor of 74, which is the whole result: a step of the same size in xy means very different things in different places.
Fig. 5 The measurement that shaped the field. Twenty-five ellipses of colours indistinguishable from their centres, at ten times actual size because at true scale most are thinner than a line. What matters is not their size but that they are ellipses, differently oriented.

The correction factors that hide the problem

Every modern difference formula carries parametric factors — usually written kL, kC and kH — that scale the three components before the distance is computed.

They exist because the formula’s answer depends on the viewing arrangement. The textile industry uses kL = 2, which halves the weight given to lightness differences, because in fabric a lightness difference of a given size matters less than a chroma difference of the same size. Automotive practice uses different values again.

Those factors are an admission. A formula that needed no correction for the viewing arrangement would not have parameters for it. What they encode is that the fitting dataset’s conditions were one arrangement among many, and that using the formula elsewhere requires adjusting it — with values determined empirically, per industry, by finding what makes the number match the complaints.

The uncomfortable part is that the factors are usually left at 1 by default, and a great deal of software does not expose them at all. So the most common use of the formula is with the correction for viewing arrangement silently set to “the arrangement of the original experiment”.

What was computed here

The comparison is done by tracing rather than by algebra, and that choice matters.

Around each MacAdam centre, the ΔE2000 = 1 contour is found by bisection along seventy-two directions in the chromaticity plane. Fifty bisection steps per direction gives a radius accurate enough to measure an axis ratio, which a coarser method would not.

The contour is not assumed to be an ellipse. ΔE2000 has discontinuous derivatives in places — its hue rotation term and its chroma weighting are piecewise — so its unit contour is not an ellipse and fitting one would smooth over exactly the structure being measured. What is measured is the longest and shortest radius of the traced contour and the direction of the longest.

The two families are then normalised to the same mean radius before comparison, so the figure compares shape and orientation only. That is the honest presentation given that the absolute scales are in different units and there is no principled conversion between them — which is the essay’s conclusion, so assuming one would beg the question.

Two assertions run in the gate: the scale ratio must vary by a wide margin across the diagram, and the orientation gap must exceed a stated angle. The second is the one that carries the argument, and it is stated separately for that reason — a scale change preserves orientation, so orientation disagreement is sufficient on its own and does not depend on the first.

The MacAdam data’s own limits

Being clear about what the threshold data is worth cuts both ways, and it is worth stating the case against relying on it too heavily.

MacAdam’s experiment used one observer. That is a real limitation and a smaller one than it sounds: the variation across the diagram is a factor of eighty, and no plausible between-observer difference is anywhere near that. The shape of the result is safe; the exact ellipse parameters are one person’s.

It measured chromaticity only, at fixed luminance — a projection that discards a dimension — so it says nothing about how discrimination varies with lightness — which is a substantial part of the space and a substantial part of what a difference formula has to handle.

And it used a matching-by-adjustment method rather than a forced choice, which is known to give slightly different thresholds from modern procedures.

None of that undermines the argument here, because the argument is about the relationship between two kinds of measurement rather than about the exact values of either. But an essay claiming that a threshold and a suprathreshold measure disagree should not be selectively sceptical about only one of them.

The scale measurement is evidence for the formula

The two measurements are presented as a pair and only one of them tells against ΔE2000. Read against what the formula was up against, the first one is a considerable success.

MacAdam’s ellipses vary in size across the diagram by a factor of eighty. Dividing each by the formula’s own unit contour leaves a ratio that varies by 4.8. So the formula removes a factor of 16.7, which in log terms is 64 per cent of the spread — nearly two thirds of the eighty-fold anisotropy in discrimination, absorbed by a formula fitted to a different kind of experiment altogether.

Quoted alone, it varies by a factor of 4.8 reads as a failure. Quoted against the eighty it started from, it reads as the formula doing most of what a formula could do and leaving a third. Both readings are true and the second is the one a reader needs in order to weigh the first.

Which leaves the orientation carrying the whole argument, exactly as the essay says it does — and makes the ordering of the two sections slightly unfortunate, since the first one is evidence in the other direction.

How much disagreement 23.6 degrees is

An orientation gap needs a scale to be read against, and there is a natural one: two axis directions chosen independently differ by 45 degrees on average, once the difference is folded into the 0-to-90 range an axis lives in.

The measured 23.6 degrees is 52 per cent of that. So the two families of contours are neither aligned nor unrelated — they agree about half as much as they could, and disagree about half as much as two random sets would.

That is a more careful statement than they do produce different geometries and it cuts both ways. The disagreement is real and large enough to rule out any scaling; the agreement is also real, and it is what one would expect if both experiments are measuring the same underlying anisotropy through different mechanisms. A formula fitted to magnitude judgements gets the orientation of discrimination half right, which is neither nothing nor enough.

The worst case is the number that settles it. Seventy degrees is 78 per cent of the maximum possible 90, so at that centre the two contours are very nearly perpendicular: the direction in which the threshold data says colours are hardest to tell apart is close to the direction in which the formula says they are easiest. One centre like that is a counterexample to any scaling, and it is stronger evidence than the mean.

The precision is in the wrong coordinate

One note about the method, because it is the kind of thing worth catching in one’s own work.

The contour is traced along 72 directions with 50 bisection steps in each. Fifty bisections resolve a radius to about one part in 10¹⁵ — full double precision — while 72 directions resolve an orientation to about 2.5 degrees, being half a step.

The measurement the argument rests on is the orientation, and it is the one measured to 2.5 degrees while the radius is measured to fifteen digits. Against a mean gap of 23.6 degrees that is a tenth, and against the individual centres nearest agreement it could be a large fraction.

Nothing here is wrong: 2.5 degrees is plenty for a 23.6-degree mean and a 70-degree extreme, and the bisection is cheap. But the effort is in the wrong coordinate by about thirteen orders of magnitude, and doubling the direction count would cost the same as removing thirty bisection steps nobody needs. A contour traced for its shape wants angular resolution and a contour traced for its size wants radial resolution, and this one is traced for its shape.

One threshold that is a unit

There is a place where a threshold genuinely functions as a unit, and noticing it sharpens what is wrong everywhere else.

The perceptual quantiser used for HDR delivery is built by integrating a threshold. Its design rule is that one step of code value equals one just-noticeable contrast step, at every luminance, and the curve is what that requirement produces. Here the threshold really is being used as a unit and the use is legitimate.

The difference is that the quantiser only ever asks a threshold question. It is deciding whether adjacent code values are distinguishable, which is exactly what a detection threshold measures, and it never asks how far apart two plainly different values look. A threshold integrated to answer threshold questions is sound; the same integration used to predict magnitude is the substitution this essay is about.

Where the model stops

The comparison here is between a measured threshold contour and a formula’s unit contour, and the second is not a measurement. ΔE2000 is a fit to suprathreshold data, so comparing against it is comparing against a summary of that data rather than against the data.

A better version of this essay would overlay MacAdam’s ellipses on contours drawn directly from a suprathreshold dataset. That is the right experiment and this site does not have the data for it, so what is established is that the formula disagrees with the threshold measurements in shape, which is one step removed from the underlying claim.

The step is small — a formula fitted to data inherits the data’s geometry, and if ΔE2000’s contours were the wrong shape for its own fitting set that would be a much bigger problem than this essay’s subject — but it is a step, and it is worth naming.

What the pictures cannot show

The hero figure normalises both contour families to the same size, which is the only way to compare their shapes and also throws away the thing a reader might most want: how big either of them actually is.

There is no way around that. The two are in different units — one is a chromaticity distance at a stated luminance, the other a distance in CIELAB — and the whole finding is that no constant converts between them. Drawing them at their raw scales would show one family dwarfing the other by whatever arbitrary factor the units happened to produce.

And neither contour can be shown as what it is. A just-noticeable difference is, by construction, at the edge of visibility, so a figure of MacAdam’s ellipses at true scale would show nothing at all — which is why they are drawn at ten times actual size everywhere on this site, and why the magnification is stated in every caption.

Why the substitution survives

A wrong sentence repeated for eighty years usually has something going for it, and this one has three things.

It is nearly right at one scale. Around ΔE 1, the threshold and suprathreshold answers are close enough that using either gives similar advice. The substitution fails at larger differences and in unusual parts of the space, and most practical work is neither.

There is no better single sentence. “A ΔE of 1 is roughly the difference the fitting dataset placed at one unit, under side-by-side viewing against a mid-grey background” is accurate and unusable. The field needs a short gloss and the short gloss is wrong.

Nobody is checking. A tolerance is set by finding what stops complaints, so the interpretation of the number never has to be right for the number to work. A specification calibrated empirically is immune to being justified incorrectly.

The last point is the honest one. This essay identifies a real conceptual error with modest practical consequences, in a field where the practical questions are settled by calibration rather than by theory. That is worth saying plainly rather than overstating the stakes: knowing which experiment a number came from matters most when extrapolating past where it was calibrated, which is exactly when nobody notices they are doing it.

Who found it, and when

MacAdam measured the ellipses at Kodak in 1942, using a single trained observer over many sessions, with an apparatus that let the observer adjust one field until it matched another. The result is one of the most consequential measurements in the subject and was made to answer a practical question about film.

Fechner established the threshold tradition in the 1860s, and Weber before him: the smallest detectable change is proportional to the magnitude, over a wide range. Stevens challenged the extension of threshold measurements to suprathreshold magnitude in the 1950s, arguing that magnitude grows as a power of intensity rather than logarithmically — and that threshold data cannot be integrated up to predict magnitude, which is precisely the assumption this essay is about, made in a different field first.

The colour-difference formulae developed from the 1970s onward were fitted to suprathreshold datasets — largely industrial acceptability judgements, which is a third kind of experiment again, since “would a customer accept this” is neither detection nor magnitude estimation. That the resulting formulae are quoted in threshold language is a historical accident that has survived three revisions.

Where this goes next

The formulae themselves, and how much they disagree with each other, are how far apart are two colours. The measurement this essay leans on is MacAdam measured it. And what happens when one of these numbers is written into a contract is a tolerance is a shape.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 31 that link here.

The objects this essay names

Each one links to every other essay that touches it.

ChromaticityΔEJust-noticeable differenceMacAdam's ellipsesPsychophysicsSuprathresholdThresholdTolerance