Matching and measuring

A unit rests on a space that was ranked

This collection ranks three colour spaces by how nearly they make MacAdam's ellipses circles, and CIELAB comes last. It then publishes every difference it computes in a formula built on CIELAB. Turning the same instrument on the formulae rather than the spaces shows the repair works — and that it buys roundness by paying in evenness.

Assumes No diagram makes them circles, The weighting is the disagreement and MacAdam measured it.

There is a circularity in this collection’s arrangements that took fourteen rounds to notice. It ranks colour spaces by a measurement of people, and it publishes every difference it computes in a formula built on the space that comes last.

MacAdam's twenty-five ellipses, measured in each unit. The uniformity instrument used here, applied to units rather than to spaces. The upper bar is anisotropy — the mean over the twenty-five of the largest radius divided by the smallest, where 1 would be a circle. The lower is spread — the largest mean radius divided by the smallest across all twenty-five, which asks whether a step of the same size means the same thing in different parts of the diagram. Reading down the three CIELAB-based formulae in the order they were published, the anisotropy falls 3.42 → 2.89 → 2.74 and the spread rises 3.24 → 3.59 → 4.12: the weighting divides a difference by the chroma it was measured at, which equalises directions at a point and unequalises magnitudes between points. Neither number is scaled, so no calibration is applied here. CAM16-UCS is ahead on both.
Fig. 1 MacAdam’s twenty-five ellipses measured in each unit rather than in each space. Anisotropy asks how round each ellipse is and spread asks how equal they are to one another. Reading the three CIELAB-based formulae in the order they were published, the first falls and the second rises.

The claim

The chroma weighting that makes CIEDE2000 work makes each of MacAdam’s ellipses rounder and makes the ellipses more unequal in size, and no published account mentions the second half.

  • CIELAB is last of the three spaces this collection can rank. Anisotropy 3.42 against CIELUV’s 2.37 and Oklab’s 2.27, on the collection’s own measurement.
  • The patch is real. Going ΔE*ab → ΔE*94 → ΔE2000, the anisotropy falls 3.42 → 2.89 → 2.74, so each step of weighting makes the ellipses rounder.
  • And it is paid for. Over the same three steps the spread rises 3.24 → 3.59 → 4.12, so the ellipses become more unequal in size across the diagram.
  • The published unit is not the best available. Three of the six units beat it on anisotropy, including a plain Euclidean distance in a space published in the same year as the one it repairs.
  • CAM16-UCS wins both, at 1.57 and 2.10, and by a margin larger than the gap between any two other units.

Two questions the ellipses answer

MacAdam’s 1942 measurement is twenty-five ellipses on the chromaticity diagram, each drawn round a colour, each enclosing the pairs one observer could not reliably tell apart. It is the single most reused measurement in this subject and it answers two questions that are usually run together.

Is a step of the same size the same thing in every direction? An ellipse that is a circle says yes at that point. The measure is anisotropy: the mean over the twenty-five of the largest radius divided by the smallest, where 1 would be perfect.

Is a step of the same size the same thing everywhere? Twenty-five circles all of the same radius would say yes. The measure is spread: the largest mean radius divided by the smallest, across all twenty-five.

Both are ratios, so neither depends on how a space happens to be scaled — which matters, because CIELAB, CIELUV and Oklab have wildly different natural magnitudes and a raw distance would measure nothing but that. This collection has ranked the spaces on both from the beginning, and the ranking is where the circularity begins.

What the ranking says about CIELAB

space anisotropy spread
Oklab 2.27 2.57
CIELUV 2.37 2.21
CIELAB 3.42 3.24

CIELAB is last on both, by a wide margin on anisotropy. That is not a scandal and it is not news — the 1976 committee recommended two spaces because it could not choose, and the failure of CIELAB to make the ellipses circles is why ΔE*94 and CIEDE2000 exist at all.

What is worth noticing is that the collection uses the loser for everything. Every difference here is CIEDE2000, CIEDE2000 is CIELAB with weighting functions in front of it, and the choice was inherited rather than made — the formula is what a colour-science library gives when asked for a colour difference.

The question is therefore not whether CIELAB is uniform, which it is not, but whether the repair works. And that is answerable, because the ellipses are a measurement of people and the repair is a formula, so the instrument can be turned on the formula.

Turning the instrument on the units

The change is small and the results are not. Instead of measuring the distance from an ellipse’s centre to each of its perimeter points in a space, measure it in a unit — with the formula’s weighting functions applied, exactly as they would be if the ellipse were a set of tolerance pairs.

unit anisotropy spread worst single ellipse
CAM16-UCS 1.57 2.10 2.83
ΔE*uv 2.37 2.21 5.30
Oklab 2.44 2.69 4.22
ΔE2000 2.74 4.12 7.41
ΔE*94 2.89 3.59 7.78
ΔE*ab 3.42 3.24 8.60

Three things fall out immediately.

The patch works on the question it was built for. ΔE*ab is exactly the CIELAB row, as it must be — the Euclidean distance in a space and the space are the same instrument. Adding ΔE*94’s divisors takes the anisotropy from 3.42 to 2.89, and CIEDE2000’s more elaborate machinery takes it to 2.74. Twenty per cent of the anisotropy removed, on the data the corrections were fitted to.

And it does not close the gap. 2.74 is still worse than a plain Euclidean distance in CIELUV, which does no weighting at all and was published in the same year, by the same committee, as an alternative nobody took.

CAM16-UCS is somewhere else entirely. 1.57 against a field running 2.37 to 3.42. The gap between it and the next-best unit is larger than the gap between the next-best and the worst.

One of MacAdam's ellipses as each unit sees it, at x 0.305, y 0.323. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 1.42; the least round is ΔE*94 at 1.90. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 2 One of the twenty-five, drawn as each unit sees it: the distance from the centre to each perimeter point, scaled to its own mean so the shapes can be compared. A perfectly uniform unit would draw a circle. What the units agree about is which direction the ellipse is long in; what they disagree about is how long.
One of MacAdam's ellipses as each unit sees it, at x 0.131, y 0.521. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 1.27; the least round is ΔE*ab at 3.19. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 3 A second ellipse, at a different place on the diagram. The ordering of the units is the same and the magnitudes are not, which is what the spread column measures.

The cost nobody mentions

The anisotropy column is the one every account of CIEDE2000 discusses. The spread column runs the other way and is not discussed anywhere this collection can find.

Over the three CIELAB-based formulae in publication order, spread rises 3.24 → 3.59 → 4.12. Each step of chroma weighting makes the twenty-five ellipses more unequal in size than the unweighted distance made them.

The mechanism is not subtle once stated. The weighting divides a difference by the chroma at which it was measured. Dividing by chroma equalises the two axes at a point — which is exactly the anisotropy question, and exactly what it improves. It also shrinks every difference in the saturated part of the diagram relative to every difference in the pale part, which unequalises the magnitudes between points — which is exactly the spread question.

The two are not independent quantities that happen to move oppositely. One operation improves the first by doing the thing that worsens the second.

Whether that is a good trade depends on what a unit is for, and there is a defensible answer: a colour-difference formula is most often used to decide whether one pair is inside a tolerance and another outside, both pairs being at roughly the same place in the space, and that is the anisotropy question. Comparing tolerances across the diagram — is a one-unit error on a pale grey the same problem as a one-unit error on a saturated red — is the spread question, and it is asked less often, mostly because the answer is known to be no.

But the trade is not stated in the recommendation, and a reader who takes “more uniform” as one property is getting one property improved and another degraded.

MacAdam's discrimination ellipses, drawn 10 times actual size. Twenty-five ellipses of colours indistinguishable from their centres. They are drawn at 10× because at true scale most are thinner than a line. Their areas vary by a factor of 74, which is the whole result: a step of the same size in xy means very different things in different places.
Fig. 4 MacAdam’s ellipses on the chromaticity diagram, magnified. The twenty-five differ in size by a factor of about twenty in the raw diagram, and every uniformity measure in this essay is an attempt to say how much of that a coordinate change removes.

The circularity, stated plainly

Assembling the pieces gives an uncomfortable sentence about this collection’s own arrangements.

It measures how uniform a space is, using MacAdam’s ellipses. It reports CIELAB last. It then computes every difference in this collection using a formula built on CIELAB, including the differences in the essays that report CIELAB as last. And when it ranks colour spaces by any other criterion, the ranking is computed in a distance defined in one of the spaces being ranked.

The circularity is not vicious — the uniformity measures above are ratios of distances within one space, so no cross-space comparison is smuggled in — but it is real, and it has a consequence that can be measured rather than argued about. Recomputing this collection’s whole inventory of published quantities in the unit that wins the ellipse test moves five of the six by between 15 and 84 per cent, and reorders two of the census’s fourteen rows.

The honest response is not to change unit. It is to say which unit, everywhere, which is the convention the figures in this round adopt and which costs one line of caption.

The two criteria together, which the repair does not survive

The essay reports the trade and does not price it. Both columns are ratios with 1 as perfection, so combining them is legitimate, and any symmetric combination gives the same answer.

unit anisotropy spread product sum
CAM16-UCS 1.57 2.10 3.30 3.67
ΔE*uv 2.37 2.21 5.24 4.58
Oklab 2.44 2.69 6.56 5.13
ΔE*94 2.89 3.59 10.38 6.48
ΔE*ab 3.42 3.24 11.08 6.66
ΔE2000 2.74 4.12 11.29 6.86

Judged on both criteria at once, CIEDE2000 is worse than the unweighted CIELAB distance it was built to repair — 11.29 against 11.08 on the product, 6.86 against 6.66 on the sum. Twenty-four years of committee work on this data, and the pair of numbers it was working to improve is two per cent further from perfection than it started.

And the best of the three is the intermediate one. ΔE*94 is six per cent better than plain CIELAB on the product, because its milder weighting buys most of the anisotropy improvement — 3.42 down to 2.89, against 2.74 for the full formula — while giving away much less spread. The second step of the repair, from ΔE*94 to CIEDE2000, buys 0.15 of anisotropy and costs 0.53 of spread, and no weighting of the two criteria that treats them as comparable makes that a gain.

None of which says CIEDE2000 is a bad formula. It says that the two criteria are not what it was optimised against: it was fitted to observer judgements of pairs, which is overwhelmingly a within-place comparison, and it does very well at that. What the table shows is the price of an objective, and the price is that the one quantity nobody measured went backwards far enough to cancel the gain in the one everybody did.

The trade runs at worse than one for one

Sizing the two movements against each other, over the three CIELAB-family units in publication order: anisotropy falls by 19.9 per cent and spread rises by 27.2.

Those are different quantities and the comparison is not a physical exchange rate. It is still the right order-of-magnitude statement to have, because the essay’s own framing — it is paid for — leaves open whether the payment is a rounding or the whole gain. It is more than the whole gain, in proportional terms, which is what the product column says in a second way.

The repair also lengthens the tail

The last column carries one more effect, and it is not the same as the worst-case observation the essay makes.

Dividing each unit’s worst ellipse by its mean gives how far the tail sits above the bulk:

unit worst / mean
Oklab 1.73
CAM16-UCS 1.80
ΔE*uv 2.24
ΔE*ab 2.51
ΔE*94 2.69
ΔE2000 2.70

The two weighted CIELAB units have the most skewed distributions of the six. Weighting does not merely improve the worst ellipse less than the mean — it makes the worst ellipse relatively worse than it was, taking the ratio from 2.51 to 2.70. So the observation that the ellipses it helps least are the ones that most need help is stronger than the 14-against-20 comparison suggests: the repair concentrates the remaining anisotropy rather than spreading it.

Oklab is the mirror image and is the reason the column is worth printing. Its mean anisotropy is fourth of six and its tail is the shortest of all, at 1.73 — a unit that is mediocre everywhere and bad nowhere, against the CIELAB family’s pattern of being decent in most places and poor in a few. Which of those two is preferable is a question about what a unit is for, and it is invisible in either column taken alone.

And two figures for the same inventory disagree

One small discrepancy worth naming, since the whole essay is about numbers that come from different places. The circularity section says recomputing the collection’s inventory in the winning unit moves five of six quantities by between 15 and 84 per cent; the figure caption at the foot of the essay says five of them move by more than 80 per cent.

Those cannot both describe the same five. A range starting at 15 contains at least one quantity well below 80, so the caption’s more than 80 per cent holds for fewer than five of them. The sentence in the text is the one that admits a spread and is the safer of the two to quote.

What the worst ellipse says

The last column of the table is the single worst ellipse under each unit, and it is worth separating from the mean because the two behave differently.

Under ΔE*ab the worst ellipse has an anisotropy of 8.60 — one direction through it is eight and a half times harder to see than the perpendicular one. Under ΔE2000 it is 7.41, under Oklab 4.22, and under CAM16-UCS 2.83. So the weighting improves the worst case by 14 per cent while improving the mean by 20, which means the ellipses it helps least are the ones that most need help.

Which ellipse it is depends on the unit, and that is informative in itself. For the two unweighted CIELAB-family units the worst is at (0.187, 0.118) — a saturated blue-violet, the region every account of CIELAB names as its weakest. For the three weighted or modelled units the worst moves to (0.160, 0.200) — a saturated blue-green, further round the diagram. The repair moves the problem rather than removing it, which is a shape this collection meets whenever a fit has fewer parameters than the thing it is fitting.

Oklab’s worst is somewhere else again, at (0.253, 0.125), and its worst case is the second best of the six despite its mean being fourth. A unit with a good worst case and a middling mean is a different kind of object from one with the reverse, and this collection has an essay on why the two rankings tend to invert.

Why not simply switch

Three reasons, and the third is the one that decides it.

The ellipses are twenty-five points at one luminance. They are a measurement of two observers in 1942 at a fixed luminance on a fixed surround, and a unit is asked to work over the whole space at every level. A unit that wins on this data may be over-fitted to it — and CAM16-UCS’s ancestor was fitted partly to it, so its winning is less independent than it looks.

An appearance unit needs a room. CAM16-UCS takes an adapting luminance, a background, a surround and a degree of adaptation, and every one of those is a further choice. Swapping a unit with no arguments for one with four is not obviously an improvement in a collection whose whole argument is that unstated arguments are the problem. The surround alone is three tabulated rooms standing in for a continuum.

And it would break every published number at once. Two hundred and eighty essays quote differences in ΔE2000. A change of unit is not an edit, it is a recomputation, and the useful thing to publish is not a second set of numbers but the exposure of the first set — which is what this round does.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 5 The same six units on a completely different instrument: how far each is from a rescaling of the published one. The two rankings are nearly reversed, so agreement with the convention and quality on the ellipses are close to opposite properties.
Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu.
Fig. 6 Where on the scale the units disagree with the published one. The ellipses are entirely at the near end of this chart, which is where the disagreement is proportionally largest — so a ranking on the ellipses and a ranking by agreement with convention are being taken at opposite ends of one curve.

One ellipse at a time is the finest grain this comparison has, and it is where a ranking of units stops being stable.

One of MacAdam's ellipses as each unit sees it, at x 0.475, y 0.300. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔEok at an anisotropy of 1.23; the least round is ΔE*94 at 3.08. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 7 A single measured ellipse in each of the six units, each scaled to its own mean radius. Which unit is least anisotropic here need not be the one that wins on the average, which is what makes a ranking of units a statement about a set.

The instrument turned on itself

There is one more measurement available here and it is uncomfortable enough to be worth taking.

MacAdam’s ellipses are the collection’s uniformity instrument. They are also twenty-five measurements by one observer, made in 1942, at one luminance, on a bipartite field in a dark surround — which is a viewing condition none of the units on the menu is normally used in. The instrument has all the properties this collection spends its time finding in other people’s instruments.

Two of them are quantifiable from what is already here. The ellipses’ own sampling is thin: twenty-five points, unevenly distributed, with a large gap in the purples and three points crowded near the green corner, so the mean anisotropy is a mean over a set nobody designed and its standard error is not small. And their luminance is a free parameter of the reanalysis rather than of the measurement — MacAdam published chromaticities, this collection realises them at Y = 0.4, and moving that changes every number in the table by a few per cent while changing no ordering.

Neither undermines the finding, because the finding is a comparison between units measured on one instrument and a common instrument’s own flaws are common-mode. What they do is bound how much the numbers deserve: the ordering of the six units is a result and the size of the gaps is not.

Where the model stops

The ellipse instrument measures one thing well and says nothing about the rest of the space. There is no equivalent dataset for lightness differences, for large differences, or for differences under a coloured adapting light, and a unit that is excellent on chromaticity discrimination at one luminance may be poor at any of those.

The anisotropy and spread numbers also depend on the luminance the ellipses are realised at, which is a free parameter of the measurement — MacAdam’s data are chromaticities without a stated luminance factor, and this collection realises them at Y = 0.4. Moving it changes every number here by a few per cent and changes no ordering.

And the twenty-five ellipses are themselves a sample of the diagram, taken by one observer, and their own uncertainty is not small. A unit ranking with error bars on it would be the right thing to publish and would need the raw trial data, which does not survive.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 8 What the circularity costs, measured: six published quantities across the whole menu, five of which move by more than 80 per cent between the unit this collection publishes in and the unit that wins the ellipse test.

Who found it, and when

MacAdam’s measurement is 1942 and its use as a uniformity criterion is immediate and universal. That CIELAB does not satisfy it was said by the committee that recommended CIELAB.

The chroma weighting is the CIE’s, 1994 and 2000, and the anisotropy improvement it buys is exactly what those recommendations are validated on. The spread degradation is not, and the reason is that spread is not usually computed: the standard validation is a correlation between formula and observer judgement over a dataset of pairs, which is dominated by within-place comparisons and does not test between-place ones. A measure that is not computed cannot be reported as worsening.

Naming it here costs nothing and is the sort of thing a collection with both numbers in one program is placed to notice. It is an advantage of arrangement, again, and not of insight.

Where the ladder goes next

A tolerance is the place where a unit stops being a description and becomes a decision. A delivery either passes or it does not, and the boundary between the two is drawn through the space of pairs by whichever formula the contract names — so two contracts quoting the same number in different units accept different sets of deliveries, and the overlap between them is measurable.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 14 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyAppearance modelCalibrationChromaCIEDE2000Colour differenceColour spaceJust-noticeable differenceMacAdam's ellipsesUniformity