Difference and uniformity

Twenty-five is a sample of the diagram

A colour space's uniformity score is the mean of twenty-five numbers, and a mean of twenty-five numbers has a standard error those twenty-five numbers determine. Nothing has to be quoted to compute it, and three of the seven adjacent pairs in this collection's ranking survive it.

Assumes How wrong would the data have to be, No diagram makes them circles and An extremum is not a sample.

The score a colour space gets here is a mean, and a mean carries an error whether or not anybody computes it.

Twenty-five ellipses is a sample, and the score has an error bar. One row per colour space this collection ranks: the mean axis ratio its ellipses come out at, with the standard error of that mean over the twenty-five ellipses it was computed from. No literature is quoted — a mean of twenty-five numbers has a standard error those twenty-five numbers determine. The bars are far from equal: the best space carries ± 0.07 and the worst ± 1.56, because a space that makes the ellipses nearly circular makes all of them nearly circular and one that does not is dominated by whichever ellipse it handles worst.
Fig. 1 Every basis in this collection’s uniformity table, with the standard error of its score over the twenty-five ellipses it was computed from. Nothing external is quoted: a mean of twenty-five numbers has an error those numbers determine.

The claim

There are two uncertainties in this ruler and only one of them needs a literature. The other is computable from the data in hand, it is the larger of the two, and it says that most of this collection’s uniformity ranking is not a ranking.

  • The score is a mean. How uniform a space is, here, is the average axis ratio of twenty-five mapped ellipses — so its uncertainty as an estimate of the diagram is a sampling error over twenty-five draws.
  • Three of the seven adjacent pairs survive a bootstrap. The other four do not, and one of them changes places in thirty-five per cent of resamples.
  • The comparison has to be paired, and getting that wrong would have declared almost the whole table unordered. The same twenty-five ellipses score both spaces, so what matters is the error on the difference, and an ellipse that is hard for everybody cancels.
  • The two instruments disagree about four pairs. An error on each ellipse leaves them safe; a resample of which ellipses were measured does not. The twenty-five-ness of the ruler matters more than the accuracy of any one of them.
  • And the instrument is most precise where the answer is smallest — a log–log correlation of 0.96 between a basis’s score and its standard error — which is the opposite of the arrangement the previous round found in the sampled estimator.

The uncertainty that quotes nobody

The other uncertainty is an error on each ellipse’s fitted axes, and it cannot be computed here: it is a fact about an experiment run in 1942. What that essay does is invert the question and report how large the error would have to be.

This one needs no inversion. Twenty-five ellipses is a sample of the chromaticity diagram, the score is their mean, and the sample’s own scatter is a complete description of how well that mean estimates what a twenty-sixth ellipse would have said.

The two answer genuinely different questions, and the difference is worth stating carefully. The first asks how well were these twenty-five measured. The second asks how much would the answer move if a different twenty-five had been chosen. Neither subsumes the other, and a table with one error bar on it would be wrong in one of the two directions.

Why the comparison must be paired

The first attempt at this gives a wrong and discouraging answer, and it is worth showing because the correction is the interesting part.

Each basis’s score carries a standard error: 0.07 for the best discrimination basis, 0.22 for CAT16 and the receptor construction, 0.66 for Bradford, 1.06 for CAT02, 1.56 for the adaptation optimum. Comparing two bases by adding their two errors in quadrature would give an error on the gap of about 0.31 for the CAT16-against-receptors pair, whose gap is 0.11 — so the pair reads as hopeless, and so does nearly every other.

That comparison is wrong because the two scores are not independent. The same twenty-five ellipses score both, and an ellipse that is hard for every basis inflates both means together. What matters is the standard error of the per-ellipse difference, and it is much smaller: 0.090 for that pair, against a gap of 0.110.

Three of the seven adjacent pairs in the table are actually ordered. One bar per adjacent pair of the uniformity ranking: the difference between the two spaces' scores divided by the standard error of that difference over the twenty-five ellipses. The comparison is paired — the same ellipses score both spaces, so an ellipse that is hard for everybody cancels — which is why the table is more informative than it looks and why treating the two errors as independent would have declared almost nothing ordered. 3 of the 7 pairs clear two; the rest do not, and one of them crosses zero in 35 per cent of resamples. The ranking's ends are real and its middle is not a ranking.
Fig. 2 Each adjacent pair as its gap divided by the standard error of that gap, computed from the per-ellipse differences. The pairing is what makes the comparison informative rather than hopeless.

This is the paired comparison every experimental design textbook is about, arriving in a place nobody would have looked for it, and the size of the correction is a factor of three or four on every pair. A version of this analysis that had treated the two scores as independent would have concluded that the entire uniformity table means nothing — a dramatic result, easy to publish, and an artefact of a missing correlation.

There is a third comparison available and it is instructive that it is worse than both. Comparing two bases by their ranks rather than their scores — asking how often a resample puts one above the other, and nothing else — throws away the size of the gap and reports a probability that says nothing about whether the difference matters. A ranking is not a measurement, and converting one into the other loses the only part with units on it.

The answers

  • The best discrimination basis against the receptor construction: a gap of 0.985 at 4.8 standard errors, and no crossing in four hundred resamples. Ordered.
  • The receptors against CAT16: 0.110 at 1.2, crossing in 9.2 per cent. Not ordered.
  • CAT16 against Hunt–Pointer–Estévez: 0.106 at 1.6, crossing in 5.5 per cent. Not ordered.
  • Hunt–Pointer–Estévez against XYZ scaling: 0.630 at 1.6, crossing in 4.8 per cent. Ordered, barely.
  • XYZ scaling against Bradford: 0.237 at 0.4, crossing in 35 per cent. Not ordered, and not close.
  • Bradford against CAT02: 0.615 at 1.0, crossing in 17 per cent. Not ordered.
  • CAT02 against the adaptation optimum: 3.402 at 2.0, crossing in 2 per cent. Ordered.

Read as a whole the table says something more useful than any of its rows. The seven gaps span a factor of thirty, from 0.106 to 3.402, and the seven standard errors span a factor of twenty-five. They are not aligned: the largest gap is not the most significant and the smallest error is not on the safest pair. What decides a pair is the ratio, and nothing about a table sorted by score makes that ratio visible.

Three of seven. The two ends of the table are ordered and its middle is a cloud of four bases whose relative positions the data do not settle — which happens to be exactly what the other instrument said about the same middle, arriving from an entirely different direction.

The XYZ-against-Bradford row is the striking one. Their scores are 3.44 and 3.68, a gap of seven per cent, and they cross in more than a third of resamples. That is not because the gap is small; it is because those two bases disagree violently about individual ellipses, so the per-ellipse differences have a standard deviation of 3.2 against a mean of 0.24.

Where the two instruments disagree

Four of the seven pairs are called differently by the two uncertainties, and all four go the same way: the sampling error says they are unsettled and an error on the ellipses says they are safe.

How wrong the ellipses would have to be for a pair to change places. One bar per adjacent pair in the uniformity table: the relative error on each ellipse's own axes at which that pair changes places in one draw in twenty. No error on the data is quoted anywhere — the question is inverted, so what is reported is how large an error would have to be, and a reader with an opinion about MacAdam's experiment can compare it with their own number. The nearest pair goes at 0.171; 2 of the 7 pairs do not reverse under any error this search covers.
Fig. 3 How large an error on each ellipse’s own axes would have to be to reverse each pair. Every one of these thresholds is above ten per cent, and four of the pairs are nonetheless unsettled by the sampling error.

The other sampling error in the same machinery is inside each ellipse rather than between them, and it is worth knowing which of the two dominates.

How many points a ratio needs is set by the ratio. Twelve curves, one per coordinate system this collection draws, each showing the axis ratio a sample of n points reports as a fraction of the exact value, against n on a logarithmic axis. Every curve rises to one, and they reach it at wildly different places: the systems whose ellipses are nearly circles are right at twenty-four points, and the CIE RGB primaries — whose worst ellipse has an axis ratio of 9 — are still 32 per cent short at forty-eight. The curve a measurement is on cannot be known until the answer is, which is what makes a fixed sample size the wrong instrument for this question.
Fig. 4 What a ring of n points reports as a fraction of the true axis ratio. This is the error inside one contour; the error bars above are the error across twenty-five of them, and the second is the larger.
Three numbers for one set of ellipses, and which of them is which. Two curves and a horizontal line, against the size the ellipses are drawn at. The line is the analytic axis ratio — the ratio of the singular values of the map's own derivative, which is what "does this space make discrimination contours circles" means. The upper curve is a very finely sampled ring, which sits 0.6 per cent above the line at full size and converges onto it as the ellipse shrinks, because the gap between them is the second-order distortion of the map across a real ellipse rather than an error. The lower curve is the forty-eight-point sample used for this until now: it does not converge onto anything, because its error is set by the sample and not by the size.
Fig. 5 And the three estimates of one contour that the whole comparison rests on. A ranking whose ordering survives both of these sampling errors is a ranking of the spaces rather than of the arithmetic.

The mechanism is the one from the correlated-error result in the neighbouring essay, seen from the other side. Perturbing every ellipse’s axes moves both bases in a pair together, because both are computed from the same perturbed ellipses; resampling which ellipses were measured moves them apart, because the two bases disagree about which ellipses are difficult.

So the two instruments are sensitive to opposite things, and the honest summary is that the composition of the set matters more than the accuracy of its members. If MacAdam had worked at twenty-five different chromaticities, chosen the same way, four of these seven pairs could plausibly have come out the other way round. If he had measured the same twenty-five twice as accurately, none of them would have moved.

That is an uncomfortable result for a table this collection uses, and it is a reassuring one for the argument the table is used to make, which is about its ends.

There is one more thing to read off the disagreement, and it is a caution about the neighbouring essay’s numbers rather than about these. A threshold of the form the ellipses would have to be seventeen per cent wrong is computed with the set held fixed. It answers a question conditional on these being the twenty-five, and the conditioning is doing more work than it looks — because if the set had been different, the pair might not have needed an error at all.

Two criteria, and the pair they disagree about

The three ordered pairs are the ones whose bootstrap crossing rate falls below five per cent, and it is worth saying that rather than the shorter thing, because the obvious alternative criterion does not give the same three.

The gap in standard errors, pair by pair down the ranking: 4.79, 1.23, 1.60, 1.62, 0.37, 1.03 and 2.03. Only two of the seven clear two standard errors. The bootstrap admits three, because the pair at 1.62 crosses in 4.8 per cent of resamples and the one at 1.60 crosses in 5.5 — two gaps of the same size, on opposite sides of a threshold, separated by seven parts in a thousand of resampling noise.

That is not a defect in either instrument; it is the finding restated one level down. A pair whose ordering depends on which of two reasonable criteria is applied is a pair that is not ordered, whatever either criterion says about it, and reading the table as though the boundary were sharp would be the same mistake as reading the ranking as though it were a ranking. The three that survive are the ones the essay’s argument uses, and they survive both rules with room to spare — 4.79 against 2.03 against everything else below 1.7.

How many ellipses would have settled the rest

The sampling error is a standard error of a mean over twenty-five draws, so it falls as the square root of the count, and the count that would settle each unsettled pair follows directly from the gap it already has.

pair gap in standard errors ellipses needed
the confusion-point basis against CAT16 1.23 67
Bradford against CAT02 1.03 95
XYZ against Bradford 0.37 742

Two of the three are within reach of an experiment somebody could actually run: sixty-seven ellipses is under three times what MacAdam measured, and ninety-five is under four. The third is not. Separating XYZ from Bradford on this ruler would need seven hundred and forty-two ellipses, thirty times the existing set, and at that point the limiting uncertainty would no longer be the sample — it would be the accuracy of each ellipse, which no amount of resampling improves and which this collection cannot compute at all.

So the honest statement about that pair is not that its order is unknown pending better data. It is that the two spaces are the same distance from uniform as far as this instrument can ever say, and a table printing them one above the other is printing an arbitrary choice with the appearance of a measurement.

Precision follows the score

One property of the instrument falls out of the same numbers and is worth having, because it is the opposite of a defect the previous round found.

A basis that scores better is also scored more precisely. The log–log correlation between a basis’s mean axis ratio and the standard error of that mean is 0.961 across the eight. The best discrimination basis is 1.61 ± 0.07; the adaptation optimum is 7.70 ± 1.56.

The mechanism is not deep and is worth stating anyway: a space that scores well does so by making every ellipse nearly circular, so its twenty-five per-ellipse ratios are all near one and have little spread. A space that scores badly is usually dominated by whichever ellipses it handles worst, so its twenty-five span an order of magnitude.

This is the reverse of the arrangement that made the sampled estimator dangerous. That estimator was accurate where the answer was boring and wrong where it mattered — biased by an amount that grew with the quantity being measured. This one is most precise where the answer is smallest, which is where the interesting comparisons are: the floors, the optima, and the distance between them.

What was computed, and how

The per-ellipse anisotropies come from the closed form — the singular values of the map’s own derivative composed with the ellipse’s shape matrix — rather than from walking round a ring. That matters here for a reason beyond accuracy: the closed form is exact per ellipse, so the only scatter in the twenty-five numbers is the scatter of the diagram itself, with none of the estimator’s own noise added.

The standard errors are computed twice, by the textbook formula and by a bootstrap over four hundred resamples with replacement, and the two are reported side by side. They agree to within about five per cent for the well-behaved bases and differ by more for the badly behaved ones, which is exactly where they should differ: the per-ellipse ratios are strongly skewed under a poor basis — one ellipse at forty and the rest near three — and the textbook formula assumes a symmetry that skew does not have.

The crossing rates are bootstrap quantities as well: resample the twenty-five per-ellipse differences with replacement, average, and count how often the average changes sign. That is the quantity a pair’s ordering depends on, and it needs no assumption of normality — which matters, because the differences are as skewed as the ratios they come from.

Where the model stops

A bootstrap over twenty-five points is a bootstrap over twenty-five points. It estimates the sampling distribution of a mean from a sample small enough that its tails are guesswork, and the crossing rates near five per cent should be read as near the line rather than as numbers.

The twenty-five are not a random sample of anything, which is the sharpest limitation and the one that cannot be repaired from inside. MacAdam chose where to work, and he chose sensibly — spread across the diagram, more densely where discrimination changes fastest. A bootstrap treats them as exchangeable draws from a population of chromaticities, which they are not. What the calculation really estimates is the variability of a mean over sets like this one, and that is a slightly different and less well-defined thing.

A resample is not a redesign. The bootstrap asks what the mean would have been over a different draw of twenty-five from this set, which is a proxy for a different set of twenty-five stimuli and is not the same thing. A genuinely different experiment would choose its chromaticities by some rule — evenly in some coordinate system, or concentrated where an application cares — and the mean over such a set is not the mean this bootstrap is estimating the error of. The distinction is invisible when the sample is large and is not invisible at twenty-five.

And a score is not a preference. A space that makes MacAdam’s ellipses round is uniform at those twenty-five chromaticities, at that luminance, for that observer, at threshold. Every one of those qualifications is load-bearing, and suprathreshold differences are not scaled threshold ones — which is a larger caveat on the whole exercise than any error bar computed here.

The generalisation

The result generalises into a rule that is easy to state and is skipped constantly, including in this collection until now.

Before reading a ranking, compare its adjacent gaps with the standard error of the quantity being ranked. A table sorted to three decimals invites an ordering; whether it has one depends on a number that is usually not printed beside it, and often not computed at all.

The second half of the rule is the one that took work here. When the same instrument scores every entry, the comparison is paired and the error on a gap is far smaller than the errors on the entries. Getting this backwards is a much more common failure than omitting the error entirely, because it produces a defensible-looking calculation with a dramatic conclusion — nothing in this table is distinguishable — that is wrong by a factor of three or four.

The third part is the one this essay would not have predicted: when there are two candidate uncertainties, check whether they order the same pairs, because a disagreement between them is more informative than either alone. Here it identifies precisely which property of the ruler the conclusions are exposed to — its composition, not its accuracy — and that is the sort of thing a single error bar can never say.

What survives

The four essays that argue from this table were checked against the result rather than assumed to be safe, and all four survive, for the same reason each time: they use the ends.

The best axes are not receptors compares the two optima with everything between them, and the optima are separated from the field by 4.8 and 2.0 standard errors. No basis is good at both is about a factor of five between two ends of a diagonal. A constraint costs what it points at uses distances from the floors, and a floor is the most precisely estimated quantity in the set.

What does not survive is any sentence of the form this published transform makes the ellipses rounder than that one. Four such comparisons are available in the middle of the table and none of them is a fact about the transforms. None was made, which is luck rather than judgement, and is now checked rather than lucky.

Who found it, and when

The paired comparison is Student’s and Fisher’s, and predates every number in this essay by a century. The bootstrap is Efron’s, from 1979, and is exactly the tool for a statistic — a mean of skewed ratios — whose sampling distribution nobody wants to write down.

What is not standard is applying either to a canonical dataset that has become an illustration. MacAdam’s ellipses are reprinted as a diagram far more often than they are used as a sample, and a diagram does not have a standard error. The move here is only to remember that the twenty-five points on the picture were twenty-five experiments, and that a mean over them is a statistic like any other.

Where the ladder goes next

Two uncertainties have now been put on this ruler and they disagree, which raises the question of what else in this collection has an uncertainty that has never been asked for. The answer is: a great deal, and it is worth doing systematically rather than one measurement at a time.

The audit that does it works through every published claim here that has a threshold attached, and asks the smallest thing that would have to be wrong for each to stop holding. The most exposed claim on this site is not the one this essay’s neighbours would have nominated.

The points a ratio needs are proportional to the ratio. A scatter of 133 points on logarithmic axes, one per MacAdam ellipse under each of six coordinate systems. The horizontal position is that ellipse's true axis ratio; the vertical is the smallest sample size, from a sequence of doublings, at which the sampled ratio comes within one per cent and stays there. A line of slope 0.94 runs through them, against a predicted 1 — the minimum's notch is 0.88 σ₂/σ₁ radians wide, so resolving it takes a number of points proportional to σ₁/σ₂, and nothing about the basis or the ellipse enters beyond that. An ellipse with a ratio of two needs seventeen points and one with a ratio of twenty-six needs a hundred and ninety-two.
Fig. 6 The sample size an axis ratio needs, plotted against the ratio itself. A different sampling question about the same object, and one with an exact answer.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyCIELABColour differenceConvergenceDegrees of freedomMacAdam's ellipsesQuadratic formSampling