Twenty-five is a sample of the diagram
Assumes How wrong would the data have to be, No diagram makes them circles and An extremum is not a sample.
The score a colour space gets here is a mean, and a mean carries an error whether or not anybody computes it.
The claim
There are two uncertainties in this ruler and only one of them needs a literature. The other is computable from the data in hand, it is the larger of the two, and it says that most of this collection’s uniformity ranking is not a ranking.
- The score is a mean. How uniform a space is, here, is the average axis ratio of twenty-five mapped ellipses — so its uncertainty as an estimate of the diagram is a sampling error over twenty-five draws.
- Three of the seven adjacent pairs survive a bootstrap. The other four do not, and one of them changes places in thirty-five per cent of resamples.
- The comparison has to be paired, and getting that wrong would have declared almost the whole table unordered. The same twenty-five ellipses score both spaces, so what matters is the error on the difference, and an ellipse that is hard for everybody cancels.
- The two instruments disagree about four pairs. An error on each ellipse leaves them safe; a resample of which ellipses were measured does not. The twenty-five-ness of the ruler matters more than the accuracy of any one of them.
- And the instrument is most precise where the answer is smallest — a log–log correlation of 0.96 between a basis’s score and its standard error — which is the opposite of the arrangement the previous round found in the sampled estimator.
The uncertainty that quotes nobody
The other uncertainty is an error on each ellipse’s fitted axes, and it cannot be computed here: it is a fact about an experiment run in 1942. What that essay does is invert the question and report how large the error would have to be.
This one needs no inversion. Twenty-five ellipses is a sample of the chromaticity diagram, the score is their mean, and the sample’s own scatter is a complete description of how well that mean estimates what a twenty-sixth ellipse would have said.
The two answer genuinely different questions, and the difference is worth stating carefully. The first asks how well were these twenty-five measured. The second asks how much would the answer move if a different twenty-five had been chosen. Neither subsumes the other, and a table with one error bar on it would be wrong in one of the two directions.
Why the comparison must be paired
The first attempt at this gives a wrong and discouraging answer, and it is worth showing because the correction is the interesting part.
Each basis’s score carries a standard error: 0.07 for the best discrimination basis, 0.22 for CAT16 and the receptor construction, 0.66 for Bradford, 1.06 for CAT02, 1.56 for the adaptation optimum. Comparing two bases by adding their two errors in quadrature would give an error on the gap of about 0.31 for the CAT16-against-receptors pair, whose gap is 0.11 — so the pair reads as hopeless, and so does nearly every other.
That comparison is wrong because the two scores are not independent. The same twenty-five ellipses score both, and an ellipse that is hard for every basis inflates both means together. What matters is the standard error of the per-ellipse difference, and it is much smaller: 0.090 for that pair, against a gap of 0.110.
This is the paired comparison every experimental design textbook is about, arriving in a place nobody would have looked for it, and the size of the correction is a factor of three or four on every pair. A version of this analysis that had treated the two scores as independent would have concluded that the entire uniformity table means nothing — a dramatic result, easy to publish, and an artefact of a missing correlation.
There is a third comparison available and it is instructive that it is worse than both. Comparing two bases by their ranks rather than their scores — asking how often a resample puts one above the other, and nothing else — throws away the size of the gap and reports a probability that says nothing about whether the difference matters. A ranking is not a measurement, and converting one into the other loses the only part with units on it.
The answers
- The best discrimination basis against the receptor construction: a gap of 0.985 at 4.8 standard errors, and no crossing in four hundred resamples. Ordered.
- The receptors against CAT16: 0.110 at 1.2, crossing in 9.2 per cent. Not ordered.
- CAT16 against Hunt–Pointer–Estévez: 0.106 at 1.6, crossing in 5.5 per cent. Not ordered.
- Hunt–Pointer–Estévez against XYZ scaling: 0.630 at 1.6, crossing in 4.8 per cent. Ordered, barely.
- XYZ scaling against Bradford: 0.237 at 0.4, crossing in 35 per cent. Not ordered, and not close.
- Bradford against CAT02: 0.615 at 1.0, crossing in 17 per cent. Not ordered.
- CAT02 against the adaptation optimum: 3.402 at 2.0, crossing in 2 per cent. Ordered.
Read as a whole the table says something more useful than any of its rows. The seven gaps span a factor of thirty, from 0.106 to 3.402, and the seven standard errors span a factor of twenty-five. They are not aligned: the largest gap is not the most significant and the smallest error is not on the safest pair. What decides a pair is the ratio, and nothing about a table sorted by score makes that ratio visible.
Three of seven. The two ends of the table are ordered and its middle is a cloud of four bases whose relative positions the data do not settle — which happens to be exactly what the other instrument said about the same middle, arriving from an entirely different direction.
The XYZ-against-Bradford row is the striking one. Their scores are 3.44 and 3.68, a gap of seven per cent, and they cross in more than a third of resamples. That is not because the gap is small; it is because those two bases disagree violently about individual ellipses, so the per-ellipse differences have a standard deviation of 3.2 against a mean of 0.24.
Where the two instruments disagree
Four of the seven pairs are called differently by the two uncertainties, and all four go the same way: the sampling error says they are unsettled and an error on the ellipses says they are safe.
The other sampling error in the same machinery is inside each ellipse rather than between them, and it is worth knowing which of the two dominates.
The mechanism is the one from the correlated-error result in the neighbouring essay, seen from the other side. Perturbing every ellipse’s axes moves both bases in a pair together, because both are computed from the same perturbed ellipses; resampling which ellipses were measured moves them apart, because the two bases disagree about which ellipses are difficult.
So the two instruments are sensitive to opposite things, and the honest summary is that the composition of the set matters more than the accuracy of its members. If MacAdam had worked at twenty-five different chromaticities, chosen the same way, four of these seven pairs could plausibly have come out the other way round. If he had measured the same twenty-five twice as accurately, none of them would have moved.
That is an uncomfortable result for a table this collection uses, and it is a reassuring one for the argument the table is used to make, which is about its ends.
There is one more thing to read off the disagreement, and it is a caution about the neighbouring essay’s numbers rather than about these. A threshold of the form the ellipses would have to be seventeen per cent wrong is computed with the set held fixed. It answers a question conditional on these being the twenty-five, and the conditioning is doing more work than it looks — because if the set had been different, the pair might not have needed an error at all.
Two criteria, and the pair they disagree about
The three ordered pairs are the ones whose bootstrap crossing rate falls below five per cent, and it is worth saying that rather than the shorter thing, because the obvious alternative criterion does not give the same three.
The gap in standard errors, pair by pair down the ranking: 4.79, 1.23, 1.60, 1.62, 0.37, 1.03 and 2.03. Only two of the seven clear two standard errors. The bootstrap admits three, because the pair at 1.62 crosses in 4.8 per cent of resamples and the one at 1.60 crosses in 5.5 — two gaps of the same size, on opposite sides of a threshold, separated by seven parts in a thousand of resampling noise.
That is not a defect in either instrument; it is the finding restated one level down. A pair whose ordering depends on which of two reasonable criteria is applied is a pair that is not ordered, whatever either criterion says about it, and reading the table as though the boundary were sharp would be the same mistake as reading the ranking as though it were a ranking. The three that survive are the ones the essay’s argument uses, and they survive both rules with room to spare — 4.79 against 2.03 against everything else below 1.7.
How many ellipses would have settled the rest
The sampling error is a standard error of a mean over twenty-five draws, so it falls as the square root of the count, and the count that would settle each unsettled pair follows directly from the gap it already has.
| pair | gap in standard errors | ellipses needed |
|---|---|---|
| the confusion-point basis against CAT16 | 1.23 | 67 |
| Bradford against CAT02 | 1.03 | 95 |
| XYZ against Bradford | 0.37 | 742 |
Two of the three are within reach of an experiment somebody could actually run: sixty-seven ellipses is under three times what MacAdam measured, and ninety-five is under four. The third is not. Separating XYZ from Bradford on this ruler would need seven hundred and forty-two ellipses, thirty times the existing set, and at that point the limiting uncertainty would no longer be the sample — it would be the accuracy of each ellipse, which no amount of resampling improves and which this collection cannot compute at all.
So the honest statement about that pair is not that its order is unknown pending better data. It is that the two spaces are the same distance from uniform as far as this instrument can ever say, and a table printing them one above the other is printing an arbitrary choice with the appearance of a measurement.
Precision follows the score
One property of the instrument falls out of the same numbers and is worth having, because it is the opposite of a defect the previous round found.
A basis that scores better is also scored more precisely. The log–log correlation between a basis’s mean axis ratio and the standard error of that mean is 0.961 across the eight. The best discrimination basis is 1.61 ± 0.07; the adaptation optimum is 7.70 ± 1.56.
The mechanism is not deep and is worth stating anyway: a space that scores well does so by making every ellipse nearly circular, so its twenty-five per-ellipse ratios are all near one and have little spread. A space that scores badly is usually dominated by whichever ellipses it handles worst, so its twenty-five span an order of magnitude.
This is the reverse of the arrangement that made the sampled estimator dangerous. That estimator was accurate where the answer was boring and wrong where it mattered — biased by an amount that grew with the quantity being measured. This one is most precise where the answer is smallest, which is where the interesting comparisons are: the floors, the optima, and the distance between them.
What was computed, and how
The per-ellipse anisotropies come from the closed form — the singular values of the map’s own derivative composed with the ellipse’s shape matrix — rather than from walking round a ring. That matters here for a reason beyond accuracy: the closed form is exact per ellipse, so the only scatter in the twenty-five numbers is the scatter of the diagram itself, with none of the estimator’s own noise added.
The standard errors are computed twice, by the textbook formula and by a bootstrap over four hundred resamples with replacement, and the two are reported side by side. They agree to within about five per cent for the well-behaved bases and differ by more for the badly behaved ones, which is exactly where they should differ: the per-ellipse ratios are strongly skewed under a poor basis — one ellipse at forty and the rest near three — and the textbook formula assumes a symmetry that skew does not have.
The crossing rates are bootstrap quantities as well: resample the twenty-five per-ellipse differences with replacement, average, and count how often the average changes sign. That is the quantity a pair’s ordering depends on, and it needs no assumption of normality — which matters, because the differences are as skewed as the ratios they come from.
Where the model stops
A bootstrap over twenty-five points is a bootstrap over twenty-five points. It estimates the sampling distribution of a mean from a sample small enough that its tails are guesswork, and the crossing rates near five per cent should be read as near the line rather than as numbers.
The twenty-five are not a random sample of anything, which is the sharpest limitation and the one that cannot be repaired from inside. MacAdam chose where to work, and he chose sensibly — spread across the diagram, more densely where discrimination changes fastest. A bootstrap treats them as exchangeable draws from a population of chromaticities, which they are not. What the calculation really estimates is the variability of a mean over sets like this one, and that is a slightly different and less well-defined thing.
A resample is not a redesign. The bootstrap asks what the mean would have been over a different draw of twenty-five from this set, which is a proxy for a different set of twenty-five stimuli and is not the same thing. A genuinely different experiment would choose its chromaticities by some rule — evenly in some coordinate system, or concentrated where an application cares — and the mean over such a set is not the mean this bootstrap is estimating the error of. The distinction is invisible when the sample is large and is not invisible at twenty-five.
And a score is not a preference. A space that makes MacAdam’s ellipses round is uniform at those twenty-five chromaticities, at that luminance, for that observer, at threshold. Every one of those qualifications is load-bearing, and suprathreshold differences are not scaled threshold ones — which is a larger caveat on the whole exercise than any error bar computed here.
The generalisation
The result generalises into a rule that is easy to state and is skipped constantly, including in this collection until now.
Before reading a ranking, compare its adjacent gaps with the standard error of the quantity being ranked. A table sorted to three decimals invites an ordering; whether it has one depends on a number that is usually not printed beside it, and often not computed at all.
The second half of the rule is the one that took work here. When the same instrument scores every entry, the comparison is paired and the error on a gap is far smaller than the errors on the entries. Getting this backwards is a much more common failure than omitting the error entirely, because it produces a defensible-looking calculation with a dramatic conclusion — nothing in this table is distinguishable — that is wrong by a factor of three or four.
The third part is the one this essay would not have predicted: when there are two candidate uncertainties, check whether they order the same pairs, because a disagreement between them is more informative than either alone. Here it identifies precisely which property of the ruler the conclusions are exposed to — its composition, not its accuracy — and that is the sort of thing a single error bar can never say.
What survives
The four essays that argue from this table were checked against the result rather than assumed to be safe, and all four survive, for the same reason each time: they use the ends.
The best axes are not receptors compares the two optima with everything between them, and the optima are separated from the field by 4.8 and 2.0 standard errors. No basis is good at both is about a factor of five between two ends of a diagonal. A constraint costs what it points at uses distances from the floors, and a floor is the most precisely estimated quantity in the set.
What does not survive is any sentence of the form this published transform makes the ellipses rounder than that one. Four such comparisons are available in the middle of the table and none of them is a fact about the transforms. None was made, which is luck rather than judgement, and is now checked rather than lucky.
Who found it, and when
The paired comparison is Student’s and Fisher’s, and predates every number in this essay by a century. The bootstrap is Efron’s, from 1979, and is exactly the tool for a statistic — a mean of skewed ratios — whose sampling distribution nobody wants to write down.
What is not standard is applying either to a canonical dataset that has become an illustration. MacAdam’s ellipses are reprinted as a diagram far more often than they are used as a sample, and a diagram does not have a standard error. The move here is only to remember that the twenty-five points on the picture were twenty-five experiments, and that a mean over them is a statistic like any other.
Where the ladder goes next
Two uncertainties have now been put on this ruler and they disagree, which raises the question of what else in this collection has an uncertainty that has never been asked for. The answer is: a great deal, and it is worth doing systematically rather than one measurement at a time.
The audit that does it works through every published claim here that has a threshold attached, and asks the smallest thing that would have to be wrong for each to stop holding. The most exposed claim on this site is not the one this essay’s neighbours would have nominated.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- An ellipse is not a ring of points anisotropy · cielab · macadam's ellipses · quadratic form · sampling
- Three numbers for one ellipse anisotropy · cielab · convergence · macadam's ellipses · sampling
- How far a quadratic can be believed anisotropy · convergence · degrees of freedom · quadratic form
- The straight line is not the shortest gradient anisotropy · cielab · colour difference · quadratic form
- What one number accepts anisotropy · colour difference · macadam's ellipses · quadratic form
- A mean has a set under it colour difference · degrees of freedom · sampling
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AnisotropyCIELABColour differenceConvergenceDegrees of freedomMacAdam's ellipsesQuadratic formSampling