How wrong would the data have to be
Assumes MacAdam measured it, No diagram makes them circles and An ellipse is not a ring of points.
Every colour space in this collection is scored against twenty-five ellipses, and the twenty-five have no error bars anywhere.
Underneath every one of those numbers is a measurement of one contour, and how it was taken decides what the ranking is a ranking of.
The claim
The ruler has an uncertainty, nothing here has ever propagated it, and the honest way to handle that is to invert the question rather than to invent a number.
- MacAdam’s twenty-five ellipses are a measurement of one person, fitted from tens of thousands of matches, and like every measurement they carry an error that the published table does not state.
- This collection cannot compute that error, because it is a fact about an experiment run in 1942 rather than about any arithmetic done here.
- So it computes the inverse: for each adjacent pair in the ranking of eight bases, how large a relative error on each ellipse’s own axes would make that pair change places one time in twenty.
- The answers run from 17 per cent to never. Two pairs survive any error inside the search; the nearest goes at seventeen.
- And an error shared by every ellipse is far gentler than an independent one of the same size — at fifteen per cent it reorders the table on five per cent of draws when independent and on none at all when common, which is the conservatism in the model measured rather than claimed.
What the ruler is
MacAdam’s ellipses are the one thing this site quotes without apology. A single observer, a bipartite field, tens of thousands of colour matches, and the result that discrimination varies by an order of magnitude across the chromaticity diagram — with the ellipses at true scale mostly smaller than a line width, which is itself the most striking fact about them.
Everything downstream is computed. The question this collection asks of a colour space is how nearly it makes them circles, the answer is a mean axis ratio over the twenty-five, and eight bases are ranked by it in a table four essays argue from.
The previous round corrected how that ratio is computed — from the map’s own derivative rather than by walking round a ring and taking the extremes — and left the data treated as exact. Treating a measurement as exact is the ordinary thing to do and is a choice nobody wrote down.
Why the question has to be inverted
The obvious move is to find MacAdam’s stated errors and propagate them. It is the wrong move here for two reasons, and the second is the one that decides it.
The first is that the published table does not carry them in a form that transfers. What is tabulated is a semi-major axis, a semi-minor axis and an angle per ellipse, fitted to standard deviations of matching along several directions; a modern reader can reconstruct an uncertainty only by making assumptions about the fitting procedure.
The second is that quoting a number here would put a measurement of somebody else’s experiment into a chain of arithmetic that is otherwise entirely this collection’s own — and every result downstream would then be partly a statement about that quoted number, with no way for a reader to separate the two.
So the question is turned round. How large would the error have to be? That is a statement about this collection’s arithmetic, it quotes nobody, and a reader with an opinion about the 1942 experiment can compare it against their own number in one step.
What an error on an ellipse looks like
A model of the error is still needed, and its shape matters more than its size.
The axes are perturbed multiplicatively and the angle additively. A semi-axis is a positive quantity whose uncertainty scales with it; an angle is not. A first version added a fixed number of chromaticity units to both axes, which put the smallest ellipse in the table — a semi-minor axis of 0.35 × 10⁻³, the smallest number in the whole set — through zero on a third of the draws, and produced infinite axis ratios that were an artefact of the noise rather than a property of anything.
The centres are held exactly. They are the stimuli MacAdam chose to work at rather than anything he fitted, so perturbing them would model a different experiment rather than the uncertainty of this one.
And the errors are independent between ellipses, which is the conservative choice rather than the realistic one. Real fitting errors on a set measured by one observer in one session are correlated — a bad day moves all of them — and a correlated error is gentler in its effect on a ranking, because a common factor cancels in a comparison.
There is a third property of the model worth stating because it is the one a reader is most likely to object to. The perturbation is applied to the fitted ellipse rather than to the matches behind it, so it models an error in the summary rather than an error in the data. That is the right level for this question — the summary is what this collection consumes — but it means the model cannot represent an ellipse that is well fitted to matches which were themselves systematically off, which is a different failure and a larger one.
The answers
The published ordering runs: the best discrimination basis at 1.61, the receptor construction at 2.60, CAT16 at 2.71, Hunt–Pointer–Estévez at 2.81, XYZ scaling at 3.44, Bradford at 3.68, CAT02 at 4.30, and the adaptation optimum at 7.70.
- The receptors against CAT16 — a gap of four per cent — reverse at a relative error of 17 per cent.
- CAT16 against Hunt–Pointer–Estévez, a gap of four per cent, at 20 per cent.
- Hunt–Pointer–Estévez against XYZ scaling, a gap of twenty-two per cent, at 32 per cent.
- The best discrimination basis against the receptors at 40 per cent, and Bradford against CAT02 at 42.
- XYZ scaling against Bradford and CAT02 against the adaptation optimum do not reverse under any error inside the search.
Two of those thresholds are worth reading twice. The receptor construction and CAT16 differ by four per cent and are separated by a seventeen per cent error; Bradford and CAT02 differ by seventeen per cent and are separated by a forty-two per cent one. The second pair is further apart and proportionally more robust, which is what a well-behaved ranking should do and is not guaranteed: two bases can differ by a wide margin on the mean and disagree violently on individual ellipses, in which case a small error moves them past each other. That is exactly what happens to XYZ scaling against Bradford under the other instrument, one rung further on.
The two extremes of the table are safe and its middle is not. That is not a coincidence and it is not about the error model: the middle of the ranking is where four bases sit within nine per cent of one another, and nine per cent is a small distance to defend against an instrument nobody has characterised.
The correlated error, measured
The independence assumption is doing real work, so its cost is measured rather than asserted.
Drawing the same fifteen-per-cent error twice — once independently per ellipse, once as a single common factor applied to all twenty-five — gives two very different answers. Independent, the ordering of the eight changes on five per cent of draws. Common, it changes on none of them at all, in a hundred and twenty draws.
The mechanism is immediate once stated: a factor on every ellipse’s axes multiplies every basis’s mean axis ratio by very nearly the same amount, so it moves the table up and down bodily. Only the part of the error that differs between ellipses can reorder anything, because a ranking is a set of comparisons and a common factor cancels in every one of them.
So the numbers above are upper bounds on the fragility. Whatever share of MacAdam’s error is common to his whole session — and some of it certainly is, since the observer and the apparatus were the same throughout — the real reversal thresholds are larger than the ones reported.
The whole curve between the two cases
Independent and fully common are the two ends of one dial, and the dial turns out to have an exact law on it.
Let c be the share of the error’s variance common to every ellipse. Bisecting for the total relative error at which the ranking of the eight changes on one draw in twenty:
| share common | error that reorders the table |
|---|---|
| 0 | 14.0% |
| 0.25 | 16.2% |
| 0.50 | 19.8% |
| 0.75 | 28.0% |
| 0.90 | 44.3% |
Those are 14.0 per cent divided by the square root of the independent share, to three figures at every point: 14.0/√0.75 = 16.2, 14.0/√0.5 = 19.8, 14.0/√0.25 = 28.0, and 14.0/√0.1 = 44.3.
The law is exact and the reason is the mechanism already stated. Only the part of the error that differs between ellipses can reorder anything; the common part multiplies every basis’s mean by the same factor and cancels in every comparison. So the tolerable total error is the tolerable independent error scaled up by however much of the total is shared — and a variance splits as a square, which makes the scaling a square root.
That turns two measurements into one number a reader can use. The independent-only threshold is 14.0 per cent. Anybody with a view about how much of MacAdam’s error was common to his session — the observer, the apparatus, the calibration of his own primaries, all unchanged throughout — divides by the square root of what is left and reads off the total error their view implies the ranking can absorb.
At a half common, which is a modest assumption for one observer on one apparatus, the table survives a twenty per cent error on every ellipse. At three quarters it survives twenty-eight, and at nine tenths, forty-four.
What was computed, and how
Each threshold is a bisection on the logarithm of the relative error, with two hundred draws at each trial value and the criterion set at a five per cent reversal rate for that pair.
Two hundred draws is not many and does not need to be. The quantity wanted is not a p-value but the size of error at which an ordering stops being an ordering, and two hundred draws puts a five per cent rate inside about two points, which is finer than the answer deserves given that the error model itself is a choice.
The axis ratios themselves are computed in closed form from the map’s own derivative composed with the ellipse’s shape matrix, which is the correction the previous round made and matters here for a reason beyond accuracy: the closed form is a smooth function of the ellipse’s parameters, so a perturbed ellipse produces a perturbed answer rather than a discontinuous one. A sampled estimator would have added its own noise to the noise being studied.
Where the model stops
The error model is a choice and its size is not measured. What is reported is a threshold, and a threshold is only useful next to somebody’s belief about the actual number. A reader who thinks MacAdam’s axes are good to five per cent should conclude that every pair in the table is safe; one who thinks they are good to twenty should conclude that three of the seven are not.
Only the shape errors are modelled, not the fitting model. MacAdam fitted ellipses to standard deviations of matching, on the assumption that a discrimination contour is an ellipse. If that assumption is wrong — if the contours have any fourth-order structure — the error is not a perturbation of an ellipse at all and none of this reaches it.
The five per cent criterion is a convention borrowed from a context that does not quite apply. A one-in-twenty reversal rate is the shape of a significance test, and nothing here is testing a hypothesis: there is no null, no sampling distribution over repetitions of an experiment, and no decision being made. What the rate does is give the threshold a definition that does not depend on how many draws were taken. A reader who prefers one in ten will find every threshold about fifteen per cent lower.
And the ranking is not the argument. Four essays here use the table, and none of them rests on a single adjacent pair: what they use is that a receptor-like basis is near the top, that the adaptation optimum is near the bottom, and that the two ends are separated by a factor of five. Those statements survive an error of any size the search covers.
The generalisation
The move here is old and is worth naming, because it is available far more often than it is used.
When a quantity’s uncertainty is unavailable, invert the question and report the threshold. Instead of the answer is 2.71 ± something, report the answer would have to be wrong by seventeen per cent for this conclusion to change. It converts an unanswerable question about somebody else’s experiment into an answerable one about the arithmetic in hand, and it hands the reader the one number they need to apply their own judgement.
It has a second virtue that matters more in a collection like this one. A propagated error bar is a number that gets quoted; a threshold is a number that gets compared, and comparison is the operation that keeps a reader thinking about where the uncertainty came from.
The same inversion is what the population’s declared widths needed, one field over, and the two are the same shape: a modelling input with no distribution attached, and a conclusion whose dependence on it can be computed exactly.
There is a cost, and it should be stated where a reader can see it. A threshold is not composable. Two error bars can be combined; two thresholds cannot, because how wrong would A have to be and how wrong would B have to be do not add into how wrong would the pair have to be. A collection that reports thresholds everywhere ends up with a set of statements that each hold one thing fixed, and the joint question — how wrong could everything be at once — needs a different apparatus and gets one.
What this does not rescue
One thing the inversion cannot do is repair a claim that depends on an adjacent pair. If an essay’s argument were CAT16 makes the ellipses rounder than the receptor construction does, this analysis would not defend it — it would say the sentence is a statement about a seventeen per cent error in a seventy-year-old measurement, and the right response would be to stop making it.
That is worth checking rather than assuming, so it was: the four essays that use the table use it for the shape of its ends. The best axes are not receptors compares the two optima with everything between them. No basis is good at both is about a factor of five. A constraint costs what it points at uses distances from the floors rather than positions in the ranking.
None of the four turns on a pair the error can move, which is a relief and is not something anybody had checked before this essay existed.
Who found it, and when
MacAdam published the twenty-five ellipses in 1942, and their fitted parameters have been reprinted, re-plotted and re-fitted ever since. Later work — Brown and MacAdam’s own extension to three dimensions, and the much larger datasets behind the modern colour-difference formulae — supersedes them as data and has not displaced them as the canonical picture, because twenty-five ellipses on one diagram are the most legible statement of the fact they demonstrate.
The uncertainty on them is discussed in the literature and rarely propagated, which is the ordinary fate of an error bar on a dataset that has become an illustration. Once a figure is reproduced enough times it stops being a measurement with a spread and becomes a diagram with lines in it.
Where the ladder goes next
The error on each individual ellipse is one of two uncertainties in this ruler, and it is the one that needs a literature. The other needs nothing at all: the twenty-five are themselves a sample of the diagram, the score is their mean, and a mean of twenty-five numbers has a standard error those numbers determine.
That second instrument is computable from the data in hand, and it does not agree with this one about which pairs of the table are settled — which is the next rung’s subject and a sharper result than either instrument alone.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- An extremum is not a sample anisotropy · convergence · macadam's ellipses · quadratic form · sampling
- Three numbers for one ellipse anisotropy · cielab · convergence · macadam's ellipses · sampling
- How far a quadratic can be believed anisotropy · convergence · declared input · quadratic form
- The straight line is not the shortest gradient anisotropy · cielab · colour difference · quadratic form
- What one number accepts anisotropy · colour difference · macadam's ellipses · quadratic form
- A deviation is not a difference cielab · colour difference · declared input
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AnisotropyCIELABColour differenceConvergenceDeclared inputMacAdam's ellipsesQuadratic formSampling