Difference and uniformity

An ellipse is not a ring of points

For eleven rounds of argument the distortion a colour space does to MacAdam's ellipses by mapping forty-eight points round each one and dividing the longest radius by the shortest. The minimum sits in a notch whose width is the reciprocal of the answer, so the method was accurate wherever the answer was small and short by a factor of two where it was large.

Assumes MacAdam measured it, No diagram makes them circles and A difference needs a basis too.

The recipe is obvious, it is what anybody with graph paper would have done, and it is wrong by a factor of two exactly where the answer matters.

The minimum sits in a notch the width of the answer's reciprocal. The distance from the centre to the boundary, all the way round one MacAdam ellipse mapped into a lightness–chroma space built on the CIE RGB primaries. The curve has two broad maxima and two very narrow minima: the dip is about 2.5 degrees wide at a third above its floor, because the width of the minimum of an ellipse's radius is the reciprocal of its axis ratio, and this ratio is 41. Forty-eight sample points, marked, are spaced 7.5 degrees apart, so none of them lands in either notch and the smallest one found is 2.2 times the true minimum. The ratio comes out 18.59 where it is 40.76.
Fig. 1 Distance from centre to boundary, right round one of MacAdam’s ellipses mapped into a lightness–chroma space. Two broad maxima, two notches a degree and a half wide, and forty-eight sample points that land in neither notch.

The claim

Measuring an ellipse’s axis ratio by sampling its boundary is accurate when the ratio is near one and useless when it is large, because the width of the minimum is the reciprocal of the ratio. The estimator’s error is set by the quantity it is estimating.

  • On the worst ellipse this site draws, forty-eight points report an axis ratio of 18.59 where it is 40.76.
  • The mechanism is arithmetic, not luck. The radius round an ellipse with semi-axes σ₁ ≥ σ₂ rises by a third within 0.88·σ₂/σ₁ radians of the minimum. At a ratio of 37 that is 1.4°, and forty-eight evenly spaced points are 7.5° apart.
  • So the sample size needed is proportional to the answer, measured at a log–log slope of 0.94 against a predicted 1, over a hundred and thirty-three ellipses.
  • The error is concentrated, not spread. Below an axis ratio of four the median sample is short by 0.35 per cent; above sixteen it is short by 26.
  • And the replacement takes no sample at all. An ellipse is a quadratic form and the map is differentiable, so the image’s semi-axes are the singular values of the derivative composed with the ellipse’s own shape matrix.

What was being measured, and why it is a ratio at all

MacAdam’s twenty-five ellipses are contours of equal discriminability measured on one observer in 1942, and they are the standard against which a colour space’s uniformity is judged. A space is uniform to the extent that the same numerical step means the same perceptual step everywhere, and the ellipses are the direct evidence: in a perfectly uniform space they would all be circles of the same size.

Two numbers come out of the comparison and this essay is about the first. The anisotropy is the mean over the ellipses of the longest radius over the shortest, and it says whether a step means the same thing in every direction at a colour. The spread is the largest mean radius over the smallest, and it says whether a step means the same thing at different colours. Neither is a property of the eye alone — both depend on the basis the space was built on, which is the entire subject of the argument these numbers serve.

The obvious way to compute the anisotropy is to take an ellipse’s boundary, map it into the space, measure the distance from the mapped centre to each mapped boundary point, and divide the largest by the smallest. It is the definition read as a recipe. It was implemented at forty-eight points per ellipse and it has been in every uniformity number this site has published.

The notch, and why its width is the answer

Take the map to be locally linear, which it is at the scale of a MacAdam ellipse — the ellipses are three thousandths of a chromaticity unit across. Then the image of the ellipse is another ellipse, with semi-axes σ₁ ≥ σ₂, and the distance from the centre to the boundary at parameter p is

r(p) = √(σ₁² cos²p + σ₂² sin²p).

Near the major axis, at p = 0, that function is flat: expand it and the leading correction is quadratic with coefficient (σ₂² − σ₁²)/2σ₁, which is small relative to σ₁. A sample anywhere within a wide arc of the major axis returns nearly the maximum, and the maximum is easy.

Near the minor axis, at p = π/2, the same expansion has coefficient (σ₁² − σ₂²)/2σ₂, which is large relative to σ₂ by exactly the ratio being measured. Setting the rise to a third of the minimum gives a half-width of 0.88 σ₂/σ₁ radians. The notch is narrow because the ellipse is long, and it is narrow in proportion.

That is the whole mechanism, and it has an immediate consequence for the method: a sample of n evenly spaced points has a spacing of 2π/n, so it resolves the notch when n is greater than about 7.1 σ₁/σ₂, and steps over it otherwise. Forty-eight points therefore handle an axis ratio of about seven and fail above it.

The points a ratio needs are proportional to the ratio. A scatter of 133 points on logarithmic axes, one per MacAdam ellipse under each of six coordinate systems. The horizontal position is that ellipse's true axis ratio; the vertical is the smallest sample size, from a sequence of doublings, at which the sampled ratio comes within one per cent and stays there. A line of slope 0.94 runs through them, against a predicted 1 — the minimum's notch is 0.88 σ₂/σ₁ radians wide, so resolving it takes a number of points proportional to σ₁/σ₂, and nothing about the basis or the ellipse enters beyond that. An ellipse with a ratio of two needs seventeen points and one with a ratio of twenty-six needs a hundred and ninety-two.
Fig. 2 The smallest sample that gets an ellipse’s ratio right, against that ellipse’s ratio, on logarithmic axes. A hundred and thirty-three ellipses under six coordinate systems, and a line of slope 0.94.

The prediction is checkable and it was checked: for each ellipse under each of six coordinate systems, the smallest sample size on a doubling ladder at which the estimate comes within one per cent and stays there was recorded, and regressed against that ellipse’s own true ratio. The slope is 0.94, against a predicted 1. An ellipse with a ratio of 1.6 needs seventeen points; one with a ratio of 26 needs a hundred and ninety-two.

The condition “and stays there” is doing real work. As the sample slides across the notch the estimate oscillates, and it crosses the true value on the way up. A first crossing is not convergence, and a ladder that recorded one would have found much smaller numbers and a much flatter slope.

How wrong the published numbers were

Averaged over the twenty-five ellipses, in a lightness–chroma space built on each of the coordinate systems this collection draws:

basis sampled at 48 exact
CAT16’s cone responses 2.69 2.71
the receptor construction 2.59 2.60
CIE xy 3.39 3.44
Bradford 3.71 3.68
CIE 1976 u′v′ 4.11 4.20
CAT02 4.07 4.29
Rec. 2020 primaries 5.80 6.40
CIE RGB primaries 6.08 8.93
Display P3 primaries 9.38 10.58
sRGB primaries 9.84 12.45

The pattern is the mechanism. Every basis whose ellipses are nearly circular was measured correctly; every basis whose ellipses are long was measured short; and the error runs from one per cent to forty-seven.

Per ellipse rather than per basis it is starker. Below an axis ratio of four — two hundred and two of the three hundred ellipse-and-basis pairs — the median sample is short by 0.35 per cent. Above sixteen, the median is short by 26 per cent and the worst by 119.

How many points a ratio needs is set by the ratio. Twelve curves, one per coordinate system this collection draws, each showing the axis ratio a sample of n points reports as a fraction of the exact value, against n on a logarithmic axis. Every curve rises to one, and they reach it at wildly different places: the systems whose ellipses are nearly circles are right at twenty-four points, and the CIE RGB primaries — whose worst ellipse has an axis ratio of 9 — are still 32 per cent short at forty-eight. The curve a measurement is on cannot be known until the answer is, which is what makes a fixed sample size the wrong instrument for this question.
Fig. 3 What a sample of n points reports as a fraction of the truth, for every coordinate system here. The curves reach one at wildly different places, and which curve a measurement is on cannot be known until the answer is.

What replaced it

An ellipse is a quadratic form. Write E for the 2×2 matrix carrying the unit circle onto the ellipse in chromaticity — a rotation by the reported angle composed with the two reported semi-axes — and J for the Jacobian of the map from a chromaticity to the lightness–chroma space. Then the image of the unit circle is the image of M = J·E, and its semi-axes are M’s singular values.

M is 3×2, and its singular values are the square roots of the eigenvalues of the 2×2 matrix MᵀM, which is one line of arithmetic. That route squares the condition number, which is exactly what a careful decomposition warns against — and the warning is about generality rather than about this case. The largest axis ratio anywhere in this collection is 40.8, so the squared ratio is 1,600 against a double’s sixteen digits, and the closed form agrees with a Jacobi decomposition of the same matrix to 10⁻¹³ on every one of the three hundred ellipses. The check is in the build rather than in this paragraph.

The Jacobian is written out rather than differenced. The map is a chromaticity lifted to tristimulus values at fixed luminance, then a basis, then a division by the white point, then the CIE compression, then three differences — every step differentiable except at the break point where the cube root hands over to its linear toe. Checked against central differences at every ellipse centre under twelve bases, the worst relative discrepancy is 2.7 × 10⁻⁸, and the closest any ellipse centre comes to the break point is 7 × 10⁻⁴ of it, so no comparison is being made across the join.

Writing the derivative out is what makes the correction more than a correction. A maximum over a sample is not a differentiable function of anything: move the basis a little and the argmax jumps from one sample point to the next, and the estimate has a kink. The closed form is analytic in the entries of the basis away from the degenerate case where the two axes coincide, and that is what makes it possible to ask what the curvature of the objective built on it looks like — a question the sampled version could not be asked at all, because its second differences scale as one over the step and diverge.

What was computed, and how

Three assertions carry this, and each is built to fail if the sample is ever quietly put back in front of the form.

The first requires the coarse sample to be wrong by a stated factor somewhere — without which nothing would have had to change — and requires the error to be concentrated where the answer is large rather than spread over the set. Both are needed: the second is the mechanism, and a version of the first on its own would be satisfied by any estimator with a bad day.

The second is the slope: the sample size a ratio needs against the ratio, at 0.94 with a bound of 0.75 to 1.25 either side of the predicted 1.

The third is the one that says the two estimators are measuring the same thing rather than two different things that happen to be close. The differential is the limit of the finite ellipse’s image as the ellipse shrinks to a point, so a very finely sampled ring at magnification m must approach the analytic answer as m falls. It does: 0.65 to 1.54 per cent above at full size, and within 0.09 per cent at a sixteenth of the size, on three bases chosen to span the range. The sample size in that test is 6,144 deliberately — a coarse sample would not converge to anything, because its own error moves with the ratio and the two effects would be inseparable.

Three numbers for one set of ellipses, and which of them is which. Two curves and a horizontal line, against the size the ellipses are drawn at. The line is the analytic axis ratio — the ratio of the singular values of the map's own derivative, which is what "does this space make discrimination contours circles" means. The upper curve is a very finely sampled ring, which sits 0.6 per cent above the line at full size and converges onto it as the ellipse shrinks, because the gap between them is the second-order distortion of the map across a real ellipse rather than an error. The lower curve is the forty-eight-point sample used for this until now: it does not converge onto anything, because its error is set by the sample and not by the size.
Fig. 4 The analytic answer as a horizontal line, a very fine ring converging onto it as the ellipse shrinks, and the forty-eight-point sample converging onto nothing.

The other three views of the same machinery say how much of the ranking this sampling error is capable of moving, which is the question a reader is really asking.

Twenty-five ellipses is a sample, and the score has an error bar. One row per colour space this collection ranks: the mean axis ratio its ellipses come out at, with the standard error of that mean over the twenty-five ellipses it was computed from. No literature is quoted — a mean of twenty-five numbers has a standard error those twenty-five numbers determine. The bars are far from equal: the best space carries ± 0.07 and the worst ± 1.56, because a space that makes the ellipses nearly circular makes all of them nearly circular and one that does not is dominated by whichever ellipse it handles worst.
Fig. 5 Each basis’s score with the standard error of that score over the twenty-five ellipses. The sampling error inside one ellipse is one term; the sample of ellipses is the other, and they are not the same size.
Three of the seven adjacent pairs in the table are actually ordered. One bar per adjacent pair of the uniformity ranking: the difference between the two spaces' scores divided by the standard error of that difference over the twenty-five ellipses. The comparison is paired — the same ellipses score both spaces, so an ellipse that is hard for everybody cancels — which is why the table is more informative than it looks and why treating the two errors as independent would have declared almost nothing ordered. 3 of the 7 pairs clear two; the rest do not, and one of them crosses zero in 35 per cent of resamples. The ranking's ends are real and its middle is not a ranking.
Fig. 6 Each adjacent pair of the ranking as its gap divided by the standard error of that gap. Two rows separated by less than one of these are two rows the data does not order.

Which of the two sampling errors dominates is settled by the third view, and it is the one a reader should be told about rather than the one that is easiest to compute.

How wrong the ellipses would have to be for a pair to change places. One bar per adjacent pair in the uniformity table: the relative error on each ellipse's own axes at which that pair changes places in one draw in twenty. No error on the data is quoted anywhere — the question is inverted, so what is reported is how large an error would have to be, and a reader with an opinion about MacAdam's experiment can compare it with their own number. The nearest pair goes at 0.171; 2 of the 7 pairs do not reverse under any error this search covers.
Fig. 7 And how large an error on the ellipses themselves would have to be to reverse each step. That is the honest form of the question, and the answer is not the same for every pair.

Which published numbers moved

The floors did not. The best lightness–chroma space built on any basis leaves a mean axis ratio of 1.611 where it was reported as 1.620, and the best space with no compression at all leaves 2.334 where it was 2.36. Both of those are optima and an optimum is where the ellipses are as round as they can be made, which is exactly where the sampled estimator was accurate. The floor below the projective floor survives at four significant figures.

The published bases moved a great deal, and always in the direction of looking worse. The basis that minimises the adaptation residual is at 7.70 on the other objective where it was reported at 6.78; the sRGB primaries at 12.45 where they were at 9.84.

The ranking did not move, and that is not luck. Which of two bases is worse is settled the same way it was. The estimator’s error is monotone in the quantity it is estimating, so it compresses an ordering without reversing one. Every comparison this collection has drawn between two bases on this axis stands; every distance between them was understated.

Where the model stops

The analytic answer is exact about the differential, which is not quite the finite ellipse. The difference between them is second-order distortion of the map across a real MacAdam ellipse — 0.65 to 1.54 per cent on the bases tested — and it is a third quantity with its own essay rather than an error in either.

The ellipses themselves are a measurement of one person in 1942. Nothing here improves them, and the between-observer variation in discrimination contours is not modelled anywhere on this site. A ratio computed exactly from an ellipse that was fitted to twenty-five thousand judgements by one observer is exact about that observer.

And the anisotropy is a mean over twenty-five ellipses. The correction changes which ellipses dominate that mean: the long ones were being under-weighted, so a basis whose failure was concentrated in a few places now looks worse relative to one that fails evenly. Whether the mean is the right summary is a separate question and the answer here is the same as it was — the per-ellipse table is drawn rather than replaced by its average.

The ranking moved by one pair

The ranking did not move, and that is not luck is the reassuring sentence in this essay, and the table three sections above it contains a counterexample.

Ordered by the exact answer the ten bases run receptor, CAT16, xy, Bradford, u′v′, CAT02, Rec. 2020, CIE RGB, P3, sRGB. Ordered by what forty-eight points reported, CAT02 and u′v′ swap: the sample puts CAT02 at 4.07 below u′v′ at 4.11, and the exact measurement puts it at 4.29 above u′v′ at 4.20. One inverted pair out of forty-five — the ordering is nearly preserved, which is a weaker and truer statement than that it is preserved.

The two are a hundredth apart in the sample and nine hundredths apart exactly, so this is not a dramatic reversal. It is the one place where the argument’s own numbers refuse the argument’s reasoning, and the reasoning is where the interest is.

The error is not monotone, so nothing protects the ordering

The claim that carries the ranking is that the estimator’s error is monotone in the quantity it is estimating, so it compresses an ordering without reversing one. The first half is what would have to be true, and the table says it is not.

Reading the shortfall down the exact ordering: 0.4, 0.7, 1.5, −0.8, 2.1, 5.1, 9.4, 31.9, 11.3, 21.0 per cent. Two of those break the sequence, and the second breaks it hard. CIE RGB, at an exact ratio of 8.93, is short by 31.9 per cent, while Display P3 at the larger ratio of 10.58 is short by 11.3. A basis with rounder ellipses lost three times as much. Whatever sets the shortfall, it is not the mean ratio alone — which is the expected consequence of averaging twenty-five ellipses, since the mean of the ratios and the mean of the sampling losses are summaries of two different distributions and a single long ellipse can dominate one without dominating the other.

So the ordering survives here as an empirical fact about ten bases rather than as a guarantee, and it survives with one exception. A compression argument needs monotonicity per ellipse to be inherited by the mean, and it is not.

Bradford’s row cannot happen

The other break in the sequence is the sharper one, because it violates a bound rather than a trend.

A sampled maximum is at most the true maximum and a sampled minimum is at least the true minimum, so a sampled axis ratio is at most the true one, on every ellipse and therefore on any mean of ellipses. Bradford is reported at 3.71 sampled against 3.68 exact. On the same object that is arithmetically impossible, which means the two columns are not measurements of the same object.

The essay says elsewhere what the difference is, without connecting it here: the sample walks a finite MacAdam ellipse and the closed form describes the differential, and a finely sampled finite ellipse sits 0.65 to 1.54 per cent above the analytic answer at full size. The sampled column therefore carries two errors of opposite sign — finite-size inflation of about a per cent, and under-sampling deflation that grows with the ratio — and Bradford is where they cross.

That fixes what the top of the table is worth. Below an exact ratio of about four the two effects are the same size, so the first four rows measure their difference and not the sampling loss: the 0.4 and 0.7 per cent shortfalls at the receptor construction and CAT16 are a percent of inflation against a percent and a half of loss, not evidence that forty-eight points nearly sufficed. The figure of 0.35 per cent quoted for the median ellipse below a ratio of four is a difference of two errors read as one.

None of this touches the conclusion that the coarse sample had to go. It sharpens what replaced it: the closed form is not a more accurate version of the sample, it is an answer to a different question, and the essay’s own three-way distinction is the place that is settled.

The generalisation

An operational description of a quantity is a recipe, and a recipe is correct in a limit that nobody runs it in. Walk round the boundary and see how far it goes defines the axis ratio perfectly and computes it badly. The gap between the two is not carelessness; it is the difference between a definition and an algorithm, and it opens exactly where the definition is most interesting.

The tell is available in advance and costs nothing. Before implementing a definition as a sweep, ask how the answer behaves near the point the sweep has to find, and compare that width against the sweep’s spacing. Here the width was 0.88 σ₂/σ₁ and the spacing was 2π/n, and the two are one line apart on paper.

A sampled maximum has no error bar, which is why it survives. Take another forty-eight points and the answer barely moves: the same wide arcs are sampled and the same narrow notch is missed. It is reproducible and wrong, and reproducibility is what a reviewer checks.

Who found it, and when

MacAdam published the ellipses in 1942 and Brown and MacAdam extended them to three dimensions in 1949. Every uniformity comparison since has needed some way of turning an ellipse into a number, and the literature’s usual route is not the sampled one — it is to transport the quadratic form, which is what the ellipse actually is, through the linearised map. That is the method here now.

The sampled route appears where a figure is being drawn as well as measured, which is the honest explanation for its appearance in this collection: the boundary points had to be computed anyway to draw the ellipse, and taking their extremes was free. A number computed as a by-product of a drawing is a number nobody chose an estimator for. That is the shape of the defect, and it is not confined to colour.

Where the ladder goes next

The correction leaves a genuinely three-way distinction that this collection had collapsed into one number, and the three are worth separating properly: what a coarse sample of a ring reports, what a fine sample of a finite ellipse reports, and what the differential says. They differ by measurable amounts and they mean different things, and only one of them is the quantity the phrase does this space make discrimination contours circles is asking about.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 11 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyBasisCIELABExtremumJacobianMacAdam's ellipsesQuadratic formSamplingSingular valueUniformity