Difference and uniformity

Three numbers for one ellipse

How far a colour space is from making a discrimination contour circular has three different answers — what a coarse sample of the boundary reports, what a converged sample of a contour of stated size reports, and what the map's own derivative says. They differ by up to a factor of two, they mean different things, and only one of them is what the question is asking.

Assumes An ellipse is not a ring of points, MacAdam measured it and A difference has no size.

A discrimination contour is an infinitesimal object drawn at a magnification, and every number taken from the drawing depends on which of those two things was measured.

Three numbers for one set of ellipses, and which of them is which. Two curves and a horizontal line, against the size the ellipses are drawn at. The line is the analytic axis ratio — the ratio of the singular values of the map's own derivative, which is what "does this space make discrimination contours circles" means. The upper curve is a very finely sampled ring, which sits 0.6 per cent above the line at full size and converges onto it as the ellipse shrinks, because the gap between them is the second-order distortion of the map across a real ellipse rather than an error. The lower curve is the forty-eight-point sample used for this until now: it does not converge onto anything, because its error is set by the sample and not by the size.
Fig. 1 The derivative’s answer as a horizontal line, a converged sample of the finite contour above it, and a forty-eight-point sample below — three numbers for one set of ellipses.

The three numbers are three readings of one object, and the object itself is what a sample is trying to trace.

The minimum sits in a notch the width of the answer's reciprocal. The distance from the centre to the boundary, all the way round one MacAdam ellipse mapped into a lightness–chroma space built on the CIE RGB primaries. The curve has two broad maxima and two very narrow minima: the dip is about 2.5 degrees wide at a third above its floor, because the width of the minimum of an ellipse's radius is the reciprocal of its axis ratio, and this ratio is 41. Forty-eight sample points, marked, are spaced 7.5 degrees apart, so none of them lands in either notch and the smallest one found is 2.2 times the true minimum. The ratio comes out 18.59 where it is 40.76.
Fig. 2 Distance from centre to boundary right round one contour. A ring of points reports the largest and smallest of what it happens to land on; the derivative reports the extremes of this curve, and the two are the same only when the sample is fine enough to have found them.
The points a ratio needs are proportional to the ratio. A scatter of 133 points on logarithmic axes, one per MacAdam ellipse under each of six coordinate systems. The horizontal position is that ellipse's true axis ratio; the vertical is the smallest sample size, from a sequence of doublings, at which the sampled ratio comes within one per cent and stays there. A line of slope 0.94 runs through them, against a predicted 1 — the minimum's notch is 0.88 σ₂/σ₁ radians wide, so resolving it takes a number of points proportional to σ₁/σ₂, and nothing about the basis or the ellipse enters beyond that. An ellipse with a ratio of two needs seventeen points and one with a ratio of twenty-six needs a hundred and ninety-two.
Fig. 3 How fine that has to be, ellipse by ellipse. The requirement rises with the axis ratio, which is exactly the case a uniformity measure is interested in — so the sample is worst where the answer matters most.

The claim

“How far is this space from making the discrimination contours circles” has three answers, and a table that does not say which one it holds is under-specified by up to a factor of two.

  • The coarse sample — the boundary walked at forty-eight points — is an artefact and is dealt with elsewhere. It is short by up to 47 per cent.
  • The converged finite contour is a real quantity and it depends on the contour’s size. At the magnification MacAdam’s ellipses are usually drawn at, the ratio is 7.6 to 20.4 per cent larger than at their tabulated size.
  • The differential — the singular values of the map’s own derivative at the ellipse’s centre — is the only one of the three that is a property of the space at a point, and it is the limit the other two approach as the contour shrinks.
  • The gap between the last two is not error. It is second-order distortion of the map across a real contour, it runs 0.65 to 1.54 per cent at the tabulated size, and it is larger where the space is worse.
  • And the question decides which is meant. Is a step of this size the same in every direction is about the finite contour. Is this space locally uniform is about the derivative. They are different questions and the discipline asks the second while measuring the first.

The three, in one picture

Take one of MacAdam’s ellipses, carry it into a lightness–chroma space, and ask for its axis ratio.

The first number comes from mapping forty-eight points round the boundary and dividing the longest resulting radius by the shortest. It is short, by an amount proportional to the answer, for the reason set out in the essay that corrected it: the minimum sits in a notch whose width is the reciprocal of the ratio.

The second number comes from doing the same thing with enough points that the answer has stopped moving — six thousand rather than forty-eight. That is a real measurement of a real object: the image, in this space, of the closed curve MacAdam tabulated.

The third number comes from not walking anything. The map from a chromaticity to the space is differentiable; its derivative at the ellipse’s centre is a 3×2 matrix; composed with the 2×2 that carries the unit circle onto the ellipse, its singular values are the semi-axes of the image of an infinitesimal ellipse of that shape. It is the local metric, and it takes no sample and no size.

At the tabulated size the second sits 0.65 to 1.54 per cent above the third. That gap is the map bending across the contour — the contour is not tiny enough for the linearisation to be exact, and the parts of it further from the centre are distorted more.

The second number moves with the size, and by a lot

This is the part that makes the distinction more than pedantry.

Shrink the contour and the second number falls onto the third; grow it and the second number rises away. Measured on the same twenty-five ellipses in a lightness–chroma space built on CAT16’s cone responses, the converged sample gives 2.707 at a sixteenth of the tabulated size, 2.723 at the tabulated size, 2.780 at four times it and 2.911 at ten times. The differential is 2.706.

In a space built on the sRGB primaries, where the ellipses are long, the same ladder runs 12.46, 12.64, 13.26, 14.99, against a differential of 12.45.

So a quoted axis ratio without a stated contour size is under-specified by up to twenty per cent, and the size most often used in print — a tenfold magnification, so that the ellipses can be seen — is the one furthest from the limit. The magnification is a drawing convention. It has no business being in a measurement, and it is.

The gap has a law

The finite contour’s excess over the differential is quoted at two sizes, and between them it follows a single power.

Measured at magnifications of 1, 2, 5, 10 and 20, with the contour converged at six thousand points, the excess for the 1931 plane runs 0.792, 1.606, 4.194, 9.072 and 22.500 per cent. For u′v′: 0.868, 1.758, 4.568, 9.797 and 23.497. For the CAT16 cone plane: 0.646, 1.314, 3.461, 7.588 and 18.970. For sRGB’s own primaries: 1.538, 2.963, 8.854, 20.358 and 47.348.

Fitted between the ends, the exponent on the magnification is 1.12 for the 1931 plane and between 1.10 and 1.17 for the other three. The excess is very nearly proportional to the contour’s size.

That is what a second-order term in the map produces in a ratio of radii, and the arithmetic is worth spelling out because it is not the obvious answer. The displacement from the linear map grows as the square of the step; the ratio of two such radii departs by only the first power, because the quadratic correction is being compared against a linear term rather than against nothing.

Two things follow. The gap extrapolates: a reader holding a contour at any magnification can multiply the tabulated-size excess by that magnification and be right to about ten per cent as far out as a factor of twenty. And it extrapolates to zero, which is the sense in which the differential is a limit rather than a fourth opinion — the finite contour converges on it linearly as the contour shrinks, at the rate the table above measures.

The ordering across the four planes holds at every size, with sRGB’s primaries always worst and the CAT16 cone plane always best. So which of the three numbers is quoted changes the value at every magnification and never changes which space is which.

Which one the question wants

The phrase this whole family of measurements exists to serve is a uniform space is one in which the same numerical step means the same perceptual step everywhere. Unpacked, that is a statement about a metric: a colour difference formula computes a distance, and the claim is that the distance agrees with discriminability. A metric is a local object, and a difference has no place in it beyond the point it is computed at. What it does at a point is a quadratic form, and the question of whether it is isotropic at that point is a question about the derivative.

So the third number is the one the sentence is about.

The second number is what the sentence is about plus a statement about how far the local approximation reaches, which is a different and also useful thing — it is the answer to how big can a step be before the local metric stops describing it, and that question has its own name and its own essay. Conflating them puts a curvature term into a number that is supposed to be about anisotropy.

The first number is about neither.

What the tabulated size even is

MacAdam’s contours are not just-noticeable differences. They are standard-deviation contours of a colour-matching distribution: one observer set one half of a bipartite field to match the other repeatedly, the settings scattered, and the ellipse is the one-standard-deviation contour of that scatter, fitted to the settings along each of several directions. The published table gives the semi-axes in units of a thousandth of a chromaticity coordinate, which makes each of them a threshold rather than a unit, and the usual figure magnifies them tenfold.

A just-noticeable difference is roughly three of those standard deviations, so the contour a discrimination argument is about is larger than the tabulated one, not smaller — which pushes the second number further from the third rather than closer. Measured at three times the tabulated size, the CAT16 basis gives 2.760 against a differential of 2.706, CIE xy gives 3.527 against 3.443, and the sRGB primaries give 13.03 against 12.45.

That is the honest position and it is not a comfortable one: the object the discipline reasons about is at a size where the local metric is already three to five per cent wrong about it. Nothing here fixes that. What this collection can do is stop quoting one number for two things.

What the three do to a comparison

The distinction would be academic if the three numbers ranked the spaces the same way, and mostly they do — but not entirely, and the exception is instructive.

Between the differential and the converged contour at tabulated size, nothing changes order: the gap is 0.65 to 1.54 per cent and no two entries in the table are that close. Between the differential and the contour at tenfold magnification the gaps reach twenty per cent, and two of the entries — the CIE 1976 u′v′ coordinates and CAT02’s cone responses, at 4.20 and 4.29 on the differential — are eight per cent apart and can be swapped by a choice of drawing scale.

That is a small reversal on a pair nobody argues about, and it is the shape of the risk rather than an important instance of it. A ranking that survives one measurement convention and not another is a ranking with a convention in it, and the only defence is to say which convention.

The coarse sample is different again, and worse in a way that is easy to miss: because its error is monotone in the answer it compresses the table without reordering it. Every comparison drawn from it was in the right direction and every distance was understated, which is exactly the failure mode that survives review — the conclusions are right, the evidence is soft, and nothing looks wrong.

What was computed, and how

The differential is computed in closed form: an analytic Jacobian of the map, composed with the ellipse’s own shape matrix, and the singular values of the resulting 3×2 taken from its 2×2 Gram matrix. The Jacobian is checked against central differences at every ellipse centre under twelve bases — worst relative discrepancy 2.7 × 10⁻⁸ — and the closed-form axes against a Jacobi decomposition of the same matrix, agreeing to 10⁻¹³ at a largest ratio of 40.8.

The convergence of the second number onto the third is the assertion that says the two are measuring the same thing. It runs a very fine ring — 6,144 points, so that the sample’s own error is far below the effect being measured — at magnifications of 1, 1/4 and 1/16 on three bases spanning the range, and requires the residual at the smallest magnification to be under 0.4 per cent and to be smaller than at full size rather than merely different. It comes back at 0.04 to 0.09 per cent.

The sample size is large deliberately, and the reason is the whole difficulty of the subject: a coarse sample would not converge onto anything. Its error is set by the axis ratio, the axis ratio moves with the magnification, and the two effects would be inseparable. Separating them is what the three numbers are for.

How many points a ratio needs is set by the ratio. Twelve curves, one per coordinate system this collection draws, each showing the axis ratio a sample of n points reports as a fraction of the exact value, against n on a logarithmic axis. Every curve rises to one, and they reach it at wildly different places: the systems whose ellipses are nearly circles are right at twenty-four points, and the CIE RGB primaries — whose worst ellipse has an axis ratio of 9 — are still 32 per cent short at forty-eight. The curve a measurement is on cannot be known until the answer is, which is what makes a fixed sample size the wrong instrument for this question.
Fig. 4 The coarse sample’s own error, against sample size, across every coordinate system here. This is the effect that has to be removed before the magnification effect can be seen at all.

The other three views of the same machinery say how much of the ranking this sampling error is capable of moving, which is the question a reader is really asking.

Twenty-five ellipses is a sample, and the score has an error bar. One row per colour space this collection ranks: the mean axis ratio its ellipses come out at, with the standard error of that mean over the twenty-five ellipses it was computed from. No literature is quoted — a mean of twenty-five numbers has a standard error those twenty-five numbers determine. The bars are far from equal: the best space carries ± 0.07 and the worst ± 1.56, because a space that makes the ellipses nearly circular makes all of them nearly circular and one that does not is dominated by whichever ellipse it handles worst.
Fig. 5 Each basis’s score with the standard error of that score over the twenty-five ellipses. The sampling error inside one ellipse is one term; the sample of ellipses is the other, and they are not the same size.
Three of the seven adjacent pairs in the table are actually ordered. One bar per adjacent pair of the uniformity ranking: the difference between the two spaces' scores divided by the standard error of that difference over the twenty-five ellipses. The comparison is paired — the same ellipses score both spaces, so an ellipse that is hard for everybody cancels — which is why the table is more informative than it looks and why treating the two errors as independent would have declared almost nothing ordered. 3 of the 7 pairs clear two; the rest do not, and one of them crosses zero in 35 per cent of resamples. The ranking's ends are real and its middle is not a ranking.
Fig. 6 Each adjacent pair of the ranking as its gap divided by the standard error of that gap. Two rows separated by less than one of these are two rows the data does not order.

Which of the two sampling errors dominates is settled by the third view, and it is the one a reader should be told about rather than the one that is easiest to compute.

How wrong the ellipses would have to be for a pair to change places. One bar per adjacent pair in the uniformity table: the relative error on each ellipse's own axes at which that pair changes places in one draw in twenty. No error on the data is quoted anywhere — the question is inverted, so what is reported is how large an error would have to be, and a reader with an opinion about MacAdam's experiment can compare it with their own number. The nearest pair goes at 0.171; 2 of the 7 pairs do not reverse under any error this search covers.
Fig. 7 And how large an error on the ellipses themselves would have to be to reverse each step. That is the honest form of the question, and the answer is not the same for every pair.

The second ratio moves too, and further

Anisotropy is one of two numbers the ellipses give. The other is the spread — the largest mean radius over the smallest, across the set — and it asks a different question: whether a step of a given size means the same thing at one colour as at another.

It is measured the same way and it inherits the same three-way distinction, and because it is a ratio of means rather than of extremes it behaves differently in one respect and identically in another. A mean radius is exactly the kind of quantity a sample estimates well, so the coarse sample’s notch problem barely touches it. But the mean radius of a finite contour depends on the contour’s size in the same way an axis does, so the spread inherits the whole of the magnification question.

On the differential, the spreads run from 2.53 for CAT16’s cone responses and 2.69 for the receptor construction, through 3.24 for CIE xy, to 13.58 for Bradford, 18.64 for the sRGB primaries, 28.19 for CIE RGB and 30.81 for CAT02. Those are much larger numbers than the anisotropies and they are much more spread out: a space can be locally round everywhere and still have its unit mean different things by a factor of thirty between one end of the diagram and the other.

Which of the two matters depends on the use, and this collection has said so before. A tolerance written as a single number is a claim about the spread; a tolerance written as a shape is a claim about the anisotropy. Both are quoted from the same table and only one of them is usually meant.

Where the model stops

Linearising a projective map is exact only in the limit. The lift from a chromaticity to tristimulus values at fixed luminance divides by y, and the compression that follows takes a cube root. Neither is linear, so the image of a finite ellipse is not an ellipse — it is a closed curve that is nearly one, and the “axis ratio” of the second number is really a longest radius over a shortest radius, which for a non-ellipse is a different quantity from a ratio of semi-axes. At the tabulated size the difference is under two per cent and at tenfold magnification it is not.

Nothing here says which size a discrimination argument should use. A tolerance is a shape whose size is part of the specification, and the same is true here. That depends on what the argument is about, and the answer is different for a tolerance in a printing specification and for a threshold in a psychophysical experiment. What is available now is that both can be computed and that they differ.

And a contour of a matching distribution is not a contour of discriminability. The two are related by a factor that is itself a modelling choice, and the standard route — a just-noticeable difference is about three standard deviations — is a convention rather than a measurement. Every number here is about the tabulated contours and their scalings.

The generalisation

A quantity defined as a limit and measured on a finite object carries the size of the object in it. That is nearly a tautology and it is skipped constantly, because the finite object is the thing that exists and the limit is the thing that was meant.

The useful discipline is not to always take the limit. It is to check how fast the finite answer moves with the size, and to report the size whenever it moves at all. Here it moves by twenty per cent over the range of sizes actually in use, which is larger than most of the differences the numbers are used to argue about — so the size is not a footnote, it is one of the inputs.

There is a second habit hidden in it. The magnification in these figures exists so that the ellipses can be seen; it is a property of the paper. It leaked into the measurement because the same array of boundary points served both purposes, which is the same mechanism that produced the coarse sample in the first place. A number computed as a by-product of a drawing inherits the drawing’s conventions, and the conventions of a drawing are chosen for legibility.

Who found it, and when

The distinction between a metric and its finite consequences is Riemann’s and is not in dispute. The colour literature’s version of it is the standing observation that a colour difference formula’s own local metric is a quadratic form — the line-element formulations of Helmholtz, Schrödinger and Stiles are exactly that — and the modern formulas are usually derived and compared as line elements, which is the differential quantity.

Where the finite contour creeps back in is the comparison of spaces rather than of formulas, because that comparison is done with a picture. A figure of transported ellipses is drawn at a magnification, its shapes are inspected, and the number that accompanies it comes from the same shapes. The published tables of ellipse anisotropy in various spaces do not generally state a magnification, and this collection’s did not either until now.

Where the ladder goes next

The differential is a smooth function of the basis, which is what makes it possible to ask the next question at all — and which a maximum over a sample could never have supported: what does the curvature of an objective built on these ellipses look like, and which directions in the space of bases does it actually see? The answer has a rank in it, and the rank is smaller than the number of parameters.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyCIELABConvergenceJacobianJust-noticeable differenceMacAdam's ellipsesMetricSamplingSingular valueUniformity