Where the model breaks

An extremum is not a sample

Three separate measurements in this collection took a maximum or a minimum over a sample of a set — forty-eight points round an ellipse, twenty-four directions out of an optimum, fourteen changes of light off a list. All three are wrong, all three are wrong in the same direction, and the error in each grows with the very quantity being measured.

Assumes A fit can be exact and empty, A constraint costs what it points at and MacAdam measured it.

A mean survives being sampled. A maximum does not, and the difference is not a matter of degree.

The points a ratio needs are proportional to the ratio. A scatter of 133 points on logarithmic axes, one per MacAdam ellipse under each of six coordinate systems. The horizontal position is that ellipse's true axis ratio; the vertical is the smallest sample size, from a sequence of doublings, at which the sampled ratio comes within one per cent and stays there. A line of slope 0.94 runs through them, against a predicted 1 — the minimum's notch is 0.88 σ₂/σ₁ radians wide, so resolving it takes a number of points proportional to σ₁/σ₂, and nothing about the basis or the ellipse enters beyond that. An ellipse with a ratio of two needs seventeen points and one with a ratio of twenty-six needs a hundred and ninety-two.
Fig. 1 Every ellipse this collection measures, under six coordinate systems, plotted as the axis ratio it has against the number of sample points it takes to find that ratio. The line through them has slope 0.94 against a predicted 1.

The claim

Three measurements in this collection were made by walking round a set, or across one, and keeping the largest or smallest thing found. Every one of them is short, every one is short by more where the answer is larger, and in each case there was an exact method available that needed no sample at all.

  • The axes of a discrimination ellipse, taken as the largest and smallest of forty-eight radii. Short by a factor of 2.19 on the worst ellipse this site draws, and right to a fraction of a per cent on the roundest.
  • The shape of an optimum, taken from twenty-four random directions walked outward until the cost rose. It reported a bowl 8.0 times longer one way than another where the answer is 29.8.
  • The worst change of light, taken as the largest of fourteen changes on a list. The family those fourteen were drawn from reaches six times further, and does not stop there.
  • In each case the error is set by the answer, not by the care taken. A sample of n points resolves an axis ratio when n is a few times the ratio; a random direction in nine dimensions reports the mean curvature whatever the spread; a list of fourteen scenes reports the worst of fourteen scenes.
  • And in each case the replacement is arithmetic rather than a bigger sample. A quadratic form has singular values. A quadratic bowl has eigenvalues. A family with no interior maximum has no worst case to find, and saying so is the result.

Why a mean is safe and a maximum is not

The reason is old and is worth stating in one line before any of the three cases, because it is the whole of them.

A sample mean converges on the population mean at a rate that depends only on how many samples were taken and how much the quantity varies — not on where the samples fell. Every point contributes. A sample maximum converges on the population maximum at a rate that depends on how much of the set is near the maximum, and in every case here that fraction is small and gets smaller as the extremum gets more extreme. The peak is narrow because it is a peak.

So the failure is not that sampling is approximate. It is that the estimator is biased in a known direction, by an amount that grows with what it is trying to report, which is the worst possible arrangement: it is accurate when the answer is boring and wrong when the answer matters. A table of such numbers looks plausible everywhere and has its extreme rows quietly halved.

The minimum sits in a notch the width of the answer's reciprocal. The distance from the centre to the boundary, all the way round one MacAdam ellipse mapped into a lightness–chroma space built on the CIE RGB primaries. The curve has two broad maxima and two very narrow minima: the dip is about 2.5 degrees wide at a third above its floor, because the width of the minimum of an ellipse's radius is the reciprocal of its axis ratio, and this ratio is 41. Forty-eight sample points, marked, are spaced 7.5 degrees apart, so none of them lands in either notch and the smallest one found is 2.2 times the true minimum. The ratio comes out 18.59 where it is 40.76.
Fig. 2 The distance from centre to boundary, right round one ellipse mapped into a lightness–chroma space. Two broad maxima, two very narrow minima, and forty-eight evenly spaced sample points that miss both minima.

The first case: an ellipse walked round

MacAdam’s ellipses are the ruler this collection measures a colour space’s uniformity with. The question asked of a space is how nearly it makes them circles, and the obvious way to answer it is to take the twenty-five ellipses, map their boundaries into the space, and take the longest radius over the shortest.

That is what was done, at forty-eight points per ellipse, for as long as this site has been asking the question.

The image of a small ellipse under a differentiable map is, to first order, another ellipse — and the radius round an ellipse with semi-axes σ₁ ≥ σ₂ is √(σ₁²cos²p + σ₂²sin²p). Near the major axis that function is flat: a sample anywhere in a wide arc returns nearly the maximum. Near the minor axis it is not. Expanding about the minimum gives a rise of a third within 0.88·σ₂/σ₁ radians, so the width of the notch is the reciprocal of the answer. At an axis ratio of 37 the notch is 1.4° wide, forty-eight points are 7.5° apart, and no sample lands in it.

The consequence is exactly what the arithmetic predicts and is larger than anybody would guess. Under the CIE RGB primaries, one ellipse has a true axis ratio of 40.76 and the forty-eight-point sample reports 18.59. Averaged over all twenty-five ellipses that basis reads 6.08 where it is 8.93; the sRGB primaries read 9.84 where they are 12.45; and CAT16’s cone responses read 2.69 where they are 2.71, which is right to one per cent.

That last row is the dangerous one. Every basis whose ellipses are nearly circular was measured correctly, so the table’s plausible rows are plausible, and the rows that were wrong are the ones the argument was about.

How many points a ratio needs is set by the ratio. Twelve curves, one per coordinate system this collection draws, each showing the axis ratio a sample of n points reports as a fraction of the exact value, against n on a logarithmic axis. Every curve rises to one, and they reach it at wildly different places: the systems whose ellipses are nearly circles are right at twenty-four points, and the CIE RGB primaries — whose worst ellipse has an axis ratio of 9 — are still 32 per cent short at forty-eight. The curve a measurement is on cannot be known until the answer is, which is what makes a fixed sample size the wrong instrument for this question.
Fig. 3 The ratio a sample reports as a fraction of the true one, against the sample size, for every coordinate system this collection draws. They reach one at wildly different places.

The replacement needs no sample. An ellipse is a quadratic form, the map is differentiable, and the image’s semi-axes are the singular values of the derivative composed with the ellipse’s own shape matrix. Two by two, in closed form, exact, and — the part that mattered downstream — a smooth function of the basis rather than a maximum over a fixed set of points.

The second case: a bowl sampled in random directions

The same collection minimises two things over the nine numbers a colour match leaves free, and having found each optimum, asked what the optimum looks like around itself. The method was twenty-four random unit directions in the nine coefficients, each walked outward by bisection until the cost had risen by five per cent, with the ratio of the longest radius to the shortest reported as the shape of the bowl.

It came back as 8.0 for one objective and 5.9 for the other, and those numbers were used to explain why one constraint of six parameters costs seventy per cent and another of three costs one.

The explanation was right and the number was not. A random unit vector in nine dimensions has an expected squared component of one ninth on every direction, so it carries a share of the stiffest direction whatever else it does. The curvature it reports is close to the average of the nine eigenvalues, and the radius it reports is close to the radius that average implies. It cannot report an extreme because it is never near one, and more samples do not help: the chance of landing within ten per cent of a given axis in nine dimensions falls off as a power of the dimension.

Taken from the eigenvalues, which needs no walking at all, the ratio is 29.8 and 27.2. The sample was short by 3.7 and 4.6.

And the sample added nothing. The same six eigenvalues predict each of the twenty-four measured radii individually, to a median of 7.6 per cent at a five per cent rise and 3.4 per cent at a rise a hundred times smaller — with nothing fitted. So the twenty-four bisections were an expensive way of confirming a quadratic form that was already there to be computed.

There is a control worth having, because the obvious objection to all of this is that the bowl is simply not a bowl five per cent above its floor, and that the sample and the model would agree closer in. Shrinking the rise by a hundred does make the model’s per-direction error fall by six — and moves the sampled ratio from 8.0 to 8.8, against a true 29.8. The loss is the dimension, and it is not recovered by measuring more carefully.

The third case: a worst case chosen from a list

The third is the one that does not have a right answer at the end of it, which is why it is the most interesting.

This collection scores an adaptation model against a census of illumination changes: fourteen named changes — two daylights, a thermal radiator, three discharge lamps, a bounce off a painted wall and a second bounce off the same wall, the eye’s own two filters. The worst row of that census is quoted as the worst case an adapted observer meets, and has been for several rounds of argument.

The fourteen are a sample of a family. A painted wall is a reflectance with a centre wavelength, a width, a depth and a base; a room applies it once or twice or three times. Searching that family rather than listing four members of it, and holding the parameters inside the range an ordinary architectural pigment occupies, the worst change reaches 21.3 ΔE00 against the census’s 3.375 — a factor of 6.3.

But the search does not come to rest anywhere interesting. It ends against the wall of the box in the centre wavelength and the width, and under a wider box it ends against three walls and returns 28.4. What that does to the ranking of the transforms it scores is a separate result and a larger one. The residual rises monotonically towards a narrower notch, at a shorter wavelength, on a darker base. There is no worst case in this family. The number anybody quotes for one is a number about the constraint they did not write down.

That is a different kind of failure from the first two, and it is a more useful one. The first two had exact replacements. This one has a correction to the question: a worst case is only meaningful relative to a stated family, and a discipline that reports one without stating the family has reported a property of its own list.

What was computed, and how

Each of the three has an assertion attached, and each assertion is built so that it would fail if somebody quietly reinstated the sampled version.

For the ellipses, the analytic Jacobian of the map from a chromaticity to a lightness–chroma space is written out rather than differenced, and is checked against central differences at every one of the twenty-five ellipse centres under twelve coordinate systems: worst relative discrepancy 2.7 × 10⁻⁸, which is a central difference’s accuracy on a cube root. The closed-form axes are checked against a Jacobi singular value decomposition of the same 2×2 on every one of those three hundred ellipses, agreeing to 10⁻¹³ at a largest axis ratio of 40.8 — the check that matters, because taking the axes from the Gram matrix squares the condition number, and this is the measurement of how much that costs at these numbers.

The claim that the two estimators are measuring the same thing is checked by shrinking the ellipse. A very finely sampled ring at full size sits 0.65 to 1.54 per cent above the analytic answer, and at a sixteenth of the size it sits within 0.09 per cent. That gap at full size is not error: it is the second-order distortion of the map across a real MacAdam ellipse, which is a third quantity worth knowing and is taken up on its own.

Three numbers for one set of ellipses, and which of them is which. Two curves and a horizontal line, against the size the ellipses are drawn at. The line is the analytic axis ratio — the ratio of the singular values of the map's own derivative, which is what "does this space make discrimination contours circles" means. The upper curve is a very finely sampled ring, which sits 0.6 per cent above the line at full size and converges onto it as the ellipse shrinks, because the gap between them is the second-order distortion of the map across a real ellipse rather than an error. The lower curve is the forty-eight-point sample used for this until now: it does not converge onto anything, because its error is set by the sample and not by the size.
Fig. 4 Three numbers for one set of ellipses: a coarse sample, a fine one, and the derivative — with the fine one converging onto the derivative as the ellipse is shrunk and the coarse one converging onto nothing.

The other three views of the same machinery say how much of the ranking this sampling error is capable of moving, which is the question a reader is really asking.

Twenty-five ellipses is a sample, and the score has an error bar. One row per colour space this collection ranks: the mean axis ratio its ellipses come out at, with the standard error of that mean over the twenty-five ellipses it was computed from. No literature is quoted — a mean of twenty-five numbers has a standard error those twenty-five numbers determine. The bars are far from equal: the best space carries ± 0.07 and the worst ± 1.56, because a space that makes the ellipses nearly circular makes all of them nearly circular and one that does not is dominated by whichever ellipse it handles worst.
Fig. 5 Each basis’s score with the standard error of that score over the twenty-five ellipses. The sampling error inside one ellipse is one term; the sample of ellipses is the other, and they are not the same size.
Three of the seven adjacent pairs in the table are actually ordered. One bar per adjacent pair of the uniformity ranking: the difference between the two spaces' scores divided by the standard error of that difference over the twenty-five ellipses. The comparison is paired — the same ellipses score both spaces, so an ellipse that is hard for everybody cancels — which is why the table is more informative than it looks and why treating the two errors as independent would have declared almost nothing ordered. 3 of the 7 pairs clear two; the rest do not, and one of them crosses zero in 35 per cent of resamples. The ranking's ends are real and its middle is not a ranking.
Fig. 6 Each adjacent pair of the ranking as its gap divided by the standard error of that gap. Two rows separated by less than one of these are two rows the data does not order.

Which of the two sampling errors dominates is settled by the third view, and it is the one a reader should be told about rather than the one that is easiest to compute.

How wrong the ellipses would have to be for a pair to change places. One bar per adjacent pair in the uniformity table: the relative error on each ellipse's own axes at which that pair changes places in one draw in twenty. No error on the data is quoted anywhere — the question is inverted, so what is reported is how large an error would have to be, and a reader with an opinion about MacAdam's experiment can compare it with their own number. The nearest pair goes at 0.171; 2 of the 7 pairs do not reverse under any error this search covers.
Fig. 7 And how large an error on the ellipses themselves would have to be to reverse each step. That is the honest form of the question, and the answer is not the same for every pair.

For the bowl, the Hessian is taken by central differences on both indices, 154 evaluations for a 9×9, at a step chosen by measuring rather than by taste: over a decade of steps the six eigenvalues the objective has move by 0.42 per cent, and the three it does not have grow by a factor of a hundred for a step ten times larger, which is the square law a truncation error obeys on a quantity that is exactly zero.

For the scene search, the objective is the same residual the census uses, and the search is a simplex with restarts over the four continuous parameters, with the bounce count taken separately because it is an integer and a simplex would interpolate it. The finding is reported as which parameters finished against a bound rather than as a number, because the number is a property of the bound.

Where the ranking moved, and where it did not

The correction to the ellipse measurement changed published numbers on this site, and it is worth being precise about which.

The floors barely moved. The best lightness–chroma space built on any basis leaves a mean axis ratio of 1.611 against a previously reported 1.620, and the best space with no compression in it leaves 2.334 against 2.36. Both are optima, both sit where the ellipses are nearly circular, and the sampled estimator was accurate there.

The published bases moved a great deal, in the direction of looking worse. The adaptation optimum’s anisotropy goes from 6.78 to 7.70; the sRGB primaries from 9.84 to 12.45; CIE RGB from 6.08 to 8.93. The ordering of the eight entries in this collection’s table of bases is unchanged, which is a relief and is not an accident — the error is monotone in the quantity, so it compresses a ranking without reversing it.

One caption did not survive. A figure in this collection said the receptor basis was the only entry in the table within a factor of two of both floors. It never was: Hunt–Pointer–Estévez is within 1.63 and 1.75, and CAT16 within 1.35 and 1.68, on the numbers as they stood at the time and on the corrected ones. That sentence had no assertion pointed at it, which is the only reason it lasted — and it is the same shape of defect as the three measurements above, one level up. A claim in prose that nothing tests is a sample of size zero.

Where the model stops

None of this says that sampling is the wrong tool. It is the right tool for a mean, an integral, a distribution, a typical value. It is the wrong tool for the largest member of a set whose largest members are rare, and the three cases here are all of that kind. Nothing about the arithmetic changes if the sample is well designed rather than uniform — a stratified sample or a low-discrepancy sequence has the same problem, because the problem is that the peak is narrow rather than that the points are badly placed.

The exact replacements are exact about something slightly different. The analytic ellipse axes are the axes of the differential of the map, which is the local metric of the space and not the finite ellipse’s image. The eigenvalues describe a quadratic bowl and the objective is not exactly quadratic. Both differences are measured here rather than waved away, and both are small compared with what the samples were missing.

And the third case has no replacement at all. Nothing here computes the worst change of light, because there is not one. What exists is a statement about a family and a bound, which is less satisfying and is what the model supports.

The generalisation

The pattern generalises past colour and past optimisation, and it is worth stating in the form that makes it checkable rather than the form that makes it obvious.

Before reporting a maximum, ask how wide the region near it is, and compare that width with the spacing of the sample. In every case here the width was computable in advance: 0.88·σ₂/σ₁ radians for the ellipse, one over the square root of the condition number for the bowl, and — for the list of scenes — not computable at all, which was itself the finding.

The reason this is skipped so reliably is that a sampled maximum has no error bar. A sampled mean announces its uncertainty: take another sample and it moves, and the movement is the estimate of the error. A sampled maximum is stable — take another forty-eight points and the answer barely changes, because the same wide arcs are being sampled and the same narrow notch is being missed. It is reproducible and wrong, which is the combination that survives review.

The same shape appears in a fit whose residual is flat: a quantity that does not move when the input is disturbed reads as robust, and is sometimes evidence that nothing is being measured. Stability is not accuracy, and the two are confused most often exactly where the check is cheapest.

Who found it, and when

The arithmetic is Cauchy’s and Rayleigh’s and is not in dispute anywhere. That the axes of a linearly transformed ellipse are the singular values of the transformation is nineteenth-century linear algebra; that the level sets of a quadratic form are ellipsoids with semi-axes proportional to the inverse square roots of the eigenvalues is the same century. Neither is difficult, and neither is what went wrong.

What went wrong is that in each case an operational description of a quantity — walk round the ellipse and see how long it is, step away from the optimum and see how far a step reaches — was implemented directly, and an operational description is a recipe rather than a definition. The recipe is correct in the limit and the limit is where nobody runs it.

The extreme-value literature has had the mathematics of this since Fisher and Tippett in 1928, and the practical form of it is stated most often in the reliability and hydrology literatures, where the cost of underestimating a maximum from a short record is a bridge. The colour literature does not appear to have a version of it, which is unsurprising: the sampled recipe is what a person with graph paper would have done, and once a number is in a table it is a number.

Where the ladder goes next

Two of the three sampled extremes had exact replacements and the third did not, and the difference between them is the useful distinction. In the first two the set being searched was a known shape — an ellipse, a quadratic bowl — and the extremum was available in closed form. In the third the set was a modelling family whose boundary somebody chose.

That suggests the question to ask of any reported worst case: is the set it was taken over a shape or a list? If it is a shape, compute; if it is a list, the number is about the list. What the shape of an optimum turns out to contain is the next step, and it starts with the surprise that three of its nine directions are not there at all.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 14 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyChromatic adaptationCondition numberConvergenceDegrees of freedomEigenvalueExtremumMacAdam's ellipsesQuadratic formSampling