An extremum is not a sample
Assumes A fit can be exact and empty, A constraint costs what it points at and MacAdam measured it.
A mean survives being sampled. A maximum does not, and the difference is not a matter of degree.
The claim
Three measurements in this collection were made by walking round a set, or across one, and keeping the largest or smallest thing found. Every one of them is short, every one is short by more where the answer is larger, and in each case there was an exact method available that needed no sample at all.
- The axes of a discrimination ellipse, taken as the largest and smallest of forty-eight radii. Short by a factor of 2.19 on the worst ellipse this site draws, and right to a fraction of a per cent on the roundest.
- The shape of an optimum, taken from twenty-four random directions walked outward until the cost rose. It reported a bowl 8.0 times longer one way than another where the answer is 29.8.
- The worst change of light, taken as the largest of fourteen changes on a list. The family those fourteen were drawn from reaches six times further, and does not stop there.
- In each case the error is set by the answer, not by the care taken. A sample of
npoints resolves an axis ratio whennis a few times the ratio; a random direction in nine dimensions reports the mean curvature whatever the spread; a list of fourteen scenes reports the worst of fourteen scenes. - And in each case the replacement is arithmetic rather than a bigger sample. A quadratic form has singular values. A quadratic bowl has eigenvalues. A family with no interior maximum has no worst case to find, and saying so is the result.
Why a mean is safe and a maximum is not
The reason is old and is worth stating in one line before any of the three cases, because it is the whole of them.
A sample mean converges on the population mean at a rate that depends only on how many samples were taken and how much the quantity varies — not on where the samples fell. Every point contributes. A sample maximum converges on the population maximum at a rate that depends on how much of the set is near the maximum, and in every case here that fraction is small and gets smaller as the extremum gets more extreme. The peak is narrow because it is a peak.
So the failure is not that sampling is approximate. It is that the estimator is biased in a known direction, by an amount that grows with what it is trying to report, which is the worst possible arrangement: it is accurate when the answer is boring and wrong when the answer matters. A table of such numbers looks plausible everywhere and has its extreme rows quietly halved.
The first case: an ellipse walked round
MacAdam’s ellipses are the ruler this collection measures a colour space’s uniformity with. The question asked of a space is how nearly it makes them circles, and the obvious way to answer it is to take the twenty-five ellipses, map their boundaries into the space, and take the longest radius over the shortest.
That is what was done, at forty-eight points per ellipse, for as long as this site has been asking the question.
The image of a small ellipse under a differentiable map is, to first order, another ellipse — and the radius round an ellipse with semi-axes σ₁ ≥ σ₂ is √(σ₁²cos²p + σ₂²sin²p). Near the major axis that function is flat: a sample anywhere in a wide arc returns nearly the maximum. Near the minor axis it is not. Expanding about the minimum gives a rise of a third within 0.88·σ₂/σ₁ radians, so the width of the notch is the reciprocal of the answer. At an axis ratio of 37 the notch is 1.4° wide, forty-eight points are 7.5° apart, and no sample lands in it.
The consequence is exactly what the arithmetic predicts and is larger than anybody would guess. Under the CIE RGB primaries, one ellipse has a true axis ratio of 40.76 and the forty-eight-point sample reports 18.59. Averaged over all twenty-five ellipses that basis reads 6.08 where it is 8.93; the sRGB primaries read 9.84 where they are 12.45; and CAT16’s cone responses read 2.69 where they are 2.71, which is right to one per cent.
That last row is the dangerous one. Every basis whose ellipses are nearly circular was measured correctly, so the table’s plausible rows are plausible, and the rows that were wrong are the ones the argument was about.
The replacement needs no sample. An ellipse is a quadratic form, the map is differentiable, and the image’s semi-axes are the singular values of the derivative composed with the ellipse’s own shape matrix. Two by two, in closed form, exact, and — the part that mattered downstream — a smooth function of the basis rather than a maximum over a fixed set of points.
The second case: a bowl sampled in random directions
The same collection minimises two things over the nine numbers a colour match leaves free, and having found each optimum, asked what the optimum looks like around itself. The method was twenty-four random unit directions in the nine coefficients, each walked outward by bisection until the cost had risen by five per cent, with the ratio of the longest radius to the shortest reported as the shape of the bowl.
It came back as 8.0 for one objective and 5.9 for the other, and those numbers were used to explain why one constraint of six parameters costs seventy per cent and another of three costs one.
The explanation was right and the number was not. A random unit vector in nine dimensions has an expected squared component of one ninth on every direction, so it carries a share of the stiffest direction whatever else it does. The curvature it reports is close to the average of the nine eigenvalues, and the radius it reports is close to the radius that average implies. It cannot report an extreme because it is never near one, and more samples do not help: the chance of landing within ten per cent of a given axis in nine dimensions falls off as a power of the dimension.
Taken from the eigenvalues, which needs no walking at all, the ratio is 29.8 and 27.2. The sample was short by 3.7 and 4.6.
And the sample added nothing. The same six eigenvalues predict each of the twenty-four measured radii individually, to a median of 7.6 per cent at a five per cent rise and 3.4 per cent at a rise a hundred times smaller — with nothing fitted. So the twenty-four bisections were an expensive way of confirming a quadratic form that was already there to be computed.
There is a control worth having, because the obvious objection to all of this is that the bowl is simply not a bowl five per cent above its floor, and that the sample and the model would agree closer in. Shrinking the rise by a hundred does make the model’s per-direction error fall by six — and moves the sampled ratio from 8.0 to 8.8, against a true 29.8. The loss is the dimension, and it is not recovered by measuring more carefully.
The third case: a worst case chosen from a list
The third is the one that does not have a right answer at the end of it, which is why it is the most interesting.
This collection scores an adaptation model against a census of illumination changes: fourteen named changes — two daylights, a thermal radiator, three discharge lamps, a bounce off a painted wall and a second bounce off the same wall, the eye’s own two filters. The worst row of that census is quoted as the worst case an adapted observer meets, and has been for several rounds of argument.
The fourteen are a sample of a family. A painted wall is a reflectance with a centre wavelength, a width, a depth and a base; a room applies it once or twice or three times. Searching that family rather than listing four members of it, and holding the parameters inside the range an ordinary architectural pigment occupies, the worst change reaches 21.3 ΔE00 against the census’s 3.375 — a factor of 6.3.
But the search does not come to rest anywhere interesting. It ends against the wall of the box in the centre wavelength and the width, and under a wider box it ends against three walls and returns 28.4. What that does to the ranking of the transforms it scores is a separate result and a larger one. The residual rises monotonically towards a narrower notch, at a shorter wavelength, on a darker base. There is no worst case in this family. The number anybody quotes for one is a number about the constraint they did not write down.
That is a different kind of failure from the first two, and it is a more useful one. The first two had exact replacements. This one has a correction to the question: a worst case is only meaningful relative to a stated family, and a discipline that reports one without stating the family has reported a property of its own list.
What was computed, and how
Each of the three has an assertion attached, and each assertion is built so that it would fail if somebody quietly reinstated the sampled version.
For the ellipses, the analytic Jacobian of the map from a chromaticity to a lightness–chroma space is written out rather than differenced, and is checked against central differences at every one of the twenty-five ellipse centres under twelve coordinate systems: worst relative discrepancy 2.7 × 10⁻⁸, which is a central difference’s accuracy on a cube root. The closed-form axes are checked against a Jacobi singular value decomposition of the same 2×2 on every one of those three hundred ellipses, agreeing to 10⁻¹³ at a largest axis ratio of 40.8 — the check that matters, because taking the axes from the Gram matrix squares the condition number, and this is the measurement of how much that costs at these numbers.
The claim that the two estimators are measuring the same thing is checked by shrinking the ellipse. A very finely sampled ring at full size sits 0.65 to 1.54 per cent above the analytic answer, and at a sixteenth of the size it sits within 0.09 per cent. That gap at full size is not error: it is the second-order distortion of the map across a real MacAdam ellipse, which is a third quantity worth knowing and is taken up on its own.
The other three views of the same machinery say how much of the ranking this sampling error is capable of moving, which is the question a reader is really asking.
Which of the two sampling errors dominates is settled by the third view, and it is the one a reader should be told about rather than the one that is easiest to compute.
For the bowl, the Hessian is taken by central differences on both indices, 154 evaluations for a 9×9, at a step chosen by measuring rather than by taste: over a decade of steps the six eigenvalues the objective has move by 0.42 per cent, and the three it does not have grow by a factor of a hundred for a step ten times larger, which is the square law a truncation error obeys on a quantity that is exactly zero.
For the scene search, the objective is the same residual the census uses, and the search is a simplex with restarts over the four continuous parameters, with the bounce count taken separately because it is an integer and a simplex would interpolate it. The finding is reported as which parameters finished against a bound rather than as a number, because the number is a property of the bound.
Where the ranking moved, and where it did not
The correction to the ellipse measurement changed published numbers on this site, and it is worth being precise about which.
The floors barely moved. The best lightness–chroma space built on any basis leaves a mean axis ratio of 1.611 against a previously reported 1.620, and the best space with no compression in it leaves 2.334 against 2.36. Both are optima, both sit where the ellipses are nearly circular, and the sampled estimator was accurate there.
The published bases moved a great deal, in the direction of looking worse. The adaptation optimum’s anisotropy goes from 6.78 to 7.70; the sRGB primaries from 9.84 to 12.45; CIE RGB from 6.08 to 8.93. The ordering of the eight entries in this collection’s table of bases is unchanged, which is a relief and is not an accident — the error is monotone in the quantity, so it compresses a ranking without reversing it.
One caption did not survive. A figure in this collection said the receptor basis was the only entry in the table within a factor of two of both floors. It never was: Hunt–Pointer–Estévez is within 1.63 and 1.75, and CAT16 within 1.35 and 1.68, on the numbers as they stood at the time and on the corrected ones. That sentence had no assertion pointed at it, which is the only reason it lasted — and it is the same shape of defect as the three measurements above, one level up. A claim in prose that nothing tests is a sample of size zero.
Where the model stops
None of this says that sampling is the wrong tool. It is the right tool for a mean, an integral, a distribution, a typical value. It is the wrong tool for the largest member of a set whose largest members are rare, and the three cases here are all of that kind. Nothing about the arithmetic changes if the sample is well designed rather than uniform — a stratified sample or a low-discrepancy sequence has the same problem, because the problem is that the peak is narrow rather than that the points are badly placed.
The exact replacements are exact about something slightly different. The analytic ellipse axes are the axes of the differential of the map, which is the local metric of the space and not the finite ellipse’s image. The eigenvalues describe a quadratic bowl and the objective is not exactly quadratic. Both differences are measured here rather than waved away, and both are small compared with what the samples were missing.
And the third case has no replacement at all. Nothing here computes the worst change of light, because there is not one. What exists is a statement about a family and a bound, which is less satisfying and is what the model supports.
The generalisation
The pattern generalises past colour and past optimisation, and it is worth stating in the form that makes it checkable rather than the form that makes it obvious.
Before reporting a maximum, ask how wide the region near it is, and compare that width with the spacing of the sample. In every case here the width was computable in advance: 0.88·σ₂/σ₁ radians for the ellipse, one over the square root of the condition number for the bowl, and — for the list of scenes — not computable at all, which was itself the finding.
The reason this is skipped so reliably is that a sampled maximum has no error bar. A sampled mean announces its uncertainty: take another sample and it moves, and the movement is the estimate of the error. A sampled maximum is stable — take another forty-eight points and the answer barely changes, because the same wide arcs are being sampled and the same narrow notch is being missed. It is reproducible and wrong, which is the combination that survives review.
The same shape appears in a fit whose residual is flat: a quantity that does not move when the input is disturbed reads as robust, and is sometimes evidence that nothing is being measured. Stability is not accuracy, and the two are confused most often exactly where the check is cheapest.
Who found it, and when
The arithmetic is Cauchy’s and Rayleigh’s and is not in dispute anywhere. That the axes of a linearly transformed ellipse are the singular values of the transformation is nineteenth-century linear algebra; that the level sets of a quadratic form are ellipsoids with semi-axes proportional to the inverse square roots of the eigenvalues is the same century. Neither is difficult, and neither is what went wrong.
What went wrong is that in each case an operational description of a quantity — walk round the ellipse and see how long it is, step away from the optimum and see how far a step reaches — was implemented directly, and an operational description is a recipe rather than a definition. The recipe is correct in the limit and the limit is where nobody runs it.
The extreme-value literature has had the mathematics of this since Fisher and Tippett in 1928, and the practical form of it is stated most often in the reliability and hydrology literatures, where the cost of underestimating a maximum from a short record is a bridge. The colour literature does not appear to have a version of it, which is unsurprising: the sampled recipe is what a person with graph paper would have done, and once a number is in a table it is a number.
Where the ladder goes next
Two of the three sampled extremes had exact replacements and the third did not, and the difference between them is the useful distinction. In the first two the set being searched was a known shape — an ellipse, a quadratic bowl — and the extremum was available in closed form. In the third the set was a modelling family whose boundary somebody chose.
That suggests the question to ask of any reported worst case: is the set it was taken over a shape or a list? If it is a shape, compute; if it is a list, the number is about the list. What the shape of an optimum turns out to contain is the next step, and it starts with the surprise that three of its nine directions are not there at all.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- How far a quadratic can be believed anisotropy · chromatic adaptation · condition number · convergence · degrees of freedom · eigenvalue · quadratic form
- Only the flat directions keep their names anisotropy · chromatic adaptation · condition number · degrees of freedom · eigenvalue · quadratic form
- An ellipse is not a ring of points anisotropy · extremum · macadam's ellipses · quadratic form · sampling
- Downhill from a published matrix chromatic adaptation · condition number · degrees of freedom · eigenvalue · quadratic form
- How long is the bowl chromatic adaptation · condition number · degrees of freedom · eigenvalue · sampling
- How wrong would the data have to be anisotropy · convergence · macadam's ellipses · quadratic form · sampling
What links here
The 8 essays that link to this one and share the most of its objects, of 14 that link here.
The objects this essay names
Each one links to every other essay that touches it.
AnisotropyChromatic adaptationCondition numberConvergenceDegrees of freedomEigenvalueExtremumMacAdam's ellipsesQuadratic formSampling