Where the model breaks

A mean is not a worst case

Every adaptation number this collection publishes is an average over objects, and the reader asking whether adaptation will fail them is asking about the object it fails on. That object costs between 1.9 and 4.0 times the published figure, and how uneven a change of light is across objects turns out to be a property of the change rather than a constant.

Assumes A mean has a set under it, The worst case is where the box stops and An extremum is not a sample.

A mean answers what happens to an average object in an average room. Almost nobody asks that. The question a specification asks, a manufacturer asks and a complaint asks is what happens to the object it goes wrong on.

A published residual is a mean, and the worst object in the room costs twice it. Three bars for each of the 14 changes of light in the adaptation census, ordered by how uneven the change is across surfaces. The first bar is the published mean residual. The second is the worst single surface in the audit's published test set. The third is the worst surface anywhere in the region that set is drawn from, found by search rather than by reading a maximum off a lattice. The mean-to-worst ratio runs from 1.90 to 4.02 and averages 2.43, so every published adaptation number has a worst case about twice it that no essay had ever quoted. The gap between the second and third bars is the other finding: a maximum over 125 sampled points understates the region's own maximum by up to 34 per cent.
Fig. 1 For each change of light in the adaptation census: the published mean, the worst surface in the site’s own test set, and the worst surface anywhere in the region that set is drawn from. The third bar is between 1.9 and 4.0 times the first.

The claim

Every published adaptation residual has a worst case about twice it, the factor varies by more than two across the census, and the variation is informative rather than noise.

  • The mean-to-worst ratio runs from 1.90 to 4.02, averaging 2.43.
  • The evenest change of light is a green wall bounced twice, at 1.90 — which is also the harshest in absolute terms.
  • The least even is the macular pigment, at 4.02 — which is nearly the mildest in absolute terms.
  • But the two rankings are not inverted, which the section below had to be written to say. Across the fourteen rows the rank correlation between the mean and the ratio is −0.19, and −0.05 once the macular row is set aside — so the apparent ordering is carried by a single row, and there is no general rule connecting how bad a change of light is with how uneven it is.
  • And the test set’s own worst member is not the worst surface, by up to 34 per cent, which is an old finding in a new place.

Why a mean was the right thing to publish

Not a criticism to open with, because the mean is defensible and was chosen for a reason.

A residual after adaptation is defined for a change of light and a surface. Reporting it per surface would be a table of 125 numbers per row, which is not a publishable object; reporting a mean summarises it in the way that survives sampling and refinement best, and it is the number that answers what does adaptation do in a room.

It is also the statistic every published treatment of chromatic adaptation reports, for the same reasons.

What was missing was the second number, not a different first one. A mean and a worst case are two answers to two questions, both cheap, and the collection had been publishing one of them and letting it be read as both.

What one change of light costs, surface by surface — daylight to a triphosphor tubeA rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a triphosphor tube is 2.323 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 4.786, which is 2.06 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.0.01.83.65.4the published mean, 2.323worst in the set, 4.7865 surfaces at exactly zerothe surfaces, sorted by what this change of light costs themΔE₀₀daylight to a triphosphor tubeCIE 1931 2° observer · the set, varied
Fig. 2 One row’s whole distribution. The published number is the horizontal line; the worst case is the right-hand end of the curve, and the search finds a surface beyond even that because the curve is a lattice’s sample of a continuum.

The measurement

The worst surface is found by search over the region the test set is drawn from, rather than by taking the largest member of the set — for reasons the next section gives. Five starts, at the region’s centre and at the four corners of its modulation square, because the residual is largest where the surface is most saturated and there are four separate places to be saturated.

change of light mean worst in the set worst in the region ratio
the macular pigment 0.368 1.104 1.478 4.02
daylight to a three-primary display 0.980 2.705 2.928 2.99
daylight to a white LED 1.697 4.083 4.675 2.75
a triphosphor tube 2.323 4.786 5.884 2.53
daylight to halogen 1.583 3.169 3.912 2.47
daylight to D100 0.508 1.082 1.230 2.42
an older lens 1.117 2.363 2.588 2.32
daylight to D50 0.459 0.948 1.026 2.24
daylight to tungsten 1.635 3.058 3.652 2.23
a red wall 1.357 2.618 2.837 2.14
daylight to D40 0.980 1.931 2.086 2.13
a green wall, one bounce 1.722 3.148 3.461 2.01
daylight to a blackbody 0.263 0.457 0.502 1.91
a green wall, two bounces 3.375 5.786 6.428 1.90

Every row is at least 1.9. There is no change of light in this collection for which the published number is within a factor of two of what an unlucky object experiences.

Why the ratio varies, and what it means

The column is not scatter. It correlates with the kind of change, and the correlation runs the opposite way from the residual itself.

A broad multiplicative change is even. A green wall multiplies the light by a smooth reflectance across the whole band, so every surface is reweighted a little and no surface is spared. Both wall rows sit at 1.90 and 2.01, the two lowest ratios in the table, and the two-bounce row is the harshest change of light in the census. It is bad for everything and it has no particular victim.

A sharp spectral change is uneven. The macular pigment is a narrow absorption in the short-wave region inside the observer. A surface with no short-wave structure barely notices it; a surface whose modulation peaks there is hit four times harder than average. The ratio is 4.02 and the mean is 0.368 — nearly the mildest row in the census.

And the tempting sentence does not survive being measured. Two rows invite it — the wall at 1.90 with the largest mean, the macular pigment at 4.02 with nearly the smallest — and the pattern they suggest is that mild changes have victims and harsh ones do not. Over all fourteen rows the rank correlation between the mean and the ratio is −0.19, which on fourteen points is nothing; set the macular row aside and it is −0.05.

The counterexample is in the table and it is the mildest row in the census. A blackbody at daylight’s own temperature has a mean of 0.263 and a ratio of 1.91 — second-lowest in the column, sitting beside the harshest row’s 1.90. The two ends of the severity ranking have almost identical evenness, so whatever the ratio measures, it is not the inverse of the mean.

What survives is the mechanism without the correlation. A broad multiplicative change is even, and both wall rows are at the bottom of the column. A sharp spectral change is uneven, and the macular pigment is at the top of it by a wide margin. Both claims are about the shape of a change of light, and the mean is not a proxy for that shape: a blackbody shift is mild and smooth, a wall bounce is harsh and smooth, and both come out even. Ranking by the mean still ranks how much an average object is disturbed, and ranking by the worst case still ranks how much the unluckiest object is — but knowing one says very little about the other, which is a weaker and more useful statement than the one this section originally made.

What one change of light costs, surface by surface — the macular pigment. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for the macular pigment is 0.368 ΔE₀₀. The curve runs from 7.0e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 1.104, which is 3.00 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 3 The macular pigment’s costs across the set, sorted. Mean 0.368, worst in the set 1.104: the most skewed distribution in the census, and the reason its worst case is four times its average.

The set’s own maximum is not the maximum

The middle column of the table is the largest value the 125-member lattice contains. The right column is the supremum over the region. They differ by 5 to 34 per cent, and the difference is not an accident of search.

A maximum read off a finite set of points is a lower bound on the maximum, and the shortfall grows with how narrow the peak is. Here the worst surfaces sit on the region’s boundary, and a lattice puts points on that boundary only where its grid happens to land — so the lattice’s best attempt at the corner is a grid point near the corner rather than the corner.

The worst offender is the macular row again, at 1.104 against 1.478 — short by a third — and it is short there for the same reason its ratio is high: the peak is narrow, and a narrow peak is what a lattice misses.

That is the same finding two rounds old, one level down. It was found then on three measurements that took extrema over samples of a set; it recurs here on the set of surfaces, which was not among the three. The general form is worth restating: any quantity reported as the largest member of a constructed collection is a lower bound, and the gap grows exactly where the quantity is most extreme.

Every worst surface sits on a number somebody typed. The region the test surfaces are drawn from, in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the requirement that the two depths sum to no more than 0.7. The 14 marked points are the worst surface for each change of light in the adaptation census, found by search over the whole region. Every one of them lies exactly on the diamond, and every one is also at the brightest level the region allows — both declared constraints active, on all 14 rows, with no interior maximum anywhere. That is the opposite of what bounding the wall gave: there the worst case turned over at a band width of six nanometres because a narrow band returns too little light, which is physics. Here the worst case is a reading of two numbers. The one constraint that is about the world — a paint's excitation purity may not exceed 0.6 — is slack everywhere: the most saturated surface the region admits reaches 0.459.
Fig. 4 Where the worst surfaces are: on the region’s boundary, all fourteen of them, at the point of deepest admissible modulation. A lattice has points near there and not on it.
No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for the macular pigment. The tallest bar is 2.40 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 94.4 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 5 Why the macular row’s ratio is 4.02: its contribution profile is the most concentrated in the census, so a minority of surfaces carry a large share and the worst of them is far above the mean.

What to publish

Both numbers, and the ratio. The mean answers what adaptation does; the worst case answers what a specification has to tolerate; the ratio says which of those two questions the row is really about. Three numbers instead of one, from arithmetic that was already being done.

And the ratio needs its own warning, now that it has been measured against the mean. It is comparative and it is stable under the region’s own uncertainty, which is why it is worth publishing — but it is not a summary of a row’s severity and cannot be substituted for one. Two rows at the same ratio can be an order of magnitude apart in what they actually cost, and this census contains that pair: a blackbody at 1.91 and a green wall bounced twice at 1.90, thirteen-fold apart in mean. A reader who took the ratio as a ranking would put those two together.

And the ratio is the one that is new. A worst case alone is a number a reader discounts, because a worst case over a constructed region is a statement about the region. A ratio is comparative and survives the region’s own uncertainty almost entirely — scaling the set’s saturation moves every row’s mean and its worst case together, so the ratio is the part of this measurement that does not depend on three numbers nobody declared.

Where the worst case comes from

Two questions the table raises and only one of them has a satisfying answer.

Which surface is worst is answered by the search: on every row it is the most deeply modulated surface the region admits, at the brightest level. The details differ — the direction of the modulation swings with the change of light, so a green wall’s worst surface and a red wall’s are nearly opposites — but the magnitude is always at the boundary.

How bad the worst case can be is not answered, because the region has no bound that is about the world. Every one of the fourteen worst surfaces sits exactly on two numbers this site declared, and the one constraint in the file that is about paint rather than about convenience never binds. That is its own finding and it means the right column above should be read as the worst case the declared region contains rather than as the worst case.

The number a specification would actually want

Neither end of the distribution is what a tolerance is written against, and the middle ground is available from the same arithmetic.

change of light mean median 90th percentile 95th worst in region
daylight to a blackbody 0.263 0.291 0.395 0.425 0.502
the macular pigment 0.368 0.367 0.616 0.728 1.478
daylight to D50 0.459 0.451 0.740 0.834 1.026
daylight to tungsten 1.635 1.731 2.410 2.656 3.652
a triphosphor tube 2.323 2.633 3.523 3.627 5.884
a green wall, two bounces 3.375 3.374 4.932 5.184 6.428

The ninety-fifth percentile is 1.5 to 2.0 times the mean, against 1.9 to 4.0 for the worst case — so it carries most of the tail’s message and none of its dependence on where the region stops.

That is why it is the number colour-management practice actually uses: an ICC profile is reported with a mean and a 95th-percentile ΔE, and printing conditions are specified the same way. The convention is not arbitrary and it is not a compromise; it is the largest statistic that is still a statistic about the distribution rather than about the set’s boundary.

It is not reported as this collection’s headline number, for one reason worth stating: a percentile over a lattice inherits the lattice’s quadrature bias more strongly than a mean does, because the tail is where a grid’s coverage is worst. The mean’s bias is one to three per cent; the 95th percentile’s has not been measured and would be larger.

Only the extreme tail is row-specific

Putting the two ratios side by side on the six rows where both are computed says something neither says alone.

change of light 95th ÷ mean worst ÷ mean
a blackbody 1.62 1.91
the macular pigment 1.98 4.02
daylight to D50 1.82 2.24
daylight to tungsten 1.62 2.23
a triphosphor tube 1.56 2.53
a green wall, two bounces 1.54 1.91
spread across the six ×1.29 ×2.11

The percentile ratio is nearly a constant and the worst-case ratio is not. Every row’s ninety-fifth percentile sits between 1.54 and 1.98 times its mean; the worst case ranges over twice that.

So these distributions are the same shape up to their ninety-fifth percentile and differ only past it. Whatever makes the macular row uneven is not a fat tail — up to the point where ninety-five objects in a hundred have been accounted for, it looks like every other row in the census. It is the last five per cent that is row-specific, and by a factor of two.

That is the strongest thing to say for the convention this essay half-defends, and also the sharpest thing against it. A ninety-fifth percentile is nearly a fixed multiple of the mean, so as a second number it carries almost nothing the first did not. What makes it safe is exactly what makes it uninformative: a statistic stable across fourteen unrelated changes of light is a statistic about distributions rather than about where somebody stopped a region.

And it locates the finding. The quantity that separates these rows is not the tail but the end of it, which is the part a lattice samples worst and a declared boundary decides. The number that varies across the census is the number least entitled to be trusted, and that is not a coincidence — it is the boundary showing through.

Where the model stops

The mean and the worst case are the two ends of a distribution and neither is what a tolerance usually wants, which is a quantile. A ninety-fifth percentile over the set would be a better specification number than either, is available from the same per-surface vector, and is not reported here because a percentile over a lattice inherits the lattice’s own quadrature bias more strongly than a mean does — the tail is where the rule is worst.

And the worst case is a worst case over surfaces at a fixed change of light. The genuinely worst thing a room can do involves choosing both, which is what the bound ladder does for the wall and what this does for the surface, and the two have never been optimised jointly.

The objects the average is over, and the region they come from. Two panels. On the left, eight of the 125 reflectance spectra in the test set, drawn as reflectance against wavelength from 380 to 780 nanometres — smooth, broad curves between about 0.02 and 0.9, with at most two gentle undulations each, because each is a level times a combination of two cosines. None of them has a narrow feature, because the family has no basis function that could make one. On the right, the region those surfaces come from, drawn in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the constraint that the two depths may not exceed 0.7 in sum, and 25 lattice points inside it. Five levels of each of those pairs is the whole test set. The square's four corners — the most saturated surfaces the two cosines could make — are outside the diamond and are not in the set at all.
Fig. 6 The region the search runs over. Its boundary is where every worst surface sits, and the boundary is two numbers rather than a physical limit.

Why the ratio is the transportable number

The column worth carrying out of this essay is the right-hand one, and the reason is that it survives what the other columns do not.

Scaling the test region’s saturation by a quarter moves every mean by about a fifth and every worst case by about the same, so the ratio moves by a few per cent — against tens of per cent for either number alone. Refining the lattice moves both together. Changing the construction rule to the clamped family moves both together on twelve of fourteen rows.

So the worst object is four times worse than the average one is a statement about the macular pigment, and the worst object costs 1.478 ΔE*₀₀ is a statement about where somebody stopped a region.

That is the same conclusion the quadrature argument reaches about levels and orderings and the same one the bound argument reaches about declarations, arrived at a third time from a third direction. Three independent routes to publish the comparative quantity is enough to make it the collection’s rule rather than an observation about one table.

Who found it, and when

The distinction is not subtle and is standard wherever tolerances are written: a mean absolute error and a maximum error are different specifications, and the ratio between them is what a process-capability index is about. The colour-management literature knows it — a profile is normally reported with both a mean and a ninety-fifth-percentile ΔE, and the ISO conditions specify both.

This collection had been reporting one of them for fourteen phases, and the reason it went unremarked is that the mean was correct. Nothing was wrong; something was missing, and a missing second number leaves no trace in any gate. It became askable at the moment the per-surface residuals were recovered for a completely different purpose, which is the third time this round that an instrument built for one question answered a better one.

No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for a green wall, two bounces. The tallest bar is 1.37 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 110.8 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 7 The evenest row in the census by the worst-to-mean ratio, at 1.90: a green wall bounced twice treats every surface badly, so its worst case is the closest to its average of anything in the table.
What one change of light costs, surface by surface — daylight to a blackbody. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a blackbody is 0.263 ΔE₀₀. The curve runs from 0.0e+0 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 0.457, which is 1.74 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 8 And the mildest change of light, a blackbody at daylight’s own temperature. Its ratio is 1.91 — nearly the same as the harshest row’s — which is why the ratio has to be read as a shape rather than as a severity.

Two questions, not one number

The essay’s practical claim is that a residual answers a question, and the two questions want different statistics — which is worth setting out as a pair rather than as a correction.

“How well does adaptation work?” wants the mean. It is the question a model is evaluated by, it is what every published treatment reports, and a mean is the right answer because it weights every object once and is stable under sampling and refinement.

“Will adaptation fail me?” wants the tail. It is the question a specification asks, a manufacturer asks and a complaint embodies, and the mean is a poor answer to it by a factor of two to four depending on the row.

The collection had been publishing the first and letting it be read as both, which is not a mistake in any number and is a mismatch between a statistic and its use. The repair is a second column, not a different first one, and the second column was available from arithmetic already being done.

The distinction generalises past adaptation. Any figure of merit averaged over a set of objects has the same two readings, and which one a reader takes depends on whether they are choosing a method or specifying a tolerance. A camera profile has exactly the same pair, with the same factor of about two between them.

Where the ladder goes next

The worst case over the set turns out to sit exactly on the region’s declared boundary on every row, and the one constraint in the file that comes from the world never binds at all.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 25 that link here.

The objects this essay names

Each one links to every other essay that touches it.

BoundChromatic adaptationColour differenceMeanReflectanceResidualSamplingTest setToleranceWorst case