Where the model breaks

The instrument named the pair that moved

A standard error over a test set flagged four steps of the adaptation census as unresolved. Rebuilding the set three different ways reversed exactly one pair, and it was one of the four. Rebuilding it to a different rule reversed a pair the error separated by nearly nine standard errors — which is not a failure of the instrument but a statement of what it is about.

Assumes Four steps the test set cannot order, A lattice is a quadrature rule and What would have to be wrong.

An error bar earns its place by being tested against something it did not compute. This one was, twice, and the two results are different.

Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.
Fig. 1 The adaptation census under five constructions of the test set. Three of the five walk the same region of surfaces differently. The fourth is the region’s own uniform measure. The fifth is built to a different rule entirely, and it is the one that reverses a pair the error had cleared.

The claim

Inside the question it is about, the standard error over the test set is a necessary condition for a reversal and not a sufficient one. Outside that question it is silent, and the silence is loud: a pair at nearly nine standard errors changes places under a set built to a different rule.

  • Every reversal inside the region was flagged. One adjacent pair changed places under a finer lattice and under the uniform measure, and it was one of the four the error had marked.
  • Most flagged pairs never reversed. Three of the four held under every construction, so the flag is a warning rather than a forecast.
  • Under the clamped realistic family, two further pairs reversed, at t = 8.8 and t = 3.7 — neither flagged, neither flaggable.
  • That is not the instrument failing. A jackknife over a set is a statistic about which members that set contains. It cannot see a set assembled by another rule, and it never claimed to.
  • The failure available here was to read the error bars as a general robustness result, and it would have produced confidence about exactly the wrong rows.

The prediction

The standard error on a gap is computed from one thing only: how the 125 per-surface residuals vary within this set. It knows nothing about lattices, measures, dimensions or clamping. So a prediction made from it is a prediction from a genuinely independent quantity, and it is worth stating before the test rather than after:

If two census rows are separated by less than about twice the paired standard error on their gap, then a different construction of the test set may put them in the other order. If they are separated by much more, it should not.

The four flagged pairs were daylight-to-D40 against a three-primary display (t = 0.0), halogen against tungsten (1.4), tungsten against a white LED (0.8), and a white LED against a green wall (0.4). Everything else in the table was at 2.0 or above, the top of it above 20.

The test, inside the region

Three constructions, all of the same region of surfaces:

construction members what it changes
coarse 39 a third the density
fine 981 eight times the density
uniform 2,000 the region’s own measure, drawn at random

Every row’s level moves by four to five per cent across the three, and the ordering comes out identical to the lattice’s except in one place. Daylight-to-D40 and daylight-to-a-three-primary-display change places between the two coarser rules and the two finer ones.

That is the pair at t = 0.0. The instrument named it, out of thirteen adjacencies, and it is the one that moved.

And the other three flagged pairs held. Halogen stayed below tungsten, tungsten below the white LED, the white LED below the green wall, under every construction of the region. So the flag is not a prediction that a pair will reverse. It is a statement that the evidence does not forbid it — which is the correct and weaker claim, and it is the shape a robustness flag has to have.

This is the same retreat the collection made a round ago about a different rule. An earlier audit predicted that the comparisons which move under perturbation are the ones with noise on them, found the prediction necessary but not sufficient, and recorded the honest form rather than the memorable one. Two independent instruments, two rounds apart, arriving at the same logical shape is worth more than either arriving at it alone, because it suggests the shape is a property of this kind of question rather than of one construction.

Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.
Fig. 2 The thirteen steps with their paired errors. The one that reverses is the shortest bar in the picture — a gap of 0.0004 against an error of 0.034.
The objects the average is over, and the region they come from. Two panels. On the left, eight of the 125 reflectance spectra in the test set, drawn as reflectance against wavelength from 380 to 780 nanometres — smooth, broad curves between about 0.02 and 0.9, with at most two gentle undulations each, because each is a level times a combination of two cosines. None of them has a narrow feature, because the family has no basis function that could make one. On the right, the region those surfaces come from, drawn in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the constraint that the two depths may not exceed 0.7 in sum, and 25 lattice points inside it. Five levels of each of those pairs is the whole test set. The square's four corners — the most saturated surfaces the two cosines could make — are outside the diamond and are not in the set at all.
Fig. 3 The region three of the four other constructions walk differently — coarsely, finely, and at random against its own measure. The fifth construction is not in this picture at all, because it is built to a different rule.

The test, outside it

The fifth construction is the clamped realistic family: 120 smooth curves of varying slope and curvature, clipped into the physical range. It has been in the census’s own machinery since the round that built it, as the check that the idealisation is honest, and it is not the same region walked differently. The clamp makes it non-linear in its parameters, so a change of light is no longer exactly a matrix on it and every row answers a slightly different question.

It reverses three pairs. One is the flagged pair, again. The other two are not:

pair t on the lattice what happens under the clamped set
a blackbody at 6500 K vs the macular pigment 8.8 change places
an older lens vs a red wall 3.7 change places

A pair separated by nearly nine standard errors changed places. Nothing in the error bar was wrong; the error bar was answering a different question.

Why those two. Both involve the macular pigment or the lens — filters inside the observer with narrow, structured absorption — and both of the rows they swap with are broad, smooth changes of the lamp. Clamping the surfaces flattens them: a curve pushed against 0 or 1 loses modulation depth where it is clipped, and the loss is largest at the extremes of the spectrum, which is exactly where a macular filter does its work. So the macular row drops by 51.7 per cent under the clamped family, against 1.1 per cent for the blackbody row, and a gap of 0.105 with an error of 0.012 is annihilated by two rows moving by fifty per cent and one per cent respectively.

A statistic about membership has no term for that. The 125 per-surface residuals of the macular row vary among themselves in a way that says how much the mean depends on which of these surfaces are present. Whether a differently shaped surface would behave differently is not in the data at all.

What one change of light costs, surface by surface — the macular pigment. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for the macular pigment is 0.368 ΔE₀₀. The curve runs from 7.0e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 1.104, which is 3.00 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 4 The macular pigment’s cost, surface by surface, on the idealised family. The distribution is the most skewed in the census — the worst surface costs three times the mean — because a narrow absorption inside the eye reaches some spectra and not others. Clamping the surfaces removes most of what it reaches.
What one change of light costs, surface by surface — daylight to a blackbody. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a blackbody is 0.263 ΔE₀₀. The curve runs from 0.0e+0 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 0.457, which is 1.74 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 5 A blackbody at daylight’s own temperature: the mildest row in the census, at 0.263, and the row that changes places with the macular pigment under the clamped family despite being 8.8 standard errors away from it.

What an error bar over a set is about

Stating it exactly is the point of the essay, because the loose version is the one that misleads.

It is about: how much the published mean depends on which members this set contains. Leave any one out, or any handful, and the mean moves by an amount this statistic quantifies. Two rows whose gap is small relative to it are two rows whose order could be changed by a different draw of members of this kind.

It is not about: whether the set is the right kind, which is a question about the rule rather than the members. Whether three basis functions is enough, whether the surfaces should be clamped, whether the region should reach further, whether a uniform measure over three parameters is the right weight. Every one of those is a question about the rule that made the set, and no statistic computed inside the set can answer a question about the rule.

The distinction has a standard name in survey work — sampling error against frame error — and the standard observation about it is that the second is usually larger and is never on the chart. That is precisely the case here. The sampling-error-analogue is a few per cent. The frame-error-analogue, measured by swapping the rule, is up to fifty-two per cent on one row.

Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.
Fig. 6 The four flagged adjacencies, of which one reverses. Three flagged pairs hold under every construction tried, so the flag is necessary and not sufficient — which is the same shape a different audit reached a round ago on a different rule.

What to publish instead

Not an error bar with a caveat attached, because a caveat under a chart does not stop the chart being read.

Two numbers, both computed, both cheap. The paired error, which says what the membership supports; and the row’s spread across constructions, which says what the rule is worth. The second is the larger of the two on eleven of fourteen rows and is larger by a factor of ten on the macular row.

And a rule for the ordering claims. An ordering that survives every construction, including the one built to a different rule, is a conclusion about changes of light. An ordering that survives only the same-region constructions is a conclusion about this family of surfaces and should say so. An ordering that survives neither is not an ordering.

By that rule the census’s ranking has three tiers, and the top and bottom of it — a blackbody at daylight’s temperature being mildest, a doubly bounced green wall being harshest — sit in the first.

What one change of light costs, surface by surface — daylight to tungsten. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to tungsten is 1.635 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 3.058, which is 1.87 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 7 What the standard error is computed from: the 125 per-surface residuals of one row. Everything the instrument knows is in this picture, and nothing about a differently shaped surface is.

Why this was worth testing at all

Because the alternative was to publish the error bars and stop, and the error bars would have been believed.

An error bar is the most persuasive object a technical figure can carry. It looks like a completed piece of statistics; it invites the reader to treat everything outside it as settled; and it says nothing about its own scope. The four flagged pairs would have been reported as the census’s soft spots, the remaining nine as established, and the macular row — moving by half under a change of rule — would have sat at the top of the table with a whisker a fortieth of its own instability.

The test that catches this is not a better statistic. It is a second constructionas a fourth dimension is for the theorem, and the second construction was already in the codebase for another purpose. That is the ordinary case: the instrument that checks an instrument is usually something already built, pointed somewhere else.

The error on a gap is not the two rows' errors added. Two bars for each of the 13 adjacent pairs in the census ranking. The upper, shorter bar is the standard error of the gap taken as a paired difference — the same 125 surfaces score both rows, so a surface that is awkward under one change of light is usually awkward under the other and the difference is quieter than either. The lower bar is the two rows' own errors added in quadrature, which is what comparing error bars by eye amounts to. Pairing is worth a factor of 1.78 on average and 3.36 on the pair it helps most, and it is the difference between 6 adjacencies unordered and 4. The gain is largest where the two rows are two daylights or two tungstens, because then the surfaces they find awkward are nearly the same surfaces.
Fig. 8 The errors themselves, paired and unpaired. Sharpening the instrument by pairing does not widen what it is about — both versions are statistics over the same 125 members.

The three tiers, stated

The rule the essay arrives at is worth writing out as a classification, because it is what the census’s ranking should be read through from now on.

Tier one: survives every construction, including the one built to a different rule. A blackbody at daylight’s own temperature is the mildest change in the table; a green wall bounced twice is the harshest; the daylight rows are ordered by how far they move along the locus. These are conclusions about changes of light and can be stated without qualification.

Tier two: survives every same-region construction and not the clamped one. The macular pigment against a blackbody, and an older lens against a red wall. These are conclusions about this family of surfaces — true of smooth cosine-combination reflectances and not of clamped ones — and should be stated with the family named.

Tier three: survives nothing. Daylight-to-D40 against a three-primary display, at a gap of four parts in ten thousand. Not an ordering at all.

The classification costs one extra column and it separates the census’s genuine findings from its printing precision. Nine of the thirteen steps are tier one, two are tier two, and one is tier three — which is a better summary of what the table establishes than either “fourteen distinct values” or “the middle is unreliable”.

The tiers add to twelve

The classification is stated as nine of the thirteen steps are tier one, two are tier two, and one is tier three, and nine plus two plus one is twelve. One step of the thirteen has been left out of its own summary.

The arithmetic that fixes it is the classification’s own. Three adjacencies reverse anywhere: the D40-against-a-three-primary-display pair, which reverses under the same-region constructions and is therefore tier three; and the blackbody-against-macular and lens-against-red-wall pairs, which reverse only under the clamped family and are therefore tier two. Every other adjacency survives every construction tried. Ten are tier one, two are tier two, one is tier three, and the three counts now sum to the number of steps there are.

The correction makes the summary slightly stronger rather than weaker — ten of thirteen orderings can be stated without qualification instead of nine — which is the direction that made it easy to miss.

The reversal is not a near-tie tipping over

The blackbody-against-macular result is easy to picture as two rows a hair apart being nudged past each other by a change of rule. It is not, and the essay’s own numbers say so.

The blackbody row sits at 0.263 and the gap to the macular row is 0.105, so the macular row is at 0.368. Applying the two published movements — the macular row down 51.7 per cent, the blackbody row down 1.1 — puts them at 0.178 and 0.260. The gap has not closed to zero and re-opened by a whisker; it has reappeared at 0.082, the other way round, at 78 per cent of its original size.

That matters for how the result should be read. A marginal reversal would be a statement that the two rows are effectively tied and that any perturbation decides between them. This is a statement that the two rows are decisively ordered under both rules and ordered oppositely — which is a worse problem for the ranking and a cleaner one for the argument, because it cannot be dismissed as noise on a boundary.

It is also a check on the account of the mechanism. The two movements were computed independently of the gap, and multiplying them through reproduces a reversal of the size the essay reports, so the “clamping removes what a narrow filter acts on” story is arithmetically closed rather than merely plausible.

The two inside-the-region instruments are the same size

The reason the flag worked inside the region and only there becomes clear when the two magnitudes are put on one scale, which the essay never quite does.

The paired standard error is 0.034 on a pair of rows near 0.98, and 0.012 on a pair near 0.37 — about 3.4 per cent of a level in both cases. The spread across the three same-region constructions is quoted as four to five per cent of a level. Those are the same number.

Two instruments that share no arithmetic agree on the size of the wobble inside the region, which is why one predicts the other: a gap smaller than twice the sampling error is also a gap smaller than the construction spread, so the same pairs are at risk on both criteria. The flag is not lucky. It is measuring the same few per cent by a different route.

And it is exactly that agreement that makes the clamped result so much larger. Fifty-two per cent against four and a half is a factor of 11.5 — an order of magnitude outside the band both in-region instruments live in. The frame error is not a bit larger than the sampling error here; it is eleven times larger, and any pair whose gap is under half a level is exposed to it while looking perfectly safe on the chart.

That gives the practical threshold the essay’s three tiers imply without stating. A step is safe against a change of rule only if its gap survives one row moving by half. On the census that is a demanding test, and it is passed by the extremes — a blackbody at daylight’s own temperature against a doubly bounced green wall is a factor of thirteen apart — and by very little in the middle.

Where the model stops

Two constructions are not many. The clamped family is the only rule-change tested here, and it changes several things at once — the clamp, the parameterisation, the member count — so which of them causes the macular row’s collapse is not separated by this experiment. It is very likely the clamp, on the argument above, and very likely is not measured.

And nothing here bounds how bad a rule-change can be. Fifty-two per cent is what one alternative rule gives; another might give more. A bound would need a statement about which rules are admissible, which is the same shape of problem as bounding the surfaces themselves and is not solved by this round either.

No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for an older lens. The tallest bar is 1.69 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 101.1 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 9 An older lens, by contribution. This row changes places with a red wall under the clamped family at 3.7 standard errors, for the same reason the macular row does: both are narrow filters inside the observer, and clamping removes what they act on.

Why the flag is still worth printing

An instrument that is necessary and not sufficient, and blind outside its own question, could reasonably be dropped. It should not be, and the reason is what it costs against what it catches.

It costs nothing. The paired standard error is one pass over per-surface residuals that are already computed for the mean. There is no second model, no second construction, no assumption added.

It caught the one thing inside its scope. One adjacency reversed under three different same-region constructions; the flag had named it, out of thirteen candidates, before any of the three was run. A cheap instrument with a true positive and no false negatives inside its domain is a good instrument.

And its silences are informative once its scope is stated. Three flagged pairs held under everything — so a flag is a do not rely on this step rather than a forecast. Two unflagged pairs reversed under a different rule — so an unflagged step is this set’s members do not threaten it rather than this is safe.

The failure available was never the instrument. It was reading an error bar as a general robustness statement, which is what an error bar looks like and is not what any error bar is. Printing the flag beside a second column that says what a change of rule does is the fix, and the second column is the one this round had to build.

Who found it, and when

The general principle is Deming’s and predates the statistics that hides it: the frame is the population an instrument can actually reach, and the difference between it and the population intended is nowhere in the standard errors. Every survey methodologist knows it and every table of survey results is read as though it were not true.

Here it arrived because the clamped family already existed. It was written into the census’s machinery at the round that built it, purely to check that the idealisation was honest, with an assertion comparing the two families’ orderings and nothing comparing their levels. Running it as a fifth column of a chart about something else is what turned a background check into this round’s sharpest result — which is the second time in two rounds that a helper written for one purpose has answered a better question than the one it was written for.

Where the ladder goes next

The rule that made the set has been changed once, by clamping. It can be changed in a way that reaches deeper than that: the family is exactly three-dimensional because three basis functions were written down, and the theorem the whole adaptation argument rests on is a theorem about three-dimensionality.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Chromatic adaptationFalsificationPredictionRankingResidualRobustnessSamplingSignificanceStandard errorTest set