The instrument named the pair that moved
Assumes Four steps the test set cannot order, A lattice is a quadrature rule and What would have to be wrong.
An error bar earns its place by being tested against something it did not compute. This one was, twice, and the two results are different.
The claim
Inside the question it is about, the standard error over the test set is a necessary condition for a reversal and not a sufficient one. Outside that question it is silent, and the silence is loud: a pair at nearly nine standard errors changes places under a set built to a different rule.
- Every reversal inside the region was flagged. One adjacent pair changed places under a finer lattice and under the uniform measure, and it was one of the four the error had marked.
- Most flagged pairs never reversed. Three of the four held under every construction, so the flag is a warning rather than a forecast.
- Under the clamped realistic family, two further pairs reversed, at t = 8.8 and t = 3.7 — neither flagged, neither flaggable.
- That is not the instrument failing. A jackknife over a set is a statistic about which members that set contains. It cannot see a set assembled by another rule, and it never claimed to.
- The failure available here was to read the error bars as a general robustness result, and it would have produced confidence about exactly the wrong rows.
The prediction
The standard error on a gap is computed from one thing only: how the 125 per-surface residuals vary within this set. It knows nothing about lattices, measures, dimensions or clamping. So a prediction made from it is a prediction from a genuinely independent quantity, and it is worth stating before the test rather than after:
If two census rows are separated by less than about twice the paired standard error on their gap, then a different construction of the test set may put them in the other order. If they are separated by much more, it should not.
The four flagged pairs were daylight-to-D40 against a three-primary display (t = 0.0), halogen against tungsten (1.4), tungsten against a white LED (0.8), and a white LED against a green wall (0.4). Everything else in the table was at 2.0 or above, the top of it above 20.
The test, inside the region
Three constructions, all of the same region of surfaces:
| construction | members | what it changes |
|---|---|---|
| coarse | 39 | a third the density |
| fine | 981 | eight times the density |
| uniform | 2,000 | the region’s own measure, drawn at random |
Every row’s level moves by four to five per cent across the three, and the ordering comes out identical to the lattice’s except in one place. Daylight-to-D40 and daylight-to-a-three-primary-display change places between the two coarser rules and the two finer ones.
That is the pair at t = 0.0. The instrument named it, out of thirteen adjacencies, and it is the one that moved.
And the other three flagged pairs held. Halogen stayed below tungsten, tungsten below the white LED, the white LED below the green wall, under every construction of the region. So the flag is not a prediction that a pair will reverse. It is a statement that the evidence does not forbid it — which is the correct and weaker claim, and it is the shape a robustness flag has to have.
This is the same retreat the collection made a round ago about a different rule. An earlier audit predicted that the comparisons which move under perturbation are the ones with noise on them, found the prediction necessary but not sufficient, and recorded the honest form rather than the memorable one. Two independent instruments, two rounds apart, arriving at the same logical shape is worth more than either arriving at it alone, because it suggests the shape is a property of this kind of question rather than of one construction.
The test, outside it
The fifth construction is the clamped realistic family: 120 smooth curves of varying slope and curvature, clipped into the physical range. It has been in the census’s own machinery since the round that built it, as the check that the idealisation is honest, and it is not the same region walked differently. The clamp makes it non-linear in its parameters, so a change of light is no longer exactly a matrix on it and every row answers a slightly different question.
It reverses three pairs. One is the flagged pair, again. The other two are not:
| pair | t on the lattice | what happens under the clamped set |
|---|---|---|
| a blackbody at 6500 K vs the macular pigment | 8.8 | change places |
| an older lens vs a red wall | 3.7 | change places |
A pair separated by nearly nine standard errors changed places. Nothing in the error bar was wrong; the error bar was answering a different question.
Why those two. Both involve the macular pigment or the lens — filters inside the observer with narrow, structured absorption — and both of the rows they swap with are broad, smooth changes of the lamp. Clamping the surfaces flattens them: a curve pushed against 0 or 1 loses modulation depth where it is clipped, and the loss is largest at the extremes of the spectrum, which is exactly where a macular filter does its work. So the macular row drops by 51.7 per cent under the clamped family, against 1.1 per cent for the blackbody row, and a gap of 0.105 with an error of 0.012 is annihilated by two rows moving by fifty per cent and one per cent respectively.
A statistic about membership has no term for that. The 125 per-surface residuals of the macular row vary among themselves in a way that says how much the mean depends on which of these surfaces are present. Whether a differently shaped surface would behave differently is not in the data at all.
What an error bar over a set is about
Stating it exactly is the point of the essay, because the loose version is the one that misleads.
It is about: how much the published mean depends on which members this set contains. Leave any one out, or any handful, and the mean moves by an amount this statistic quantifies. Two rows whose gap is small relative to it are two rows whose order could be changed by a different draw of members of this kind.
It is not about: whether the set is the right kind, which is a question about the rule rather than the members. Whether three basis functions is enough, whether the surfaces should be clamped, whether the region should reach further, whether a uniform measure over three parameters is the right weight. Every one of those is a question about the rule that made the set, and no statistic computed inside the set can answer a question about the rule.
The distinction has a standard name in survey work — sampling error against frame error — and the standard observation about it is that the second is usually larger and is never on the chart. That is precisely the case here. The sampling-error-analogue is a few per cent. The frame-error-analogue, measured by swapping the rule, is up to fifty-two per cent on one row.
What to publish instead
Not an error bar with a caveat attached, because a caveat under a chart does not stop the chart being read.
Two numbers, both computed, both cheap. The paired error, which says what the membership supports; and the row’s spread across constructions, which says what the rule is worth. The second is the larger of the two on eleven of fourteen rows and is larger by a factor of ten on the macular row.
And a rule for the ordering claims. An ordering that survives every construction, including the one built to a different rule, is a conclusion about changes of light. An ordering that survives only the same-region constructions is a conclusion about this family of surfaces and should say so. An ordering that survives neither is not an ordering.
By that rule the census’s ranking has three tiers, and the top and bottom of it — a blackbody at daylight’s temperature being mildest, a doubly bounced green wall being harshest — sit in the first.
Why this was worth testing at all
Because the alternative was to publish the error bars and stop, and the error bars would have been believed.
An error bar is the most persuasive object a technical figure can carry. It looks like a completed piece of statistics; it invites the reader to treat everything outside it as settled; and it says nothing about its own scope. The four flagged pairs would have been reported as the census’s soft spots, the remaining nine as established, and the macular row — moving by half under a change of rule — would have sat at the top of the table with a whisker a fortieth of its own instability.
The test that catches this is not a better statistic. It is a second construction — as a fourth dimension is for the theorem, and the second construction was already in the codebase for another purpose. That is the ordinary case: the instrument that checks an instrument is usually something already built, pointed somewhere else.
The three tiers, stated
The rule the essay arrives at is worth writing out as a classification, because it is what the census’s ranking should be read through from now on.
Tier one: survives every construction, including the one built to a different rule. A blackbody at daylight’s own temperature is the mildest change in the table; a green wall bounced twice is the harshest; the daylight rows are ordered by how far they move along the locus. These are conclusions about changes of light and can be stated without qualification.
Tier two: survives every same-region construction and not the clamped one. The macular pigment against a blackbody, and an older lens against a red wall. These are conclusions about this family of surfaces — true of smooth cosine-combination reflectances and not of clamped ones — and should be stated with the family named.
Tier three: survives nothing. Daylight-to-D40 against a three-primary display, at a gap of four parts in ten thousand. Not an ordering at all.
The classification costs one extra column and it separates the census’s genuine findings from its printing precision. Nine of the thirteen steps are tier one, two are tier two, and one is tier three — which is a better summary of what the table establishes than either “fourteen distinct values” or “the middle is unreliable”.
The tiers add to twelve
The classification is stated as nine of the thirteen steps are tier one, two are tier two, and one is tier three, and nine plus two plus one is twelve. One step of the thirteen has been left out of its own summary.
The arithmetic that fixes it is the classification’s own. Three adjacencies reverse anywhere: the D40-against-a-three-primary-display pair, which reverses under the same-region constructions and is therefore tier three; and the blackbody-against-macular and lens-against-red-wall pairs, which reverse only under the clamped family and are therefore tier two. Every other adjacency survives every construction tried. Ten are tier one, two are tier two, one is tier three, and the three counts now sum to the number of steps there are.
The correction makes the summary slightly stronger rather than weaker — ten of thirteen orderings can be stated without qualification instead of nine — which is the direction that made it easy to miss.
The reversal is not a near-tie tipping over
The blackbody-against-macular result is easy to picture as two rows a hair apart being nudged past each other by a change of rule. It is not, and the essay’s own numbers say so.
The blackbody row sits at 0.263 and the gap to the macular row is 0.105, so the macular row is at 0.368. Applying the two published movements — the macular row down 51.7 per cent, the blackbody row down 1.1 — puts them at 0.178 and 0.260. The gap has not closed to zero and re-opened by a whisker; it has reappeared at 0.082, the other way round, at 78 per cent of its original size.
That matters for how the result should be read. A marginal reversal would be a statement that the two rows are effectively tied and that any perturbation decides between them. This is a statement that the two rows are decisively ordered under both rules and ordered oppositely — which is a worse problem for the ranking and a cleaner one for the argument, because it cannot be dismissed as noise on a boundary.
It is also a check on the account of the mechanism. The two movements were computed independently of the gap, and multiplying them through reproduces a reversal of the size the essay reports, so the “clamping removes what a narrow filter acts on” story is arithmetically closed rather than merely plausible.
The two inside-the-region instruments are the same size
The reason the flag worked inside the region and only there becomes clear when the two magnitudes are put on one scale, which the essay never quite does.
The paired standard error is 0.034 on a pair of rows near 0.98, and 0.012 on a pair near 0.37 — about 3.4 per cent of a level in both cases. The spread across the three same-region constructions is quoted as four to five per cent of a level. Those are the same number.
Two instruments that share no arithmetic agree on the size of the wobble inside the region, which is why one predicts the other: a gap smaller than twice the sampling error is also a gap smaller than the construction spread, so the same pairs are at risk on both criteria. The flag is not lucky. It is measuring the same few per cent by a different route.
And it is exactly that agreement that makes the clamped result so much larger. Fifty-two per cent against four and a half is a factor of 11.5 — an order of magnitude outside the band both in-region instruments live in. The frame error is not a bit larger than the sampling error here; it is eleven times larger, and any pair whose gap is under half a level is exposed to it while looking perfectly safe on the chart.
That gives the practical threshold the essay’s three tiers imply without stating. A step is safe against a change of rule only if its gap survives one row moving by half. On the census that is a demanding test, and it is passed by the extremes — a blackbody at daylight’s own temperature against a doubly bounced green wall is a factor of thirteen apart — and by very little in the middle.
Where the model stops
Two constructions are not many. The clamped family is the only rule-change tested here, and it changes several things at once — the clamp, the parameterisation, the member count — so which of them causes the macular row’s collapse is not separated by this experiment. It is very likely the clamp, on the argument above, and very likely is not measured.
And nothing here bounds how bad a rule-change can be. Fifty-two per cent is what one alternative rule gives; another might give more. A bound would need a statement about which rules are admissible, which is the same shape of problem as bounding the surfaces themselves and is not solved by this round either.
Why the flag is still worth printing
An instrument that is necessary and not sufficient, and blind outside its own question, could reasonably be dropped. It should not be, and the reason is what it costs against what it catches.
It costs nothing. The paired standard error is one pass over per-surface residuals that are already computed for the mean. There is no second model, no second construction, no assumption added.
It caught the one thing inside its scope. One adjacency reversed under three different same-region constructions; the flag had named it, out of thirteen candidates, before any of the three was run. A cheap instrument with a true positive and no false negatives inside its domain is a good instrument.
And its silences are informative once its scope is stated. Three flagged pairs held under everything — so a flag is a do not rely on this step rather than a forecast. Two unflagged pairs reversed under a different rule — so an unflagged step is this set’s members do not threaten it rather than this is safe.
The failure available was never the instrument. It was reading an error bar as a general robustness statement, which is what an error bar looks like and is not what any error bar is. Printing the flag beside a second column that says what a change of rule does is the fix, and the second column is the one this round had to build.
Who found it, and when
The general principle is Deming’s and predates the statistics that hides it: the frame is the population an instrument can actually reach, and the difference between it and the population intended is nowhere in the standard errors. Every survey methodologist knows it and every table of survey results is read as though it were not true.
Here it arrived because the clamped family already existed. It was written into the census’s machinery at the round that built it, purely to check that the idealisation was honest, with an assertion comparing the two families’ orderings and nothing comparing their levels. Running it as a fifth column of a chart about something else is what turned a background check into this round’s sharpest result — which is the second time in two rounds that a helper written for one purpose has answered a better question than the one it was written for.
Where the ladder goes next
The rule that made the set has been changed once, by clamping. It can be changed in a way that reaches deeper than that: the family is exactly three-dimensional because three basis functions were written down, and the theorem the whole adaptation argument rests on is a theorem about three-dimensionality.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The error on a gap is not the errors at its ends chromatic adaptation · residual · sampling · significance · standard error · test set
- Two instruments and one ranking chromatic adaptation · residual · sampling · standard error · test set
- A mean has a set under it chromatic adaptation · residual · sampling · test set
- A mean is not a worst case chromatic adaptation · residual · sampling · test set
- An extremum is still not a sample chromatic adaptation · residual · sampling · test set
- A partial correction is worth its fraction chromatic adaptation · residual · test set
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Chromatic adaptationFalsificationPredictionRankingResidualRobustnessSamplingSignificanceStandard errorTest set