Four steps the test set cannot order
Assumes The error on a gap is not the errors at its ends, The census is a construction too and Which changes of light pay for it.
A table of fourteen numbers printed to four decimal places asserts fourteen distinct values. Four of the thirteen steps between them are steps the evidence does not take.
The claim
The census’s ranking is firmly established at both ends and unresolved in the middle, and the unresolved middle is the group of ordinary indoor lamps.
- Nine of thirteen steps are resolved, seven of them at four standard errors or better and several at more than twenty. The remaining two sit at 3.7 and 2.0, and the second does not survive being asked the question the way this table invites it — which is the section below.
- Four are not. Daylight-to-D40 against a three-primary display, at t = 0.0; halogen against tungsten at 1.4; tungsten against a white LED at 0.8; a white LED against a green wall at 0.4.
- Three of the four are consecutive, so it is not four separate near-ties but one unordered block of three lamps with a wall attached.
- The block is halogen, tungsten and a white LED — which is to say, three of the four things a room is actually lit by.
- A different construction of the test set reverses exactly one pair, and it is the pair at t = 0.0.
What the census is asked
Fourteen changes of illumination, each scored by how much of it an adapted observer fails to remove. The number is what is left after a diagonal gain in CAT16’s basis, averaged over a set of surfaces, in ΔE*₀₀.
The table is read as a ranking. It is the evidence for statements like a coloured wall is worse than any lamp, and for the specific claim that the worst thing in the census is a green wall bounced twice. Those statements are about order, not about level, and order is what an error on a gap is for.
What is established
| step | gap | t |
|---|---|---|
| daylight to D100 → daylight to D40 | 0.4720 | 22.6 |
| a red wall → daylight to halogen | 0.2258 | 9.1 |
| a blackbody at 6500 K → the macular pigment | 0.1050 | 8.8 |
| a triphosphor tube → a green wall bounced twice | 1.0516 | 8.0 |
| a green wall → a triphosphor tube | 0.6015 | 5.3 |
| daylight to D50 → daylight to D100 | 0.0491 | 5.2 |
| the macular pigment → daylight to D50 | 0.0908 | 4.0 |
| an older lens → a red wall | 0.2404 | 3.7 |
| a three-primary display → an older lens | 0.1368 | 2.0 |
Nine steps, and the top four of them are not close. The two ends of the ranking are not in doubt at all. That a blackbody at daylight’s own temperature is the mildest change in the table, and a saturated wall bounced twice is the harshest, are conclusions no plausible variation of the test set touches — and nor does a different construction of the set entirely.
That matters more than the failures below, because it is what the census is for. The collection’s arguments about adaptation rest on the shape of this table at its extremes, and the shape is solid.
What is not
| step | gap | t |
|---|---|---|
| daylight to D40 → a three-primary display | 0.0004 | 0.0 |
| daylight to halogen → daylight to tungsten | 0.0522 | 1.4 |
| daylight to tungsten → a white LED | 0.0619 | 0.8 |
| a white LED → a green wall, one bounce | 0.0244 | 0.4 |
The first is a curiosity and the other three are the finding. These four are the steps the test set cannot order — the gaps are smaller than the noise on them, and no way of reading the table changes that. A fifth step is demoted further down for an unrelated reason, and keeping the two kinds apart is the point of that section.
Daylight-to-D40 and daylight-to-a-three-primary-display come out at 0.9796 and 0.9800. Four parts in ten thousand, printed as two different numbers to four figures, with a standard error on the gap of 0.034 — eighty-five times the gap itself. There is no sense in which the table orders them, and refining the test set or changing its measure swaps them, which is the outcome a t of 0.0 predicts.
The other three run together. Taking every pair inside the block rather than only the adjacent ones:
| pair | t |
|---|---|
| halogen vs tungsten | 1.4 |
| tungsten vs a white LED | 0.8 |
| halogen vs a white LED | 2.5 |
| a white LED vs a green wall | 0.4 |
| tungsten vs a green wall | 2.8 |
So {halogen, tungsten, a white LED} is an unordered triple — every internal comparison inside two and a half standard errors — with a green wall’s single bounce sitting inside it as well on two of the three comparisons. Four rows, one cloud.
Counting the comparisons, which this essay had not done
The section on provenance names Tukey’s caution — that a table inviting every pairwise comparison must be judged as every pairwise comparison — and then the table above applies a threshold appropriate to one. Applying the caution to this essay’s own numbers costs it a step.
Thirteen adjacent steps tested at the ordinary two-tailed five per cent gives about a two in three chance that at least one of thirteen true ties is called a difference. Holding the family-wise rate at five per cent instead raises the bar:
| how the question is asked | critical t | steps resolved |
|---|---|---|
| one comparison, chosen in advance | 1.96 | 9 |
| the thirteen adjacent steps | 2.89 | 8 |
| all ninety-one pairs the table invites | 3.46 | 8 |
One step falls out: a three-primary display against an older lens, at t = 2.0. Every other verdict is unchanged, which is the useful part — the eight strongest steps clear even the ninety-one-comparison bar, and the four already-unresolved ones are nowhere near any of these thresholds.
So the corrected report is eight steps and five ties, and the fifth tie joins the four for a different reason from theirs. The unordered block of lamps is unordered because the underlying quantities are genuinely close; the display-against-lens step is demoted because the table asks so many questions that a two-sigma answer to one of them is what chance produces.
That distinction is worth keeping rather than collapsing. A gap smaller than its own error is a statement about the world; a gap that fails a multiplicity correction is a statement about how the table is read. The first would not improve with a better instrument and the second would — a comparison specified in advance, on its own, at t = 2.0, is evidence, and the same comparison picked out of ninety-one is not.
And it sharpens what the previous section says about sharpening the instrument. Twenty-five times the surfaces would take the lens step from 2.0 to 10 and settle it honestly. The same twenty-five-fold set applied to the block of lamps would move t from 0.4 to 2, which is under the corrected bar and, as that section says, would be measuring the lattice by then anyway. The one step worth more surfaces is the one that is not in the block, which is the opposite of where a reader’s attention goes.
Why it is that block
Not chance, and worth an explanation because the explanation is the useful part.
The four rows in the cloud sit at 1.583, 1.635, 1.697 and 1.722 ΔE*₀₀ — a span of nine per cent. The table’s whole range is 0.263 to 3.375, a factor of thirteen. So the cloud is four points inside a two-hundredth of the table’s range, and the question is why four unrelated light sources land within nine per cent of one another.
The answer is that they are not unrelated. Halogen and tungsten are the same physics — a hot filament, one running hotter than the other — so their spectra differ by a smooth tilt and their residuals are near-identical by construction. A white LED is a blue emitter under a broad phosphor, which is a different mechanism producing a spectrum that is, at the resolution three cone channels have, a similar smooth warm-shifted continuum. And a green wall’s single bounce multiplies daylight by a broad reflectance, which is again a smooth reweighting.
All four are broad, smooth changes of light of about the same magnitude, and three cone channels cannot tell smooth changes apart very finely. That is not a defect of the test set. It is the finding, stated the other way round: adaptation deals with all four about equally badly, and asking which is worst is asking a question the difference between them cannot support.
The rows that are separated are separated by mechanism. The triphosphor tube has three narrow emission lines. The macular pigment is a narrow absorption inside the observer. D40 against D100 is a large move along the daylight locus. Each of those is doing something structurally different — the macular pigment is an absorption inside the observer, the tube is three lines — , and the census sorts them cleanly.
How to report it
The temptation is to sharpen the instrument until the block resolves, and it is worth saying why that is the wrong response.
More surfaces would not fix it, and a finer lattice would move the numbers while chasing them. The standard error falls as the square root of the count, so resolving a gap of 0.024 at t = 2 from t = 0.4 needs twenty-five times the set — three thousand surfaces. That is affordable. But the gap itself moves by a few per cent when the set’s construction changes, which is the same size as the gap, so a t of 2 obtained that way would be measuring the lattice rather than the lamps.
And a bigger set would not answer the question anybody has. The reason to want halogen ranked against a white LED is a practical one — which lamp to put in a room where colour matters — and a difference of 0.06 ΔE*₀₀ is far below any threshold at which a person notices anything. A ranking that resolved it would be a true statement of no use.
So the report is the honest one: nine steps, four rows unordered, and the unordered rows are equivalent for every purpose the ranking has. The table now prints its rows in ranked order with the four unresolved steps marked, which costs nothing and stops the four decimal places from making a claim.
Why nothing caught it
Because a table of means has no place to put the fact.
Every gate on this site checks the census: that a change of light is exactly a matrix on the family, that the white is a fixed point, that a dimmer is exactly a gain, that the worst listed row is not the worst possible one. Those are assertions about arithmetic and about extremes, and every one of them passes on a table whose middle is a cloud.
The specific assertion that came closest was the one that requires the ordering to survive perturbing the census’s own constructed constants — the wall’s centre wavelength, its width, the lens’s age. It found, a round ago, that the winner never changes and that a pair separated by six parts in a thousand does change places. That was the same finding, on a different perturbation, and it was recorded as a fact about two adaptation transforms rather than as a fact about how the table should be read. The generalisation — that a table with a resolved top and an unresolved middle is what a continuous measurement over similar candidates produces — was available then and was not drawn until the set itself was audited.
What a table should print instead
Four decimal places on fourteen numbers is a claim about fourteen distinct values, and the table can say what it knows without saying less.
Rank order with the unresolved steps marked. The rows stay in ranked order — that is what the table is for — and the steps the evidence does not take are drawn as ties rather than as steps. Eight steps and five ties reads as more informative than thirteen steps, because it is.
Significant figures set by the error rather than by the arithmetic. A residual of 0.9796 with a standard error of 0.03 has two significant figures, not four. Printing four is what made a gap of 0.0004 look like a difference; printing 0.98 makes it look like what it is.
And the error on the gap, not on the row. The rows’ own errors are the wrong comparison and are three times too pessimistic. A table of ranked rows with a column of gap errors beside the steps carries exactly the information a reader needs to make the comparisons the table invites.
None of the three is new practice. All three are what a table of measured quantities in any experimental field looks like, and this collection’s census had none of them because its numbers had never had an error attached.
Where the model stops
The four unresolved steps are unresolved by this test set. Two things could change that and neither is a bigger lattice.
A test set with narrow spectral features would separate them, because the differences between a white LED’s phosphor spectrum and a halogen’s continuum are at wavelength scales that a smooth three-parameter surface cannot interrogate. A fourth reflectance dimension does exactly that and moves these four rows apart by up to ninety per cent — which is a real result and is a result about a different set of surfaces, not a sharper measurement of this one.
And a different metric would reorder them, which is one of the structural choices no sweep reaches. ΔE*₀₀ averaged over surfaces is one summary; the worst surface rather than the mean is another, and it ranks the same fourteen rows differently. Neither is more correct. The census answers what it asks, and what it asks has a four-row tie in it.
What the resolved steps establish
Worth spending a paragraph on the nine, because an essay about four failures can leave the impression that the table is weak and it is not.
The largest step in the table is 22.6 standard errors. Daylight to D40 costs 0.472 ΔE*₀₀ more than daylight to D100, and no construction of the test set, no refinement of the lattice, no change of measure and no fourth dimension moves it. That is a fact about how far along the daylight locus a change goes, and it is the kind of finding the census exists for.
The two ends are secure at eight standard errors or better. A blackbody at daylight’s own temperature is the mildest change of light this collection models; a saturated wall bounced twice is the harshest, by a factor of thirteen over the mildest. Both survive every perturbation tried in this round.
And the count survives the correction where it matters. Seven of the nine originally-resolved steps are at four standard errors or better and clear even the ninety-one-comparison bar with room to spare; only the weakest was demoted. A multiplicity correction that removed half a table would be evidence the table had been over-read, and this one removes a thirteenth of it.
And the ordering by kind survives. Ocular filters and mild daylight shifts at the bottom, discharge sources and coloured walls at the top. The unordered block sits entirely inside one band of that ordering rather than crossing it, so every claim this collection makes about which sort of change adaptation handles worst is untouched by the four unresolved steps.
A table with eight established steps and five ties is a table that establishes a great deal. What it does not establish is which of three ordinary indoor lamps is worst, and that turns out to be a question with no answer rather than a question the table got wrong.
Who found it, and when
The statistical point is Fisher’s least-significant-difference and the caution about it is Tukey’s: a table of pairwise comparisons invites the reader to make every comparison and report the interesting ones, which is why the honest report is the whole matrix rather than the winning row.
Here it is a consequence of the previous rung. The paired error on a gap did not exist in this collection until the per-surface residuals were recovered, so there was no quantity to compare a gap against and no way to notice that four of them were smaller than their own noise. The table was not read too confidently; it was read with no error bar at all — because the terms of the mean had never been recovered, which is a different failure and a more common one — a printed number with no uncertainty beside it does not look uncertain.
Where the ladder goes next
Four steps are unresolved and one of them reverses under a different construction of the set. That the flag caught the reversal is worth having; that most flagged pairs never reverse, and that a pair the flag cleared comfortably reverses under a set built to another rule, is worth more.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The census in six units chromatic adaptation · colour difference · illuminant · residual · test set
- A partial correction is worth its fraction chromatic adaptation · colour difference · residual · test set
- A choice with no magnitude colour difference · residual · test set
- A dial through a discrete menu colour difference · residual · test set
- A fourth dimension has a shape chromatic adaptation · illuminant · test set
- An extremum is still not a sample chromatic adaptation · residual · test set
What links here
The 8 essays that link to this one and share the most of its objects, of 10 that link here.
The objects this essay names
Each one links to every other essay that touches it.
Chromatic adaptationColour differenceHalogenIlluminantLight-emitting diodeRankingResidualSignificanceStandard errorTest set