The error on a gap is not the errors at its ends
Assumes A lattice is a quadrature rule, A mean has a set under it and Twenty-five is a sample of the diagram.
Draw a ranking with an error bar on each row and a reader will compare the bars. That comparison is not the one the table is asking about, and on this collection’s own tables it is between one and a half and three and a half times too pessimistic.
The claim
Two rows of the adaptation census are two numbers computed over the same set of surfaces. The error on their difference is therefore not the two errors combined, and treating it as though it were declares half again as much of the table unordered as is actually unordered.
- The gain from pairing averages 1.78 across the thirteen adjacencies and reaches 3.36 on the pair it helps most.
- It is the difference between four adjacencies unresolved and six.
- The size of the gain is itself a measurement. It is largest where the two changes of light are alike — two daylights, a tungsten against a halogen — because then the surfaces one finds awkward are the surfaces the other finds awkward.
- This is the second time the same arithmetic has been needed in two rounds, on two different tables, for the same reason, and the gain is smaller here for a reason worth stating.
What the wrong comparison is
Each census row is a mean over 125 surfaces, and a mean over a set has a standard error — with the caveat about what a standard error over a lattice licenses, which does not affect this argument at all because both rows carry it equally.
Write the two rows as A and B, with per-surface values aᵢ and bᵢ. The question the table asks is whether mean(b) − mean(a) is greater than zero. Its variance is var(b̄ − ā) = var(ā) + var(b̄) − 2 cov(ā, b̄), and the third term is the whole essay. Adding the two errors in quadrature — which is what the eye does with two error bars — is exactly this expression with the covariance dropped. Dropping it is right when the two quantities were measured independently. Here they were measured on the same objects, one after the other, in the same loop.
And the covariance is large, because a surface that is awkward under one change of light is usually awkward under another. A deeply modulated reflectance has spectral structure the observer’s three channels cannot follow, and that is true whichever lamp is swapped in. So aᵢ and bᵢ rise and fall together across the set, the difference bᵢ − aᵢ is far quieter than either, and its standard error is correspondingly smaller.
What it is worth
| adjacent pair | gap | paired error | unpaired | gain |
|---|---|---|---|---|
| a red wall → daylight to halogen | 0.2258 | 0.0249 | 0.0835 | 3.36 |
| daylight to D50 → daylight to D100 | 0.0491 | 0.0095 | 0.0295 | 3.11 |
| daylight to halogen → daylight to tungsten | 0.0522 | 0.0378 | 0.0870 | 2.30 |
| daylight to D100 → daylight to D40 | 0.4720 | 0.0209 | 0.0432 | 2.07 |
| daylight to D40 → a three-primary display | 0.0004 | 0.0340 | 0.0662 | 1.95 |
| a blackbody → the macular pigment | 0.1050 | 0.0119 | 0.0217 | 1.81 |
| daylight to a white LED → a green wall | 0.0244 | 0.0684 | 0.1098 | 1.61 |
| daylight to tungsten → a white LED | 0.0619 | 0.0764 | 0.1079 | 1.41 |
| the macular pigment → daylight to D50 | 0.0908 | 0.0229 | 0.0263 | 1.15 |
| an older lens → a red wall | 0.2404 | 0.0653 | 0.0710 | 1.09 |
| a green wall → a triphosphor tube | 0.6015 | 0.1143 | 0.1178 | 1.03 |
The consequence for the table is direct. Ask which adjacencies have a gap smaller than twice their own error, which is the loosest reading of not ordered by this evidence:
- paired: four of thirteen;
- unpaired: six of thirteen.
Two steps of the census ranking are established by the evidence and would have been reported as undecided.
The gain is a measurement of similarity
The right-hand column is not noise, and reading it as a ranking of how much pairing helps misses what it is.
The gain is √(var(a) + var(b)) / sd(b − a), which is a function of the correlation between the two rows across the set. High correlation means a big gain. So the column says, row by row, how alike two changes of light are in what they find difficult.
And the column can be turned into the quantity it stands for, in one line, without computing anything further. When the two rows have similar spreads the gain reduces to a function of the correlation alone — the shared variance cancels out of both numerator and denominator — and the relation inverts to
so a gain of 3.36 is a correlation of 0.911, a gain of 1.78 is 0.684, and a gain of 1.03 is 0.057. That covers the whole table: a red wall against halogen at 0.91, the two daylight shifts at 0.90, halogen against tungsten at 0.81, tungsten against a white LED at 0.50, a lens against a red wall at 0.16, and a green wall against a triphosphor tube at 0.06.
The typical pair of rows in this census correlates at 0.68, which is the mean gain of 1.78 read through the same relation, and it is the single number that says why pairing is worth having here and not worth as much as it was on the uniformity table. That table’s gains of three to four invert to correlations of 0.89 and 0.94 — and the essay’s own statement that its ellipses correlate at better than 0.9 comes back out of the arithmetic without being put in, which is the check that the identity is the right one.
Two corrections fall out of doing it. The range is quoted above as 0.1 to 0.95 and the inversion gives 0.06 to 0.91 — lower at both ends, and lower at the top by enough to matter, because a correlation of 0.95 would put the census’s most alike pair level with the uniformity table’s and it is not. And the approximation has a stated cost: the relation above holds exactly when the two rows have equal variance, and inflates the correlation by two per cent at a variance ratio of 1.5, six per cent at 2 and fifteen per cent at 3. The census’s rows differ in spread by less than a factor of two on every adjacent pair, so every figure above is good to a few per cent — which is well inside what a correlation on 125 points supports anyway.
Publishing the correlation rather than the gain is the better choice, and the reason is that a gain is a property of a comparison and a correlation is a property of a pair. A gain of 1.03 sounds like a failed technique; a correlation of 0.06 sounds like what it is, which is two changes of light that have nothing to say about each other. The first invites a reader to conclude that pairing did not work on that row, and the second invites the conclusion the section above wants — that the row is a measurement of dissimilarity and the pairing reported it correctly.
It also makes the similarity matrix usable as one. Gains are not comparable across tables, because they depend on how alike the things being compared happen to be in that table; correlations are, and the two numbers this collection now has — 0.68 typical among census rows, 0.9 among colour spaces on shared ellipses — sit on one scale and can be quoted beside each other. That is what makes the second arrival of the same arithmetic informative rather than merely repetitive, and it is available for the cost of a reciprocal.
The two largest are the two most alike pairs in the table. Daylight to D50 against daylight to D100 is two daylight phases against each other — the same illuminant family at two correlated colour temperatures, whose spectra differ by a smooth tilt. A red wall against daylight to halogen is two broad, smooth reddenings of the light with different causes. In both cases the surfaces that suffer are almost identical, and the difference between the two rows is nearly free of the surfaces altogether.
The three smallest are pairs that are not alike at all: a green wall against a triphosphor tube is a broad multiplicative filter against a source with three narrow emission lines, and there is no reason a surface awkward under one should be awkward under the other. The gain there is 1.03 — pairing buys essentially nothing, and correctly so.
So the column is worth publishing on its own account. A table of gains from pairing is a similarity matrix for the rows, computed for free while doing something else.
The second time, and why it is smaller
The same arithmetic settled a different table a round ago. Comparing two colour spaces by how nearly they make MacAdam’s ellipses circular has the same structure — the same twenty-five ellipses score both spaces — and adding the two standard errors in quadrature declared almost the whole uniformity ranking unordered, while pairing left three of seven adjacent pairs established.
There the gain was three to four. Here it averages 1.78. The difference is not a defect in either measurement and it is exactly what the paragraph above predicts.
Two colour spaces scored on the same twenty-five ellipses are two nearly identical questions: the ellipses are mapped through two similar transformations and the per-ellipse anisotropies correlate at better than 0.9. Two rows of the adaptation census are two genuinely different changes of light, and their per-surface residuals correlate at anything from 0.1 to 0.95 depending on the pair. The pairing gain is bounded by how alike the two things being compared are, and the census’s rows are less alike than the uniformity table’s columns.
That is worth stating because the tempting inference from the first case — pair everything and tables become three times sharper — is wrong. Pairing recovers whatever similarity there is, and where there is none it recovers nothing.
What it does not fix
The pairing removes the part of the uncertainty the two rows share. Everything the two rows share in their bias it removes as well, which is a second and quieter benefit: the two per cent by which the lattice overstates the region’s own integral is nearly common-mode, so a gap is a better-estimated quantity than a level in that sense too.
What it cannot touch is anything the two rows do not share. In particular:
It says nothing about a different set. A paired standard error is a statistic about which members this set contains. A set built to a different rule — the clamped realistic family, say — can reverse a pair the paired error separates comfortably, and does, at nearly nine standard errors. The pairing sharpens the instrument without widening what the instrument is about.
And it does not make an unresolved pair resolved. Four adjacencies remain inside twice their own error after pairing, and the shortest of them is a gap of four parts in ten thousand. No amount of correct statistics turns that into an ordering; the honest report is that the census does not order those rows.
Why nothing caught it
For the same reason as last time, which is that both quantities exist and only one of them was ever computed.
The census’s machinery returns a mean per row. Nothing in it has ever returned the per-surface terms, because nothing needed them — a figure draws the mean, an assertion compares two means, and the vector was summed and discarded inside the function that produced it. Recovering it required a second implementation of the same loop, deliberately not shared with the fast path that the basis search calls thousands of times.
So the covariance was not overlooked in the sense of being considered and dismissed. It was unavailable, and a quantity that is unavailable is not a quantity anybody decides about — the same shape as a helper written for one caller. The general lesson is the one the round before this one drew about a helper written for one caller: a function that sums and returns is a function whose terms nobody can audit, and where the terms are what an audit would want, returning them is not an extravagance.
What the similarity column says about the census
The gains from pairing form a similarity matrix over the census’s rows, computed for nothing, and reading it is worth a section because it groups the rows in a way nothing else in this collection does.
The two daylight-to-daylight comparisons pair at 3.1 and 2.1. Those rows are the same illuminant family at different correlated colour temperatures, and their per-surface residuals correlate at better than 0.9 — a surface awkward under one daylight shift is awkward under another, because the two shifts are nearly the same smooth tilt of the spectrum.
A red wall against halogen pairs at 3.4, the highest in the table, and the two have nothing in common mechanically: one is a reflectance between the lamp and the surface, the other is a hotter filament. What they share is the shape of what they do to the light — a broad warm reweighting — and the pairing gain says so without anybody having to notice it.
The wall against the triphosphor tube pairs at 1.03. A broad multiplicative filter and a three-line source find entirely different surfaces difficult, and the correlation is near zero.
So the column is a clustering of changes of light by what they find difficult rather than by what they are — and it puts a red wall beside a halogen lamp, which no taxonomy by mechanism would. That is a small finding and it is free, which is the argument for computing per-surface terms as a matter of course rather than only when a covariance is wanted.
Where the model stops
The paired error assumes the 125 per-surface differences are exchangeable enough for their spread to describe the spread of their mean. They are not independent — a lattice’s neighbours are similar by construction — so the error is, if anything, understated, and the effective member count bounds the correction, and a formal treatment would want an effective sample size rather than 125.
The size of that correction can be bounded from the participation ratio, which puts the effective count at 90 to 111 rather than 125, and inflates every error in the table above by at most eight per cent. That does not change any of the four/six counts, so the conclusion survives it, and stating the bound is more useful than pretending the assumption holds.
Why the gaps are drawn rather than the rows
A drawing decision that follows from the argument and is worth stating, because the alternative is what every ranking chart does.
The natural picture is fourteen rows with an error bar on each. It is compact, it is familiar, and it invites exactly the comparison this essay is about: a reader looks at whether two bars overlap, which is the unpaired test, which is three times too pessimistic here. A chart cannot carry a caption saying do not compare these bars and expect to be obeyed.
So the family draws the gaps. Thirteen bars, one per adjacent step, each with twice the paired error on its end, and a bar shorter than its own whisker marked. The comparison the picture invites is then the right one, and the wrong one is not available because the rows’ individual errors are not drawn at all.
The cost is that the picture no longer shows the rows’ values, which have to be read from the ranked labels instead. That is a real loss and it is the right trade: the levels are the part a change of construction moves by tens of per cent and the gaps are the part it does not, so a picture that makes gaps easy to read and levels hard is a picture weighted towards what the measurement establishes.
The same rule applies wherever a ranking is drawn from paired measurements, and it is the reason a forest plot of differences reads better than two columns of means.
Who found it, and when
Pairing is Student’s, in the 1908 paper that introduced the t distribution, and the paired test appears there for exactly this reason: two treatments applied to the same material have a difference whose variance is smaller than the sum of the two variances, and treating them as independent throws away the design. It is the first thing an experimental statistician checks and the thing a table of published means makes invisible, because a published mean has no material attached to it.
Here it arrived twice within two rounds, on two tables that were both being read wrong, and the second arrival is the more instructive: knowing that pairing had mattered once was not enough to look for it again, because the census’s rows were not thought of as scores on shared material until the set they share was named.
Where the ladder goes next
With the right error in hand, the census’s ranking can be read for what it actually establishes — and four of its thirteen steps turn out to be steps the evidence does not take.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The instrument named the pair that moved chromatic adaptation · residual · sampling · significance · standard error · test set
- Two instruments and one ranking chromatic adaptation · colour difference · residual · sampling · standard error · test set
- The census in six units chromatic adaptation · colour difference · mean · residual · test set
- A choice with no magnitude colour difference · mean · residual · test set
- A partial correction is worth its fraction chromatic adaptation · colour difference · residual · test set
- An extremum is still not a sample chromatic adaptation · residual · sampling · test set
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Chromatic adaptationColour differenceCorrelationMeanPaired comparisonResidualSamplingSignificanceStandard errorTest set