Difference and uniformity

The error on a gap is not the errors at its ends

Comparing two rows of a table by looking at whether their error bars overlap is the wrong comparison, and here it is wrong by a factor of up to 3.4. The same 125 surfaces score both rows, so the difference between them is quieter than either — and how much quieter is a measurement of how alike the two rows are.

Assumes A lattice is a quadrature rule, A mean has a set under it and Twenty-five is a sample of the diagram.

Draw a ranking with an error bar on each row and a reader will compare the bars. That comparison is not the one the table is asking about, and on this collection’s own tables it is between one and a half and three and a half times too pessimistic.

The error on a gap is not the two rows' errors added. Two bars for each of the 13 adjacent pairs in the census ranking. The upper, shorter bar is the standard error of the gap taken as a paired difference — the same 125 surfaces score both rows, so a surface that is awkward under one change of light is usually awkward under the other and the difference is quieter than either. The lower bar is the two rows' own errors added in quadrature, which is what comparing error bars by eye amounts to. Pairing is worth a factor of 1.78 on average and 3.36 on the pair it helps most, and it is the difference between 6 adjacencies unordered and 4. The gain is largest where the two rows are two daylights or two tungstens, because then the surfaces they find awkward are nearly the same surfaces.
Fig. 1 For each of the thirteen adjacent pairs in the adaptation census, the standard error of the gap taken two ways: as a paired difference over the same 125 surfaces, and as the two rows’ own errors added in quadrature. The second is what comparing error bars amounts to.

The claim

Two rows of the adaptation census are two numbers computed over the same set of surfaces. The error on their difference is therefore not the two errors combined, and treating it as though it were declares half again as much of the table unordered as is actually unordered.

  • The gain from pairing averages 1.78 across the thirteen adjacencies and reaches 3.36 on the pair it helps most.
  • It is the difference between four adjacencies unresolved and six.
  • The size of the gain is itself a measurement. It is largest where the two changes of light are alike — two daylights, a tungsten against a halogen — because then the surfaces one finds awkward are the surfaces the other finds awkward.
  • This is the second time the same arithmetic has been needed in two rounds, on two different tables, for the same reason, and the gain is smaller here for a reason worth stating.

What the wrong comparison is

Each census row is a mean over 125 surfaces, and a mean over a set has a standard error — with the caveat about what a standard error over a lattice licenses, which does not affect this argument at all because both rows carry it equally.

Write the two rows as A and B, with per-surface values aᵢ and bᵢ. The question the table asks is whether mean(b) − mean(a) is greater than zero. Its variance is var(b̄ − ā) = var(ā) + var(b̄) − 2 cov(ā, b̄), and the third term is the whole essay. Adding the two errors in quadrature — which is what the eye does with two error bars — is exactly this expression with the covariance dropped. Dropping it is right when the two quantities were measured independently. Here they were measured on the same objects, one after the other, in the same loop.

And the covariance is large, because a surface that is awkward under one change of light is usually awkward under another. A deeply modulated reflectance has spectral structure the observer’s three channels cannot follow, and that is true whichever lamp is swapped in. So aᵢ and bᵢ rise and fall together across the set, the difference bᵢ − aᵢ is far quieter than either, and its standard error is correspondingly smaller.

What it is worth

adjacent pair gap paired error unpaired gain
a red wall → daylight to halogen 0.2258 0.0249 0.0835 3.36
daylight to D50 → daylight to D100 0.0491 0.0095 0.0295 3.11
daylight to halogen → daylight to tungsten 0.0522 0.0378 0.0870 2.30
daylight to D100 → daylight to D40 0.4720 0.0209 0.0432 2.07
daylight to D40 → a three-primary display 0.0004 0.0340 0.0662 1.95
a blackbody → the macular pigment 0.1050 0.0119 0.0217 1.81
daylight to a white LED → a green wall 0.0244 0.0684 0.1098 1.61
daylight to tungsten → a white LED 0.0619 0.0764 0.1079 1.41
the macular pigment → daylight to D50 0.0908 0.0229 0.0263 1.15
an older lens → a red wall 0.2404 0.0653 0.0710 1.09
a green wall → a triphosphor tube 0.6015 0.1143 0.1178 1.03

The consequence for the table is direct. Ask which adjacencies have a gap smaller than twice their own error, which is the loosest reading of not ordered by this evidence:

  • paired: four of thirteen;
  • unpaired: six of thirteen.

Two steps of the census ranking are established by the evidence and would have been reported as undecided.

Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.
Fig. 2 The same thirteen steps drawn as gaps rather than as rows, each with twice the paired error on its end. Drawing the gaps is the point: a chart of rows with error bars invites exactly the comparison this essay is about.
What one change of light costs, surface by surface — daylight to D100. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to D100 is 0.508 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 1.082, which is 2.13 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 3 Daylight to D100, surface by surface. Set beside the picture below it for D50 the two are nearly the same curve, which is the correlation the pairing exploits.

The gain is a measurement of similarity

The right-hand column is not noise, and reading it as a ranking of how much pairing helps misses what it is.

The gain is √(var(a) + var(b)) / sd(b − a), which is a function of the correlation between the two rows across the set. High correlation means a big gain. So the column says, row by row, how alike two changes of light are in what they find difficult.

And the column can be turned into the quantity it stands for, in one line, without computing anything further. When the two rows have similar spreads the gain reduces to a function of the correlation alone — the shared variance cancels out of both numerator and denominator — and the relation inverts to

ρ  =  11g2\rho \;=\; 1 - \frac{1}{g^{2}}

so a gain of 3.36 is a correlation of 0.911, a gain of 1.78 is 0.684, and a gain of 1.03 is 0.057. That covers the whole table: a red wall against halogen at 0.91, the two daylight shifts at 0.90, halogen against tungsten at 0.81, tungsten against a white LED at 0.50, a lens against a red wall at 0.16, and a green wall against a triphosphor tube at 0.06.

The typical pair of rows in this census correlates at 0.68, which is the mean gain of 1.78 read through the same relation, and it is the single number that says why pairing is worth having here and not worth as much as it was on the uniformity table. That table’s gains of three to four invert to correlations of 0.89 and 0.94 — and the essay’s own statement that its ellipses correlate at better than 0.9 comes back out of the arithmetic without being put in, which is the check that the identity is the right one.

Two corrections fall out of doing it. The range is quoted above as 0.1 to 0.95 and the inversion gives 0.06 to 0.91 — lower at both ends, and lower at the top by enough to matter, because a correlation of 0.95 would put the census’s most alike pair level with the uniformity table’s and it is not. And the approximation has a stated cost: the relation above holds exactly when the two rows have equal variance, and inflates the correlation by two per cent at a variance ratio of 1.5, six per cent at 2 and fifteen per cent at 3. The census’s rows differ in spread by less than a factor of two on every adjacent pair, so every figure above is good to a few per cent — which is well inside what a correlation on 125 points supports anyway.

Publishing the correlation rather than the gain is the better choice, and the reason is that a gain is a property of a comparison and a correlation is a property of a pair. A gain of 1.03 sounds like a failed technique; a correlation of 0.06 sounds like what it is, which is two changes of light that have nothing to say about each other. The first invites a reader to conclude that pairing did not work on that row, and the second invites the conclusion the section above wants — that the row is a measurement of dissimilarity and the pairing reported it correctly.

It also makes the similarity matrix usable as one. Gains are not comparable across tables, because they depend on how alike the things being compared happen to be in that table; correlations are, and the two numbers this collection now has — 0.68 typical among census rows, 0.9 among colour spaces on shared ellipses — sit on one scale and can be quoted beside each other. That is what makes the second arrival of the same arithmetic informative rather than merely repetitive, and it is available for the cost of a reciprocal.

The two largest are the two most alike pairs in the table. Daylight to D50 against daylight to D100 is two daylight phases against each other — the same illuminant family at two correlated colour temperatures, whose spectra differ by a smooth tilt. A red wall against daylight to halogen is two broad, smooth reddenings of the light with different causes. In both cases the surfaces that suffer are almost identical, and the difference between the two rows is nearly free of the surfaces altogether.

The three smallest are pairs that are not alike at all: a green wall against a triphosphor tube is a broad multiplicative filter against a source with three narrow emission lines, and there is no reason a surface awkward under one should be awkward under the other. The gain there is 1.03 — pairing buys essentially nothing, and correctly so.

So the column is worth publishing on its own account. A table of gains from pairing is a similarity matrix for the rows, computed for free while doing something else.

What one change of light costs, surface by surface — daylight to a triphosphor tube. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a triphosphor tube is 2.323 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 4.786, which is 2.06 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 4 A triphosphor tube, surface by surface. Nothing about this curve resembles the two daylight curves above, which is why pairing a daylight row against this one buys a factor of 1.03 rather than 3.11.

The second time, and why it is smaller

The same arithmetic settled a different table a round ago. Comparing two colour spaces by how nearly they make MacAdam’s ellipses circular has the same structure — the same twenty-five ellipses score both spaces — and adding the two standard errors in quadrature declared almost the whole uniformity ranking unordered, while pairing left three of seven adjacent pairs established.

There the gain was three to four. Here it averages 1.78. The difference is not a defect in either measurement and it is exactly what the paragraph above predicts.

Two colour spaces scored on the same twenty-five ellipses are two nearly identical questions: the ellipses are mapped through two similar transformations and the per-ellipse anisotropies correlate at better than 0.9. Two rows of the adaptation census are two genuinely different changes of light, and their per-surface residuals correlate at anything from 0.1 to 0.95 depending on the pair. The pairing gain is bounded by how alike the two things being compared are, and the census’s rows are less alike than the uniformity table’s columns.

That is worth stating because the tempting inference from the first case — pair everything and tables become three times sharper — is wrong. Pairing recovers whatever similarity there is, and where there is none it recovers nothing.

What one change of light costs, surface by surface — daylight to D50. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to D50 is 0.458 ΔE₀₀. The curve runs from 1.7e-13 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 0.948, which is 2.07 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 5 The per-surface costs of daylight to D50, sorted. The picture for daylight to D100 is very nearly the same picture, which is why the gap between them has an error a third the size of either.
Three of the seven adjacent pairs in the table are actually ordered. One bar per adjacent pair of the uniformity ranking: the difference between the two spaces' scores divided by the standard error of that difference over the twenty-five ellipses. The comparison is paired — the same ellipses score both spaces, so an ellipse that is hard for everybody cancels — which is why the table is more informative than it looks and why treating the two errors as independent would have declared almost nothing ordered. 3 of the 7 pairs clear two; the rest do not, and one of them crosses zero in 35 per cent of resamples. The ranking's ends are real and its middle is not a ranking.
Fig. 6 The same arithmetic on the uniformity table a round ago, where the two things compared were two colour spaces scored on the same twenty-five ellipses. There the gain was three to four; here it is 1.78, and the difference is how alike the two things compared are.

What it does not fix

The pairing removes the part of the uncertainty the two rows share. Everything the two rows share in their bias it removes as well, which is a second and quieter benefit: the two per cent by which the lattice overstates the region’s own integral is nearly common-mode, so a gap is a better-estimated quantity than a level in that sense too.

What it cannot touch is anything the two rows do not share. In particular:

It says nothing about a different set. A paired standard error is a statistic about which members this set contains. A set built to a different rule — the clamped realistic family, say — can reverse a pair the paired error separates comfortably, and does, at nearly nine standard errors. The pairing sharpens the instrument without widening what the instrument is about.

And it does not make an unresolved pair resolved. Four adjacencies remain inside twice their own error after pairing, and the shortest of them is a gap of four parts in ten thousand. No amount of correct statistics turns that into an ordering; the honest report is that the census does not order those rows.

No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for a red wall. The tallest bar is 1.54 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 106.0 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 7 A red wall’s contribution profile. Pairing this row against daylight-to-halogen — two broad, smooth reddenings with different causes — is worth a factor of 3.36, the largest gain in the table.

Why nothing caught it

For the same reason as last time, which is that both quantities exist and only one of them was ever computed.

The census’s machinery returns a mean per row. Nothing in it has ever returned the per-surface terms, because nothing needed them — a figure draws the mean, an assertion compares two means, and the vector was summed and discarded inside the function that produced it. Recovering it required a second implementation of the same loop, deliberately not shared with the fast path that the basis search calls thousands of times.

So the covariance was not overlooked in the sense of being considered and dismissed. It was unavailable, and a quantity that is unavailable is not a quantity anybody decides about — the same shape as a helper written for one caller. The general lesson is the one the round before this one drew about a helper written for one caller: a function that sums and returns is a function whose terms nobody can audit, and where the terms are what an audit would want, returning them is not an extravagance.

Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.
Fig. 8 Why the pairing matters beyond the error bars: the constructions move every row together, so a gap carries far less of the set’s uncertainty than a level does.

What the similarity column says about the census

The gains from pairing form a similarity matrix over the census’s rows, computed for nothing, and reading it is worth a section because it groups the rows in a way nothing else in this collection does.

The two daylight-to-daylight comparisons pair at 3.1 and 2.1. Those rows are the same illuminant family at different correlated colour temperatures, and their per-surface residuals correlate at better than 0.9 — a surface awkward under one daylight shift is awkward under another, because the two shifts are nearly the same smooth tilt of the spectrum.

A red wall against halogen pairs at 3.4, the highest in the table, and the two have nothing in common mechanically: one is a reflectance between the lamp and the surface, the other is a hotter filament. What they share is the shape of what they do to the light — a broad warm reweighting — and the pairing gain says so without anybody having to notice it.

The wall against the triphosphor tube pairs at 1.03. A broad multiplicative filter and a three-line source find entirely different surfaces difficult, and the correlation is near zero.

So the column is a clustering of changes of light by what they find difficult rather than by what they are — and it puts a red wall beside a halogen lamp, which no taxonomy by mechanism would. That is a small finding and it is free, which is the argument for computing per-surface terms as a matter of course rather than only when a covariance is wanted.

Where the model stops

The paired error assumes the 125 per-surface differences are exchangeable enough for their spread to describe the spread of their mean. They are not independent — a lattice’s neighbours are similar by construction — so the error is, if anything, understated, and the effective member count bounds the correction, and a formal treatment would want an effective sample size rather than 125.

The size of that correction can be bounded from the participation ratio, which puts the effective count at 90 to 111 rather than 125, and inflates every error in the table above by at most eight per cent. That does not change any of the four/six counts, so the conclusion survives it, and stating the bound is more useful than pretending the assumption holds.

Why the gaps are drawn rather than the rows

A drawing decision that follows from the argument and is worth stating, because the alternative is what every ranking chart does.

The natural picture is fourteen rows with an error bar on each. It is compact, it is familiar, and it invites exactly the comparison this essay is about: a reader looks at whether two bars overlap, which is the unpaired test, which is three times too pessimistic here. A chart cannot carry a caption saying do not compare these bars and expect to be obeyed.

So the family draws the gaps. Thirteen bars, one per adjacent step, each with twice the paired error on its end, and a bar shorter than its own whisker marked. The comparison the picture invites is then the right one, and the wrong one is not available because the rows’ individual errors are not drawn at all.

The cost is that the picture no longer shows the rows’ values, which have to be read from the ranked labels instead. That is a real loss and it is the right trade: the levels are the part a change of construction moves by tens of per cent and the gaps are the part it does not, so a picture that makes gaps easy to read and levels hard is a picture weighted towards what the measurement establishes.

The same rule applies wherever a ranking is drawn from paired measurements, and it is the reason a forest plot of differences reads better than two columns of means.

Who found it, and when

Pairing is Student’s, in the 1908 paper that introduced the t distribution, and the paired test appears there for exactly this reason: two treatments applied to the same material have a difference whose variance is smaller than the sum of the two variances, and treating them as independent throws away the design. It is the first thing an experimental statistician checks and the thing a table of published means makes invisible, because a published mean has no material attached to it.

Here it arrived twice within two rounds, on two tables that were both being read wrong, and the second arrival is the more instructive: knowing that pairing had mattered once was not enough to look for it again, because the census’s rows were not thought of as scores on shared material until the set they share was named.

Where the ladder goes next

With the right error in hand, the census’s ranking can be read for what it actually establishes — and four of its thirteen steps turn out to be steps the evidence does not take.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Chromatic adaptationColour differenceCorrelationMeanPaired comparisonResidualSamplingSignificanceStandard errorTest set