Concept

Standard error — where it appears

How much a mean would move if the set it was taken over had different members, estimated from how much the quantity varies across that set. It is a statistic about membership, so it says nothing about a set assembled by a different rule.

Named by 4 essays across 3 fields — each of them below, with the objects they name alongside it.

The error on a gap is not the two rows' errors added. Two bars for each of the 13 adjacent pairs in the census ranking. The upper, shorter bar is the standard error of the gap taken as a paired difference — the same 125 surfaces score both rows, so a surface that is awkward under one change of light is usually awkward under the other and the difference is quieter than either. The lower bar is the two rows' own errors added in quadrature, which is what comparing error bars by eye amounts to. Pairing is worth a factor of 1.78 on average and 3.36 on the pair it helps most, and it is the difference between 6 adjacencies unordered and 4. The gain is largest where the two rows are two daylights or two tungstens, because then the surfaces they find awkward are nearly the same surfaces.

The error on a gap is not the errors at its ends

Comparing two rows of a table by looking at whether their error bars overlap is the wrong comparison, and here it is wrong by a factor of up to 3.4. The same 125 surfaces score both rows, so the difference between them is quieter than either — and how much quieter is a measurement of how alike the two rows are.

difference · Metric
Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.

Four steps the test set cannot order

The adaptation census prints fourteen numbers to four figures and its ranking is asked to say which lamps adaptation handles worst. Nine of its thirteen steps are established beyond any doubt the test set can raise; the other four are not, and three of them are consecutive — a tungsten lamp, a halogen lamp and a white LED are simply not ordered.

light · Light
Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.

The instrument named the pair that moved

A standard error over a test set flagged four steps of the adaptation census as unresolved. Rebuilding the set three different ways reversed exactly one pair, and it was one of the four. Rebuilding it to a different rule reversed a pair the error separated by nearly nine standard errors — which is not a failure of the instrument but a statement of what it is about.

limits · Limits
Which steps of the census ranking a change of unit reverses. Every adjacent pair in the published census ranking that at least one unit puts the other way round. The bar counts how many of the five other units reverse it. The marker on the left says whether the test set had already declared the pair unresolved — a gap smaller than twice its own paired standard error, which is a statement about sampling over 125 surfaces and shares no arithmetic with a change of ruler. The two pairs every unit reverses are both flagged, which is the agreement. The pair at the bottom is the disagreement: the test set resolves it at 9.1 standard errors and four of the five units reverse it anyway, because a sampling error cannot see a change of ruler and a change of ruler cannot see a sampling error.

Two instruments and one ranking

A sampling error over a hundred and twenty-five surfaces and a change of colour-difference formula share no arithmetic at all, and they were asked the same question of the same table. Every adjacency the whole menu reverses had already been flagged as unresolved. And one the test set settles at nine standard errors is reversed by four of the five formulae, which is what makes them two instruments rather than one.

limits · Limits

Named alongside it

The objects these essays reach for when they reach for this one.

Chromatic adaptationResidualTest setColour differenceSamplingSignificanceRankingCalibrationCorrelationFalsificationHalogenIlluminant

All concepts