What the eye does

The surfaces that answer nothing

Five of the hundred and twenty-five test surfaces contribute exactly zero to every number the adaptation census reports — not approximately, exactly — and the reason is the one fact about von Kries adaptation that makes it worth having at all. Counting the set by how much it contributes gives about a hundred members rather than a hundred and twenty-five.

Assumes A mean has a set under it, The three numbers a gain cannot see and What no adaptation can remove.

The most useful members of a test set are usually the hardest ones. Five of the members here are the easiest possible, they contribute nothing at all, and they are the reason the whole model is worth writing down.

No one surface carries the answer, and the set is smaller than it looksA falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for daylight to tungsten. The tallest bar is 1.50 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 108.5 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.the mean, 1.635108.5 effectivesurfaces of 1255 contribute nothingthe surfaces, ordered by contributionΔE₀₀daylight to tungstenCIE 1931 2° observer · the set, varied
Fig. 1 The 125 test surfaces ordered by how much each contributes to one published residual. The tail is flat on the axis: five surfaces contribute under one per cent of the mean, and the smallest of them contributes 4 × 10⁻¹⁴.

The claim

A flat reflectance costs an adapted observer exactly zero under any change of illumination, and that is not an approximation.

  • The residual on a grey is 4 × 10⁻¹⁴ ΔE*₀₀ — machine zero — for every one of the fourteen changes of light in the census.
  • The reason is a two-line identity, not a coincidence, and it is what a von Kries gain is.
  • Five of the 125 surfaces are flat or nearly so, so a mean advertised as being over 125 objects has five terms that were never going to say anything.
  • Counting the set by contribution gives 89.6 to 110.8 effective members, and which end a row sits at is a property of the change of light rather than of the set.
  • The right response is not to remove them. A test set with no members the model handles exactly is a test set nobody can sanity-check.

The identity

Take a flat reflectance: ρ(λ) = k for every wavelength. Under a light E its tristimulus is ∫ E(λ) · k · x̄(λ) dλ, which is k · W where W is the tristimulus of the perfect diffuser under E — which is to say, the white. A grey is the white scaled, exactly, for any light and any observer.

Now apply a von Kries gain. The adapted observer computes the two whites in a chosen basis and multiplies each channel by their ratio, so that the second light’s white is carried onto the first’s. Applied to a grey under the second light, the gain carries k · W′ to k · W, which is the grey under the first light.

Exactly. No residual, no approximation, no dependence on the basis. Every von Kries transform in existence — Hunt–Pointer–Estévez, Bradford, CAT02, CAT16, a plain XYZ scaling — agrees on greys, because getting the white right is what all of them are constructed to do and a grey is the white scaled.

That is the machine zero in the table. Not a small number that happens to round to nothing: an algebraic identity evaluated in floating point.

What it says about adaptation

This collection has an assertion about it, and the assertion is a warning rather than a result: the white is a fixed point of the residual operator by construction, whatever the basis — which is why an adaptation transform that gets the white right proves nothing.

That sentence is the whole of why the census exists. If getting the white right were the job, every transform would score identically and there would be nothing to measure. What separates them is what they do to everything that is not the white, which is what a gain cannot see, and the amount of separation available is exactly the amount of departure-from-flat in the test set.

So the five flat surfaces are not a defect in the set and they are not padding. They are the control, and they have the property a control needs: the answer on them is known in advance, it is zero, and it is zero for a reason that would still hold if every other number in the file were wrong. A run in which they came back at 0.02 rather than 10⁻¹⁴ would be a run with a bug in it, and nothing else in the census would necessarily show it.

What one change of light costs, surface by surface — daylight to a triphosphor tube. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a triphosphor tube is 2.323 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 4.786, which is 2.06 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 2 A triphosphor tube’s costs across the set, the harshest but one in the census. It too begins at exactly zero, because a grey is a grey under a three-line source as much as under daylight.
What one change of light costs, surface by surface — daylight to D50. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to D50 is 0.458 ΔE₀₀. The curve runs from 1.7e-13 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 0.948, which is 2.07 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 3 A mild daylight change, sorted. It too starts at exactly zero, and the jump from the five zeros to the smallest non-flat surface is abrupt because the set is a lattice rather than a random draw.

Which surfaces are flat

The set is a lattice indexed by a brightness and two modulation depths. A surface is flat when both depths are zero, and that happens once per brightness level — five levels, five flat surfaces.

The count is worth checking against the arithmetic rather than taken from the sorted list, because “under one per cent of the mean” catches near-flat members too. It does not here: the five members with both depths at zero return between 10⁻¹⁴ and 10⁻¹³, and the next smallest surface in the set returns about 0.3 ΔE*₀₀ for a tungsten change — four per cent of the largest. There is nothing in between, so the set has five surfaces that answer nothing and a hundred and twenty that answer something.

That is a consequence of the lattice being a lattice. A randomly drawn set would have members arbitrarily close to flat, which is one more reason a lattice was chosen and a continuum of near-zero contributions, which is another small reason the deterministic construction was the right decision even though it makes the set a quadrature rule.

How many surfaces the mean is over

The honest count for a mean is not how many terms it has but how evenly they contribute. The standard one-number answer is the participation ratio (Σ xᵢ)² / Σ xᵢ², which equals the member count exactly when every term is equal and falls as the distribution concentrates.

change of light n_eff of 125 largest single share
a green wall, two bounces 110.8 1.37%
daylight to tungsten 108.5 1.50%
a green wall, one bounce 108.3 1.46%
daylight to D40 106.7 1.58%
an older lens 101.1 1.69%
daylight to D100 99.5 1.71%
the macular pigment 94.4 2.40%
daylight to a three-primary display 89.6 2.21%

Two things to read off it.

The set is worth about a hundred members rather than a hundred and twenty-five, which is a modest correction and inflates every standard error in the collection by at most eight per cent. It does not change any of the resolved/unresolved counts, which is why the correction is stated as a bound rather than applied.

And the column is a measurement in its own right. How evenly a change of light is spread over objects is a property of the change: a broad multiplicative filter — a painted wall — treats everything a little and sits at 110, while a sharp spectral change — a three-primary display, a macular pigment — picks out the surfaces whose own structure lines up with it and sits at 90. That is the same pairing argument a fourth dimension makes at a different scale, visible here without adding anything.

No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for daylight to a three-primary display. The tallest bar is 2.21 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 89.6 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 4 The lowest participation ratio in the census, 89.6 of 125. A source with three narrow lines is a sharp instrument and finds a minority of the surfaces much harder than the rest.
No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for a green wall, two bounces. The tallest bar is 1.37 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 110.8 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 5 The most evenly spread row in the census, at 110.8 effective members of 125. A broad multiplicative filter treats every surface a little, which is what an even contribution profile looks like.

The ceiling is 120, not 125

Five terms of a hundred and twenty-five are identically zero, so the participation ratio cannot reach 125 whatever the other hundred and twenty do. Its ceiling is 120, and every row should be read against that.

change of light n_eff of 125 of 120
a green wall, two bounces 110.8 88.6 % 92.3 %
daylight to tungsten 108.5 86.8 % 90.4 %
daylight to D100 99.5 79.6 % 82.9 %
the macular pigment 94.4 75.5 % 78.7 %
a three-primary display 89.6 71.7 % 74.7 %

The correction decomposes the error inflation the essay reports. √(125/n_eff) runs from 1.062 to 1.181, and 1.021 of that is the five zeros alone — the floor a perfectly even set of the other hundred and twenty would still show. So on the most even row the genuine unevenness contributes 1.041 and the structural zeros contribute 1.021, and the two are the same size; on the least even row the unevenness contributes 1.157 and dominates.

Which makes the two halves of the correction different kinds of thing. The 2 per cent from the zeros is a fact about the lattice and applies to every row identically; the rest is a measurement of the change of light. Reporting one number for both hides that the smallest correction in the census is mostly bookkeeping.

The two concentration statistics disagree about which row is worst

The essay reads the table two ways — by n_eff and by the largest single share — and treats them as saying the same thing. On its own numbers they do not.

Dividing each row’s largest share by the share an even set would give, 1/n_eff:

change of light largest share even share ratio
a green wall, two bounces 1.37 % 0.90 % 1.52
daylight to tungsten 1.50 % 0.92 % 1.63
a three-primary display 2.21 % 1.12 % 1.98
the macular pigment 2.40 % 1.06 % 2.27

By n_eff the three-primary display is the most concentrated row at 89.6; by largest share it is the macular pigment at 2.40 per cent, and by the ratio between them the macular row is furthest out at 2.27. The two statistics pick different winners because they measure different shapes of concentration: a low n_eff can mean a broad minority carrying more than its share, while a high peak ratio means one surface carrying much more than the rest of that minority.

The mechanism reading survives and gets sharper. The three-primary source concentrates its cost on many surfaces — the ones with structure near its three lines — which drops n_eff without producing a single dominant term. The macular pigment concentrates on one, which barely moves n_eff and gives the census its largest individual contribution. So sharp spectral changes sit at 89 to 95 groups two rows that are unlike each other, and the peak ratio is the statistic that separates them.

It also says where dropping a single surface would matter most, which is the macular row — and it matters there not because that row is hard, since it is nearly the mildest in the census, but because its mechanism is narrow enough to have a victim.

Why no single surface carries the answer

The worry the influence profile is drawn to settle is the obvious one: a mean over a lattice might really be a mean over three awkward corners with a crowd behind them.

It is not. The largest single contribution anywhere in the census is 2.40 per cent of the total, on the macular row, and the typical figure is between 1.4 and 1.7. Leaving out any one surface moves the published mean by well under a per cent, so a leave-one-out is not an alarm worth raising and the spread is a genuine spread rather than a few outliers.

That is what makes the participation ratio the right summary and the jackknife the right error. A mean dominated by one term would need neither — it would need the term named — and a mean with a hundred nearly equal terms is exactly the case both instruments are for.

Where the flat surfaces do matter

One place, and it is this round’s own headline.

Everything the census measures is a departure from flat, so the amount of answer available is the amount of modulation in the set. That is why the elasticity to the surfaces’ saturation is 0.7 and the elasticity to their brightness is 0.10: brightness scales a grey into another grey, which is still worth zero, while saturation is the axis along which the set stops being flat.

Put the other way: the five flat surfaces are the origin of the axis the whole audit turns out to be about. They contribute nothing to any number and they explain why one of the set’s three declared parameters carries seven times the weight of another.

How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one.
Fig. 6 The elasticities, with the flat surfaces’ explanation visible in them: how saturated the set is decides how far it is from all-flat, and how far it is from all-flat is how much answer there is.
Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.
Fig. 7 The effective count’s practical consequence: it inflates every error in this chart by at most eight per cent, which moves none of the four unresolved adjacencies into or out of the resolved set.

Why nothing caught the count

Because five terms of zero in a mean of a hundred and twenty-five do nothing detectable.

They lower the mean by four per cent relative to a set without them, which is inside every other uncertainty in the neighbourhood; they raise the standard deviation slightly; they do not move any ordering. Every gate on this site passes identically with or without them, and no figure looked different.

They became visible only when the terms of the mean were recovered rather than summed — the recovery a paired comparison needed for a different reason. The per-surface residuals did not exist as a quantity until this round, because nothing needed them — and the first thing a sorted list of them shows is a flat tail on the axis.

That is a small instance of a general point about instrumentation. A sum reports what it sums to. A list of terms reports what kind of object the sum is, and the two are different amounts of information at almost the same cost.

Where the model stops

The identity holds for a von Kries gain, which is what the census measures. It does not hold for a full appearance model: CIECAM16 applies a nonlinear response compression after the gain, and a grey under two very different adapting luminances does not have the same lightness even when the chromatic adaptation is exact — which is the Hunt effect and is a real prediction rather than a residual.

Nor does it hold for incomplete adaptation. At a degree of adaptation below one the gain is pulled towards the identity, the white is no longer carried exactly onto the white, and a grey acquires a residual — which is small, and is the mechanism behind every adaptation number in this collection assuming complete adaptation being a correction of 1.74 on the mean.

So the flat surfaces are exactly zero for the quantity the census reports and not for every quantity in the neighbourhood, which is worth stating because the control is only a control for the thing it is exact about.

What the profile shape says about a change of light

The influence profile was drawn to settle a worry and it turns out to carry a measurement, which is worth taking rather than leaving.

Its shape — how fast it falls from the largest contributor to the tail — is the participation ratio in a picture, and the participation ratio separates the census’s rows by mechanism more cleanly than the residual itself does.

Broad multiplicative changes sit at 106 to 111. Both wall rows, daylight-to-D40, the two-bounce row. A filter that multiplies the whole spectrum by a smooth function reweights every surface a little, so the contributions are nearly equal and the profile is nearly flat.

Sharp spectral changes sit at 89 to 95. A three-primary display, the macular pigment. These pick out the surfaces whose own structure lines up with theirs and leave the rest nearly alone, so a minority of surfaces carries a disproportionate share.

And the ordering by participation ratio is not the ordering by residual. The macular pigment is nearly the mildest row in the census and among the most concentrated; a green wall bounced twice is the harshest and the most even. So the two statistics say different things about a change of light, and the second says which changes have a victim.

That is a second measurement from a picture drawn to answer a different question, and it cost nothing beyond drawing it.

Who found it, and when

The identity is von Kries’s, in the sense that it is what a von Kries transform is designed to do, and the observation that it makes greys uninformative is old enough to be folklore — it is why colour-constancy experiments use Mondrians of many chromatic patches and why a grey card is a calibration target rather than a test target.

Its appearance here is the ordinary one: a control that nobody put in on purpose. The five flat surfaces are in the set because a lattice of modulation depths includes zero, and zero was included because a lattice that skipped its own centre would have been a stranger construction than one that did not.

The objects the average is over, and the region they come from. Two panels. On the left, eight of the 125 reflectance spectra in the test set, drawn as reflectance against wavelength from 380 to 780 nanometres — smooth, broad curves between about 0.02 and 0.9, with at most two gentle undulations each, because each is a level times a combination of two cosines. None of them has a narrow feature, because the family has no basis function that could make one. On the right, the region those surfaces come from, drawn in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the constraint that the two depths may not exceed 0.7 in sum, and 25 lattice points inside it. Five levels of each of those pairs is the whole test set. The square's four corners — the most saturated surfaces the two cosines could make — are outside the diamond and are not in the set at all.
Fig. 8 The five flat surfaces are the centre column of the diamond, one per brightness level. They are in the set because a lattice includes its own origin, not because anybody decided a control was wanted.

A control nobody has to maintain

The five flat surfaces are the best kind of check because nothing has to be done to keep them working, and it is worth saying why that is unusual.

Most checks in this collection are written: an assertion states a property, is checked whenever the collection is rebuilt, and has to be maintained as the machinery moves. The flat surfaces are a check by construction. They are in the set because a lattice of modulation depths contains zero, they return exactly zero because of an algebraic identity, and both facts survive any change to the lattice’s size, its levels, its basis functions or the change of light being measured.

That gives them a property no written assertion has: they check things nobody thought to check. A bug in the white-point calculation, in the basis inversion, in the adaptation gain, in the CIELAB conversion — any of those breaks the identity and shows up as a non-zero value where a zero belongs, whatever the bug was.

The one thing they do not check is the thing the census is actually about, since getting the white right proves nothing. So the set’s five easiest members verify the machinery and its hundred and twenty hard ones carry the argument, which is a reasonable division of labour and was arrived at by nobody deciding it.

What the effective count changes

The participation ratio corrects the standard errors and it is worth being precise about how little.

An error computed as sd/√125 should be sd/√n_eff, which inflates it by √(125/n_eff) — between 1.06 and 1.18 across the census, so at most eighteen per cent and typically eight. Applied to the thirteen adjacent gaps, it moves no adjacency across the resolved/unresolved line: the smallest resolved t is 2.0 and would become 1.9, which is the one borderline case, and the largest unresolved is 1.4 and would become 1.3.

So the correction is real, is bounded, and changes no conclusion — which is why it is stated as a bound rather than applied. Applying it would put a second uncertain quantity inside every error bar, and a correction smaller than the thing it corrects for is better reported than absorbed.

Where the ladder goes next

A mean over a set has now been characterised from several directions. The question it does not answer is the one a reader actually has — not what adaptation does on average, but what it does to the object it fails on.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

BasisChromatic adaptationDegrees of freedomThe grey-world assumptionMeanReflectanceResidualTest setThe von Kries transformWhite point