What a scene does

A mean has a set under it

Every adaptation number this collection publishes is an average over a hundred and twenty-five surfaces that were written down once, in one file, with no argument for how many there should be or how saturated. The average runs from exactly zero to twice itself across them, and the set has never been varied.

Assumes The census is a construction too, What would have to be wrong and Three numbers cannot see a line.

Every number this collection publishes about chromatic adaptation is an average, and the thing it is an average over has never been named in an essay.

What one change of light costs, surface by surface — daylight to tungstenA rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to tungsten is 1.635 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 3.058, which is 1.87 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.0.01.12.33.4the published mean, 1.635worst in the set, 3.0585 surfaces at exactly zerothe surfaces, sorted by what this change of light costs themΔE₀₀daylight to tungstenCIE 1931 2° observer · the set, varied
Fig. 1 The cost of one change of light — daylight to tungsten — on each of the 125 surfaces the published figure is an average of, sorted. The average is 1.635 ΔE*₀₀. The curve runs from exactly zero to 3.058.

The claim

The adaptation census reports fourteen numbers. Each is a mean over one set of surfaces, that set is a declaration rather than a measurement, and the quantity being averaged varies across it by a factor of two.

  • The set is 125 smooth reflectances, built from three basis functions on a lattice of five brightnesses and seven modulation depths each way.
  • Its coefficient of variation is between 0.36 and 0.63 on every row of the census. The published number is the middle of a wide distribution, not a property of the change of light.
  • Five of the 125 contribute nothing at all — exactly nothing, not approximately — and counting the set by how evenly it contributes gives about 100 effective members rather than 125.
  • The set has three declared numbers in it and none of them has ever been quoted with a range, varied, or defended.
  • An earlier round audited every declared scalar in the collection and said in as many words that it could reach no structural choice. This is the structural choice.

Where the number comes from

The census asks one question fourteen times. Take a change of illumination — daylight to tungsten, daylight to a triphosphor tube, a green wall between the lamp and the surface — and ask how much of it an observer’s adaptation actually removes.

The arithmetic is settled and was settled three rounds ago. A change of light acts on this family of surfaces as an exact 3×3 matrix; an adapted observer applies a diagonal gain in a named basis; the residual is what the second fails to undo of the first. It is reported in ΔE*₀₀ — a metric whose own weighting is inside every number here — and it is the difference between an object looking the same under two lights and looking different.

What is not settled, and what has never been stated, is which objects. The residual is not defined for a change of light alone. It is defined for a change of light and a surface, and the census reports the mean over a set.

The objects the average is over, and the region they come from. Two panels. On the left, eight of the 125 reflectance spectra in the test set, drawn as reflectance against wavelength from 380 to 780 nanometres — smooth, broad curves between about 0.02 and 0.9, with at most two gentle undulations each, because each is a level times a combination of two cosines. None of them has a narrow feature, because the family has no basis function that could make one. On the right, the region those surfaces come from, drawn in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the constraint that the two depths may not exceed 0.7 in sum, and 25 lattice points inside it. Five levels of each of those pairs is the whole test set. The square's four corners — the most saturated surfaces the two cosines could make — are outside the diamond and are not in the set at all.
Fig. 2 Left: eight of the 125 reflectance spectra the average is taken over. Right: the region they are drawn from, in its own two modulation coordinates — a square of allowed depths with the L¹ bound inscribed in it, and the 25 admitted depth pairs marked. Five brightness levels of each pair is the whole set.

What the set is

The file that builds it is honest about what the three basis functions are and silent about everything else. They are a constant, a half-cosine and a full cosine across the visible band, and the docstring says plainly that they are not a measured principal-component basis, that this site has no measured reflectance collection, and that the family they span is exactly three-dimensional so that the matrix result is a theorem rather than an approximation. All of that is true and all of it is stated.

What is nowhere is any account of the numbers:

what it declares the value where it comes from
how bright the surfaces are five levels, 0.10 to 0.54 nothing
how saturated they are seven depths, ±0.6 nothing
how far two modulations may go together 0.7 in sum keeps the surfaces physical
how many members 125 five times twenty-five, which is what the first three give

Only the third has a reason, and the reason is a constraint rather than a choice: a surface whose two modulations are too deep reflects a negative amount of light at the blue end, which the file refuses outright. The first two are round numbers. The fourth is not a decision at all — it is what falls out of the first three.

A set of 125 objects looks like a sample. It is not one, and the difference matters more than the count. It is not one. It is a lattice, deliberately, so that a number quoted in a figure does not carry a random seed with it — and that is the right decision, made for the right reason, and it is the only decision about the set that anybody wrote down.

What the spread is

The picture at the top of this essay is the whole finding in one line, and it is worth having the numbers rather than the shape.

change of light mean worst surface in the set ratio surfaces at zero
daylight to a blackbody at 6500 K 0.263 0.457 1.74 5
the macular pigment 0.368 1.104 3.00 5
daylight to D50 0.459 0.948 2.07 5
daylight to tungsten 1.635 3.058 1.87 5
a green wall, twice 3.375 5.786 1.71 5

The right-hand column is the one that is not obvious. A flat reflectance is a grey; a grey under a changed light is the white scaled; and an adapted observer’s gain is exactly the ratio of the whites. So the residual on a flat surface is zero — not small, zero, to machine precision — whatever the change of light is. Five of the 125 members of the lattice are flat or nearly so, and they carry none of the answer.

That is not a defect. It is what the control row of the census is for at a larger scale, and a test set with no members it handles exactly would be a test set nobody could sanity-check. But it means “a mean over a hundred and twenty-five surfaces” was always a slightly generous description, and the honest count is worth having.

No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for daylight to tungsten. The tallest bar is 1.50 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 108.5 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 3 Each of the 125 surfaces, ordered by how much it contributes to the published mean. The tallest is 1.50 per cent of the total, so no single object carries the answer; the flat tail is the surfaces an adaptation gain handles exactly. By contribution the set is worth 108.5 effective members of 125.
Every change of light this site models, and how much of it a gain removes. Each row is a change of illumination. The pale bar is how far it moves an ordinary surface for an observer who does not adapt; the solid bar at its left end is what is left after the observer has applied the one gain adaptation gives them, which is the ratio of the two whites in the CAT16 basis and is not fitted to anything. Sorted by the fraction left rather than by the size of the change, because the two orderings are different: the largest change here is removed almost entirely and the worst row is a change less than a third its size.
Fig. 4 The census as it is published: fourteen changes of light, each a single mean residual. Nothing in the table says what the mean is over, and nothing in it could.

How many surfaces the mean is really over

The count that matters for a mean is not how many terms it has but how evenly they contribute, and the standard way to say that in one number is the participation ratio — the sum squared over the sum of squares. It answers how many equally-contributing members would give this much concentration, and for a perfectly even set it is the member count exactly.

Across the census it runs from 89.6 to 110.8, out of 125.

The low end is worth naming: daylight to a three-primary display at 89.6 and the macular pigment at 94.4. Both are changes of light with sharp spectral structure, and a sharp change picks out the surfaces whose own structure lines up with it while leaving the rest nearly alone. The high end is the two wall rows, at 108 and 111, because a broad multiplicative filter treats everything a little.

So how evenly a change of light is spread over objects is itself a property of the change, which is the useful half of the statistic and the reason it is worth reporting alongside a mean — with a qualification the section below has to add about what “evenly” can and cannot be made to mean.

What one change of light costs, surface by surface — daylight to a three-primary display. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a three-primary display is 0.980 ΔE₀₀. The curve runs from 8.1e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 2.705, which is 2.76 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 5 The same picture for a three-primary display. Its participation ratio is the lowest in the census at 89.6 effective surfaces of 125, because a source with three narrow lines picks out the surfaces whose own structure lines up with them.

The two statistics are one statistic

Two numbers have now been quoted as separate characterisations of the same set — a coefficient of variation running 0.36 to 0.63, and an effective member count running 89.6 to 110.8 — and they are not independent. They are algebraically the same quantity.

For nn terms with mean μ\mu and standard deviation σ\sigma, the participation ratio is

(x)2x2  =  n2μ2n(μ2+σ2)  =  n1+c2\frac{\left(\sum x\right)^2}{\sum x^2} \;=\; \frac{n^2\mu^2}{n(\mu^2 + \sigma^2)} \;=\; \frac{n}{1 + c^2}

with cc the coefficient of variation. Substituting the two ends of the published range: 125/(1+0.362)=110.7125/(1+0.36^2) = 110.7 and 125/(1+0.632)=89.5125/(1+0.63^2) = 89.5, against the 110.8 and 89.6 printed above. The agreement is to the last digit the coefficients were rounded to.

So the effective member count adds no information to the coefficient of variation. Reporting both is reporting one measurement twice, in two vocabularies — which is worth knowing before either is quoted as corroborating the other.

The identity does not make the statistic useless; it makes clear which of the two to publish. The coefficient of variation is portable, since it does not depend on the set’s size; the effective count is interpretable, since it is in units a reader can picture. Neither is a second opinion about the first.

And it removes a reading the essay had put on it. The section above explains a low participation ratio as a sharp change of light picking out the surfaces whose structure lines up with it — a claim about concentration, about a few members carrying the answer. A participation ratio cannot support that, because it is a function of the dispersion alone: a distribution with a long thin tail and one with a broad symmetric spread have the same ratio if their coefficients of variation match. The statistic is blind to skew, which is exactly the property “picks out” is about.

The essay’s own numbers make the point. The macular row’s worst surface costs three times its mean while tungsten’s costs 1.87 — a plain difference in skew — and the two rows’ participation ratios are 94.4 and 108.5, which is what their dispersions are and says nothing about the tail. What supports the concentration claim is the worst-to-mean ratio, which is measured elsewhere on the site and which no function of the variance can replace.

One decomposition is worth having, because it separates a fact about the set from a fact about the light. A set with five exact zeros and a hundred and twenty identical values has a participation ratio of exactly 120 — so five of the missing effective members are the flat surfaces, on every row, always. The rest is genuine dispersion among the surfaces that do contribute: nine more on the mildest row and thirty more on the harshest. That second number is the one that varies with the change of light, and it is the one worth reporting, because the first is a property of the lattice and is the same fourteen times.

What this is not

Two things this essay is careful not to claim, because both are available and both would be wrong.

It is not that no single surface carries the answer. That worry is the obvious one — a mean over a lattice might really be a mean over three awkward corners — and it is measurably false. The largest single contribution anywhere is 2.40 per cent of the total, and leaving out any one surface moves the published mean by well under a per cent. The spread is a genuine spread, and a leave-one-out is the wrong alarm to raise about it.

And it is not that the mean is the wrong statistic. A mean over a set of objects is exactly the right thing to report when the question is what adaptation does in a room, and a mean survives being sampled in a way a maximum does not. The point is not that the mean should be something else. It is that a mean is a statement about a set, the set here is a declaration, and until now the declaration has been invisible in every figure and every sentence that quoted one of these numbers.

Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.
Fig. 6 The consequence of the spread, drawn ahead of its own essay: with a standard error attached, four of the census’s thirteen ranked steps turn out not to be steps the set establishes.

Why nothing caught it

The gates on this site are good at asking whether a number is right and have no way to ask what it is a number about.

colourcheck runs every assertion the libraries carry, and the assertions about the census are assertions about its arithmetic: that a change of light is exactly a matrix on this family, that the white is a fixed point of the residual operator, that a dimmer is exactly a gain. All true, all checked, none of them a question about the surfaces. overlap_check asks whether two essays make the same argument. link_density_check counts links.

Nothing in the fleet asks what a mean is over, because a mean does not know. It arrives as one number with the same shape whether it summarises a hundred and twenty-five carefully chosen objects or five arbitrary ones, and the sentence that reports it reads identically either way.

The nearest thing to a warning was in the file itself, and it was a warning about something else: the docstring says the family is not a measured collection and says so wherever the question arises. That protects against the reader who thinks the surfaces are data. It does not protect against the reader — including every essay on this site for fourteen phases — who thinks the choice of surfaces is a detail.

What one change of light costs, surface by surface — the macular pigment. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for the macular pigment is 0.368 ΔE₀₀. The curve runs from 7.0e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 1.104, which is 3.00 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 7 The same picture for a filter inside the observer rather than a change of lamp. The distribution is far more skewed: the worst surface costs three times the mean, against 1.87 for tungsten, because a narrow absorption inside the eye reaches some spectra and not others.
No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for a green wall, two bounces. The tallest bar is 1.37 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 110.8 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 8 The contribution profile for the harshest row in the census. A broad multiplicative filter treats every surface a little, which is why this row has the highest participation ratio in the table at 110.8 of 125.

What follows

Three things, and the rest of this round is them.

A set has more than one thing wrong with it at once. A width has a magnitude and one instrument: multiply it and see. A set has membership, it has a measure, it has a shape and it has a dimension, and those are four separate questions with four separate answers. Walking the same region differently is not the same question as giving the surfaces one more degree of freedom.

The spread has a use. A quantity that varies by a factor of two across a set has a standard error on its mean, and a standard error on a mean is what says whether two rows of a table are in a different order or the same order twice. Four of the census’s thirteen steps turn out not to be resolved by the set they were measured over.

And a mean is not a worst case. The reader asking whether adaptation will fail them is asking about the object it fails on, and that object costs about twice the published number on every row.

One reflectance, two illuminants, two coloursA reflectance peaking near 550 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.400450500550600650700wavelength / nmunder D650.339, 0.501under A0.419, 0.515reflectance is a fraction, 0 to 1CIE 1931 2° observer
Fig. 9 One surface, integrated under two lights. Every number in the census is this operation repeated 125 times and averaged, and the average is what gets published.

Where the model stops

This essay establishes that the census’s numbers depend on a set that was declared rather than measured. It cannot say what the right set is, and it should not be read as saying the numbers are wrong.

There is no measured reflectance collection here. Getting one would not settle the question either, for reasons the list of what an audit cannot reach sets out, because a measured collection is also a set somebody assembled — the Munsell chips, a paint manufacturer’s fan deck and a hyperspectral scene are three different declarations with three different saturations, and choosing between them is the same choice made in a different vocabulary. The honest position is the one the rest of the round takes: measure how much the answer depends on the set, publish that alongside the answer, and stop pretending the set is not there.

Who found it, and when

The shape of the argument is old and belongs to statistics rather than to colour. R. A. Fisher’s distinction between a statistic and the population it estimates is exactly this distinction, and the phrase that fits best is one econometricians use — a sampling frame, the list of things that could have been in the sample. A frame nobody wrote down is a frame nobody can criticise.

Its arrival here is more mundane. The round before this one spent itself auditing declared scalars and ended by listing what it had not reached, in words worth repeating: that the population is a pigment template rather than the physiological fundamentals, that adaptation is a diagonal at all, that a colour difference is CIEDE2000 — none of those has a multiplier to sweep. It named three structural choices and missed the fourth, which was sitting in the same file as the numbers it was auditing.

What is now printed beside a residual

The repair is a convention rather than a number, and it is small enough to state in full.

Every figure in the new family names the set in its caption strip. The right-hand slot, which on this site carries the observer, carries the set, varied — because the observer is held in all of them and the set is what changes. That costs one line and makes a residual with no set beside it visibly incomplete.

And the machinery now returns the terms, not only the sum. perSurface gives the 125 values a census row is a mean of, and everything else in the audit is built on having them: the standard error on a gap, the participation ratio, the influence profile, the worst case. None of those was available while the function summed and discarded.

And the audit found one of its own statistics to be redundant, which is a result about auditing rather than about colour. Two numbers computed from the same 125 terms, presented in different sections, reading as independent corroboration and related by an identity — that is what happens when the terms become available all at once and every summary anybody knows how to compute gets computed. The discipline that would have caught it is the one this collection applies to inputs and had not applied to outputs: ask what a statistic is a function of before quoting it beside another one.

That second change is the transferable one. A function that returns a mean is a function whose terms nobody can audit, and where the terms are what an audit would want, returning them costs one line of allocation and buys every statistic in this round.

Where the ladder goes next

The set has four things about it that can be varied, and this essay has varied none of them — it has only shown that the mean is a summary of something wide. The next rung takes the first: whether walking the same region of surfaces more finely gives a better answer, and what it means that the answer moves when it does.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 31 that link here.

The objects this essay names

Each one links to every other essay that touches it.

BasisChromatic adaptationColour differenceDegrees of freedomMeanReflectanceResidualSamplingTest setThe von Kries transform