The census under another observer
Assumes The census is a construction too, A mean has a set under it and An observer is a contract.
An audit that finds a term worth several colour differences owes an accounting against the collection’s own largest body of numbers. This is that accounting, and it comes out better than it might have.
The claim
The adaptation census’s comparisons survive the observer audit and several of its absolute numbers do not, and the difference is whether the observer appears on both sides of a subtraction.
- The census reports differences between adaptation models and between changes of light, both of which hold the observer fixed on both sides, so the term cancels to first order.
- Its absolute levels do not. A statement that a given change of light costs 2.4 ΔE₀₀ carries an observer term of comparable size that is not in the figure.
- The census’s own set-sensitivity work has the same structure, and this round’s departures are one more input the census was never varied against.
- And the audit does not change any of the census’s conclusions, which is worth stating as plainly as the caveats.
What the census is
The census is a construction too established what the object is: a hundred and twenty-five constructed surfaces, fourteen changes of light, and for each combination the colour difference between what an adaptation model predicts and what a full computation gives. It is the largest single table this collection has produced and most of its adaptation results rest on it.
Everything in it is computed through the 1931 observer. That was never a choice anybody made; it is what this collection’s spectral machinery does by default, and the census inherited it the way every other calculation here does.
The audit’s question is what that inheritance costs. And the answer has two halves, because the census reports two kinds of number.
The comparisons survive
Most of what the census is used for is comparative. Which of two adaptation transforms predicts better; which change of light is hardest; whether a correction helps; how the ranking moves under a different colour-difference formula.
In every one of those the observer appears on both sides. Two adaptation models evaluated on the same surfaces under the same lights through the same observer differ by a quantity in which the observer’s own departures largely cancel — not exactly, since the departures interact with the surfaces, but to first order and by a large factor.
That is the same argument the previous round used about its four departures and the same one that protects most of this collection’s scene work. A comparison inherits the difference of a shared term rather than the term, and the difference of a shared term is second order.
So the census’s rankings, its identification of which lights are expensive, its finding about which parameters the answers are sensitive to — all of that stands.
The absolute levels do not survive. The census also publishes levels: a change of light costs so many ΔE₀₀, a correction recovers so much of it, a residual of such a size remains.
Those are single numbers about a single configuration and each carries the observer term in full. A census level of 2.4 ΔE₀₀ computed through the 1931 observer would be a different number computed through a seventy-year-old’s, and the difference is of the same order as the level.
That does not make the levels wrong. They are exactly what they claim: what the adaptation model costs for the standard observer, which is a well-defined and useful quantity, and it is the quantity a specification written in the same terms would need.
What it means is that reading a census level as what a person would see adds a term of comparable size, and the census’s captions do not distinguish the two readings. That is the same distinction matching and appearance have always had in this collection, arriving in a table of numbers rather than in a caution.
That linearity is what makes the estimate below possible. Both the census’s levels and this round’s departures scale with the same property of a surface — how far it sits from the light it is seen under — so their ratio is roughly constant across a set even though each varies by an order of magnitude across it.
Estimating the ratio is therefore a matter of comparing two numbers on comparable surfaces rather than of matching two test sets patch by patch. That is a weaker comparison than a proper joint computation and it is available now, which is the trade this essay makes throughout.
The one place the two interact
There is a case where the observer term does not cancel out of a comparison, and it is worth finding because it is the census’s most interesting result.
The census asks which changes of light are expensive, and an adaptation model’s job is to discount a change of illumination. Its errors are largest when the two lights differ most in spectral shape — which is exactly when a stimulus’s deviation from its adapting white is largest, which is exactly when the observer departures are largest.
So the two are correlated: the changes of light the census finds hardest are the changes under which observers disagree most. That is not a cancellation and it is not an addition; it is a common cause, and it means the census’s difficult cases are difficult twice over.
The size of the correlation has not been computed. Doing so means running the census under two observers and comparing, which is an afternoon with the machinery this round built and is not in this round. It is recorded as owed.
What a two-observer census would show
Sketching the experiment is worth doing because it says what would be learned.
Running the census twice, once through the standard observer and once through a seventy-year-old’s, and comparing the census results rather than the colours, would answer a question nobody has asked: does an adaptation model’s error depend on whose eye it is being evaluated through?
There are two possible outcomes and they are both interesting. If the census’s rankings are unchanged, the model’s failures are properties of the spectra rather than of the observer, which strengthens every conclusion the census reaches. If they change, then an adaptation model tuned on one observer’s data is tuned for that observer, which would be a substantial finding about the whole corresponding-colour literature.
The corresponding-colour data are absent — this collection has recorded that for six rounds — and the absence bites here too, because the published transforms were fitted to matching data from panels of observers whose composition is not reported.
The census’s own sensitivity work
There is a pleasing symmetry between what this round is doing to the census and what the census already does to itself.
A mean has a set under it audited the census’s own declared inputs: the hundred and twenty-five surfaces, their brightness and saturation distributions, the dimension of the basis they were built from. It found that varying those moves the published levels by more than any declared width in the collection.
The observer is one more such input, and it was not on the list because it was not visible as an input at all — it lives in a module rather than in a function signature. That is the same shape as the widths that were declared and never varied, and the same shape as the reflectance that was treated as data rather than as a choice.
A parameter that is imported rather than passed is a parameter nobody varies, and three rounds of this collection have now found one each.
Where the census’s surfaces sit
The identity gives a way of estimating the census’s exposure without rerunning it, and the estimate is reassuring.
The census’s surfaces are constructed with declared brightness and saturation distributions, and the saturation distribution was found to be nearly everything — the census’s levels are dominated by how saturated its surfaces are. The observer departure is proportional to the same quantity.
So the two scale together, and the ratio between a census level and its observer term is roughly constant across the set rather than varying by the factor of sixteen the surface family shows. That is worth having: it means the census’s exposure can be summarised by one ratio rather than needing a per-surface accounting.
Estimating that ratio properly needs the two-observer run. From the numbers available it is somewhere near one — the census’s levels and this round’s departures are of comparable size on comparable surfaces — which is the least comfortable value it could take.
Three sentences would repair the census’s captions. Three sentences would repair the census’s captions and none of them requires recomputing anything.
Its levels are for the standard observer, which is the contract a specification is written against and is not a person. A reader wanting what a viewer would see should add an observer term of comparable size.
Its rankings are robust to the observer to first order, because the term appears on both sides. Those are the results the census is mostly used for.
And its hardest cases are hardest twice, because the changes of light that defeat an adaptation model are the changes under which observers disagree most. That is a common cause rather than a compounding, and it means the census’s tail is more exposed than its median.
Writing those three sentences is what an audit produces when the audited work turns out to be sound: not a correction, but a statement of what its numbers are numbers about. The same treatment was owed to the scene work for a different assumption, and the shape of the answer is the same.
What was computed, and how
Nothing is recomputed here. The census’s numbers are as published and the observer departures are this round’s, over a different but comparable surface family.
The claim that the census’s comparisons survive is an argument rather than a measurement, and it is the standard first-order argument about a shared term. It could be checked by the two-observer run and has not been.
The claim that its absolute levels carry the term in full is not an argument; it follows from the levels being single numbers about single configurations, with no subtraction to cancel anything.
An audit of a large table is for this. It is worth stepping back and saying what has and has not been achieved here, because the answer is unglamorous and is the usual answer.
Nothing in the census has been shown to be wrong. Nothing in it has been recomputed. What has changed is that a term which was invisible — because it lived in an imported module rather than in a declared input — is now visible, sized, and sorted into the claims it affects and the claims it does not.
That is what an audit produces most of the time, and it is worth more than it sounds. The census’s numbers were already correct and were being read in two ways, one of which they support and one of which they do not, with nothing in the presentation to separate them. Separating them costs three sentences and turns a table that was silently ambiguous into one that is not.
The alternative outcome — finding a table’s numbers wrong — is rarer and, when it happens, easier to act on. The common case is a result that is right and incompletely stated, and the reason audits are worth running is that the incompleteness is invisible from inside the work that produced it.
Where the model stops
The census’s surfaces and this round’s are different constructions — a hundred and twenty-five with declared distributions against forty-two on a regular grid — so the comparison between their sizes is an order-of-magnitude one.
The first-order cancellation argument assumes the observer departures are small compared with the quantities being compared, and on the census’s hardest cases they are not: a departure of 2.4 units against a level of 2.4 units is not a perturbation. So the cancellation is weakest exactly where the census’s results are most interesting.
And nothing here touches the adaptation transform’s own basis, which is worth up to fifteen colour differences against the physiology and is a separate inherited term.
The warm-light case deserves separating out because it is where the census’s exposure is largest and where its results are most used. A great many of its fourteen changes involve a tungsten or a warm fluorescent source at one end, and under those the observer departures are between 0.79 and 4.47 ΔE₀₀ against 1.20 to 2.38 under daylight.
So the census’s warm-light rows carry roughly twice the observer term its daylight rows do, and its warm-light rows are the ones a lighting designer or a retail specification would consult. The audit’s term is largest exactly where the census is most consulted, which is a coincidence of subject matter rather than of arithmetic and is worth knowing before a number is quoted.
That also sharpens what the two-observer run would settle. Running it under daylight would find a smaller effect than running it under tungsten, so a single run would understate or overstate depending on which light was chosen — and the honest version is the full fourteen changes under two observers, which is a larger job than an afternoon and is what is actually owed.
The generalisation
The habit is about auditing a large result against a newly-measured term.
The temptation is to withdraw the result, and it is nearly always wrong: most of a large result is comparisons, comparisons survive shared terms, and withdrawing is expensive and unnecessary. The opposite temptation — to note the term and move on — leaves the reader unable to tell which parts are affected.
The move is to sort the result’s claims into those with the term on both sides and those with it on one, which is usually a quick reading and produces a short list. The short list is what to check and everything else is safe.
The failure mode is to sort by importance rather than by structure. The claims that matter most are not the claims most exposed, and here the census’s most-cited results are its rankings, which are the safest, while its least-cited absolute levels are the exposed ones.
One last note about what the census would look like if it were rebuilt today rather than audited. It would take the observer as an argument, the way the CIE’s own 2006 fundamental observer does, and its published levels would come with a spread rather than as points.
That is not a criticism of how it was built. Nothing in this collection had an observer as an argument until this round wrote one, and a table cannot be parameterised by a variable that does not exist. What it does mean is that the next large table here should be, and the machinery to do it is now in the tree.
Who found it, and when
The census is this collection’s own and dates from its adaptation work several rounds ago. Its self-audit — the finding that its declared surface distributions move its published levels — is from the round that examined what a mean is a statement about.
The observer’s contribution to adaptation model evaluation is not, so far as this collection can tell, quantified anywhere. The corresponding-colour data sets the transforms were fitted to were collected from panels of observers and averaged, so the observer variation is inside the data rather than measured beside it.
Where the ladder goes next
The audit’s terms have now been carried into a tolerance, an instrument, a display, a camera profile and a census. What is left is the place where a colour is finally judged, which is a room with a person in it — and a booth and an eye disagree about more than the audit has counted so far.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A departure is not a unit audit · declared input · sensitivity · test set
- A discount nobody measured cat16 · chromatic adaptation · declared input · the von kries transform
- A gain is not an observer chromatic adaptation · individual variation · standard observer · the von kries transform
- A gain needs a basis cat16 · chromatic adaptation · standard observer · the von kries transform
- One person is two observers chromatic adaptation · individual variation · standard observer · the von kries transform
- The census in six units chromatic adaptation · mean · test set · the von kries transform
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AuditCAT16Chromatic adaptationDeclared inputIndividual variationMeanSensitivityStandard observerTest setThe von Kries transform