Concept

Mean — where it appears

The arithmetic average of a quantity over a set of objects, and the statistic almost every figure of merit is reported as. It answers what happens to an average member and says nothing about the member a model fails on, which is usually a factor of two worse.

Named by 11 essays across 7 fields — each of them below, with the objects they name alongside it.

What one change of light costs, surface by surface — daylight to tungsten. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to tungsten is 1.635 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 3.058, which is 1.87 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.

A mean has a set under it

Every adaptation number this collection publishes is an average over a hundred and twenty-five surfaces that were written down once, in one file, with no argument for how many there should be or how saturated. The average runs from exactly zero to twice itself across them, and the set has never been varied.

scene · Scene
Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.

A lattice is a quadrature rule

Walking a set of test surfaces more finely does not converge on a better answer, because refining a lattice under a constraint changes which corners of the region get sampled and not only how densely. The lattice used here turns out to be a two per cent biased estimate of the integral it stands for.

difference · Metric
The error on a gap is not the two rows' errors added. Two bars for each of the 13 adjacent pairs in the census ranking. The upper, shorter bar is the standard error of the gap taken as a paired difference — the same 125 surfaces score both rows, so a surface that is awkward under one change of light is usually awkward under the other and the difference is quieter than either. The lower bar is the two rows' own errors added in quadrature, which is what comparing error bars by eye amounts to. Pairing is worth a factor of 1.78 on average and 3.36 on the pair it helps most, and it is the difference between 6 adjacencies unordered and 4. The gain is largest where the two rows are two daylights or two tungstens, because then the surfaces they find awkward are nearly the same surfaces.

The error on a gap is not the errors at its ends

Comparing two rows of a table by looking at whether their error bars overlap is the wrong comparison, and here it is wrong by a factor of up to 3.4. The same 125 surfaces score both rows, so the difference between them is quieter than either — and how much quieter is a measurement of how alike the two rows are.

difference · Metric
No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for daylight to tungsten. The tallest bar is 1.50 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 108.5 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.

The surfaces that answer nothing

Five of the hundred and twenty-five test surfaces contribute exactly zero to every number the adaptation census reports — not approximately, exactly — and the reason is the one fact about von Kries adaptation that makes it worth having at all. Counting the set by how much it contributes gives about a hundred members rather than a hundred and twenty-five.

eye · Cones
A published residual is a mean, and the worst object in the room costs twice it. Three bars for each of the 14 changes of light in the adaptation census, ordered by how uneven the change is across surfaces. The first bar is the published mean residual. The second is the worst single surface in the audit's published test set. The third is the worst surface anywhere in the region that set is drawn from, found by search rather than by reading a maximum off a lattice. The mean-to-worst ratio runs from 1.90 to 4.02 and averages 2.43, so every published adaptation number has a worst case about twice it that no essay had ever quoted. The gap between the second and third bars is the other finding: a maximum over 125 sampled points understates the region's own maximum by up to 34 per cent.

A mean is not a worst case

Every adaptation number this collection publishes is an average over objects, and the reader asking whether adaptation will fail them is asking about the object it fails on. That object costs between 1.9 and 4.0 times the published figure, and how uneven a change of light is across objects turns out to be a property of the change rather than a constant.

limits · Limits
What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.

A choice with no magnitude

An audit can multiply a width by 1.25 and report an elasticity. It cannot multiply CIEDE2000 by anything. Auditing a structural choice needs a different instrument, and building one shows that six published numbers in this collection each carry a factor of about two of unit-choice — after the change of scale has been taken out.

limits · Limits
The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.

The census in six units

Recomputing every change of light in the adaptation census under six colour-difference formulae, with the scale factor divided out, leaves a table whose levels move by up to a factor of three point seven. The rows that move most are the mild ones, which is the opposite of what a reader would guess and is a property of where each formula was fitted.

light · Light
Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.

The ranking is not stable

On a red pigment under daylight the six observer departures run from 2.38 down to 1.20 ΔE₀₀. Over forty-two surfaces two of them change places, the top two separate, and every one spans between a factor of ten and a factor of thirty-five. A chart of six bars is a chart of one example.

eye · Cones
Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.

The census under another observer

This collection's largest computed result is an adaptation census — fourteen changes of light judged over a hundred and twenty-five constructed surfaces. Every number in it was computed through one observer, and the observer's own departures are between one and two and a half units on the same surfaces, which is the size of the effects the census reports.

brain · Appearance
How far the answer for the average is from the average answer, over eight spreads. For each spread of inputs, the distance between the mean of the model's answers and its answer for the mean input, as a share of the spread of the answers. Over a population of observers it is 1.4 per cent. Over surfaces it is 21 on smooth natural reflectances, 27 on a banded family with lightness in it, and 9 on a set of pale surfaces. Over the light one room sees in a day it is 59.

The average surface does not look average

Over a population of observers the appearance model is so nearly linear that the mean of its answers is its answer for the mean, to 1.4 per cent of the spread. Over the surfaces in a scene it is not. On 240 smooth reflectances the gap is 21 per cent of the spread, and the mean surface looks 4.2 units lighter than the surfaces look on average. The grey that matches the average light is a 47 per cent reflectance; the grey that matches the average look is a 43 per cent one.

brain · Appearance
Noise clipped at zero, averaged over a shadow, under tungsten. A grey ramp from black to ten per cent reflectance under tungsten, captured at three illustrative noise levels, with every negative raw reading set to zero before the readings are averaged over an area. Each line is the colour difference between that average and the noiseless grey. At high gain a half per cent grey is 1.04 off and a black frame 0.79; at very high gain the worst is 2.95, at 1.0 per cent. The same readings averaged before any clip come back exactly, at every level. The tint is gone once every channel sits several deviations above zero.

Clipped noise does not average away

Noise on a raw reading is as likely to fall below the true value as above it, which is why averaging an area removes it. A converter that sets negative readings to zero keeps the upper half and throws the lower away, and the mean of what is left is the signal plus a pedestal. With no black level error anywhere, a half per cent grey under a tungsten lamp comes out 1.04 colour differences off at high gain, 8.98 after a four-stop push — and a blur that removes every trace of the noise leaves the tint where it was.

imaging · Capture

Named alongside it

The objects these essays reach for when they reach for this one.

Test setChromatic adaptationResidualColour differenceReflectanceSamplingThe von Kries transformBasisCalibrationCIEDE2000Declared inputDegrees of freedom

All concepts