Concept

Test set — where it appears

The collection of surfaces, patches or stimuli a figure of merit is averaged over, and the thing a published error is therefore a statement about. It is nearly always constructed rather than measured, and how saturated its members are decides most of what the average comes out as.

Named by 46 essays across 9 fields — each of them below, with the objects they name alongside it.

What one change of light costs, surface by surface — daylight to tungsten. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to tungsten is 1.635 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 3.058, which is 1.87 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.

A mean has a set under it

Every adaptation number this collection publishes is an average over a hundred and twenty-five surfaces that were written down once, in one file, with no argument for how many there should be or how saturated. The average runs from exactly zero to twice itself across them, and the set has never been varied.

scene · Scene
Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.

A lattice is a quadrature rule

Walking a set of test surfaces more finely does not converge on a better answer, because refining a lattice under a constraint changes which corners of the region get sampled and not only how densely. The lattice used here turns out to be a two per cent biased estimate of the integral it stands for.

difference · Metric
The error on a gap is not the two rows' errors added. Two bars for each of the 13 adjacent pairs in the census ranking. The upper, shorter bar is the standard error of the gap taken as a paired difference — the same 125 surfaces score both rows, so a surface that is awkward under one change of light is usually awkward under the other and the difference is quieter than either. The lower bar is the two rows' own errors added in quadrature, which is what comparing error bars by eye amounts to. Pairing is worth a factor of 1.78 on average and 3.36 on the pair it helps most, and it is the difference between 6 adjacencies unordered and 4. The gain is largest where the two rows are two daylights or two tungstens, because then the surfaces they find awkward are nearly the same surfaces.

The error on a gap is not the errors at its ends

Comparing two rows of a table by looking at whether their error bars overlap is the wrong comparison, and here it is wrong by a factor of up to 3.4. The same 125 surfaces score both rows, so the difference between them is quieter than either — and how much quieter is a measurement of how alike the two rows are.

difference · Metric
Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.

Four steps the test set cannot order

The adaptation census prints fourteen numbers to four figures and its ranking is asked to say which lamps adaptation handles worst. Nine of its thirteen steps are established beyond any doubt the test set can raise; the other four are not, and three of them are consecutive — a tungsten lamp, a halogen lamp and a white LED are simply not ordered.

light · Light
Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.

The instrument named the pair that moved

A standard error over a test set flagged four steps of the adaptation census as unresolved. Rebuilding the set three different ways reversed exactly one pair, and it was one of the four. Rebuilding it to a different rule reversed a pair the error separated by nearly nine standard errors — which is not a failure of the instrument but a statement of what it is about.

limits · Limits
What a fourth reflectance dimension costs the theorem that a change of light is a matrix. Four rising curves on axes of the fourth dimension's amplitude, left to right, against what is left of daylight to tungsten after the exact 3×3 change-of-light matrix has been applied, in ΔE₀₀. All four begin at exactly zero: on the three-dimensional family the matrix is solved rather than fitted and there is no remainder at all, which is the theorem this collection's adaptation argument is built on. Adding a fourth reflectance dimension breaks it, and how badly depends far more on the fourth function's shape than on its size — at five per cent amplitude the four shapes cost 0.329, 0.572, 0.063, 0.124 ΔE₀₀ respectively, a factor of 9.1 between the dearest and the cheapest. For scale, the smallest von Kries residual anywhere in the census is 0.26 ΔE₀₀, so the cheapest of the four is a quarter of it and the dearest is twice it.

A theorem about a family

A change of light acts on the test surfaces used here as an exact 3×3 matrix with no residual whatsoever, and the whole adaptation argument is built on that being exact. It is exact because the surfaces span exactly three dimensions, and they span exactly three dimensions because three basis functions were written down.

scene · Scene
Which fourth dimensions are expensive, and how fine is too fine to matter. Two curves on axes of how many half-cycles a cosine fourth basis function makes across the visible band, against what it costs the matrix theorem in ΔE₀₀, at a fixed ten per cent amplitude. Both curves touch zero at exactly one and two half-cycles: those are the family's own second and third basis functions, so a fourth coefficient along them adds no dimension and a change of light stays exactly a matrix. Between them the cost climbs, reaches a maximum, and — for the smooth source — falls away again, because structure finer than the scale on which three broad cone sensitivities differ integrates to nearly nothing. The two curves part company at the fine end. Under daylight-to-tungsten the cost has fallen by a factor of 2.2 from its peak; under daylight-to-a-triphosphor-tube it has barely fallen at all, because a source with three narrow emission lines has structure of its own at that scale for the surface's structure to beat against. The observer is identical in both curves.

A fourth dimension has a shape

How much a fourth reflectance dimension costs spans a factor of nine across four equally plausible shapes at one amplitude, and the expensive ones are not the shapes a variance figure would identify. The band that hurts is set by the illuminant rather than by the eye, which is why a triphosphor tube and a tungsten lamp disagree about it.

light · Light
The census as the surfaces stop being three-dimensional. A slope chart with three columns — a test set with no fourth reflectance dimension, one with a fourth dimension at ten per cent amplitude, and one at twenty — and a line per change of light. Almost every line rises: a surface with structure the observer's three channels cannot follow is a surface an adaptation gain handles worse. Two lines are drawn heavy. daylight to a three-primary display rises fastest, by 90 per cent, because a source made of three narrow lines is precisely the instrument that cannot see a fourth reflectance dimension. And daylight to a triphosphor tube falls — the only row that does — because a triphosphor tube already samples the spectrum at three places, so extra structure in the surface is partly averaged away rather than added. The order of the middle of the table is not the same at the two ends; the extremes do not move.

The row a fourth dimension improves

Giving the test surfaces one more degree of freedom makes almost every change of light harder for an adapted observer — but not all of them, and which one it helps depends entirely on what the extra dimension looks like. A triphosphor tube is improved by one shape and hurt more than anything else in the census by another.

light · Light
How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one.

Saturation is nearly everything

The set of test surfaces has three numbers describing it, and only one of them matters. How saturated the surfaces are carries an elasticity of about 0.7 on every result computed over them; how bright they are carries 0.10. A test chart's chroma range decides its answer and its lightness range does not.

difference · Metric
How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one.

The input nobody declared

An audit that swept every declared width in this collection found the largest elasticity anywhere to be about a half. The most elastic input turns out to be one that was never declared, never quoted with a range and never varied — and being undeclared is exactly why it escaped the audit that was looking for it.

limits · Limits
No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for daylight to tungsten. The tallest bar is 1.50 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 108.5 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.

The surfaces that answer nothing

Five of the hundred and twenty-five test surfaces contribute exactly zero to every number the adaptation census reports — not approximately, exactly — and the reason is the one fact about von Kries adaptation that makes it worth having at all. Counting the set by how much it contributes gives about a hundred members rather than a hundred and twenty-five.

eye · Cones
A published residual is a mean, and the worst object in the room costs twice it. Three bars for each of the 14 changes of light in the adaptation census, ordered by how uneven the change is across surfaces. The first bar is the published mean residual. The second is the worst single surface in the audit's published test set. The third is the worst surface anywhere in the region that set is drawn from, found by search rather than by reading a maximum off a lattice. The mean-to-worst ratio runs from 1.90 to 4.02 and averages 2.43, so every published adaptation number has a worst case about twice it that no essay had ever quoted. The gap between the second and third bars is the other finding: a maximum over 125 sampled points understates the region's own maximum by up to 34 per cent.

A mean is not a worst case

Every adaptation number this collection publishes is an average over objects, and the reader asking whether adaptation will fail them is asking about the object it fails on. That object costs between 1.9 and 4.0 times the published figure, and how uneven a change of light is across objects turns out to be a property of the change rather than a constant.

limits · Limits
A published residual is a mean, and the worst object in the room costs twice it. Three bars for each of the 14 changes of light in the adaptation census, ordered by how uneven the change is across surfaces. The first bar is the published mean residual. The second is the worst single surface in the audit's published test set. The third is the worst surface anywhere in the region that set is drawn from, found by search rather than by reading a maximum off a lattice. The mean-to-worst ratio runs from 1.90 to 4.02 and averages 2.43, so every published adaptation number has a worst case about twice it that no essay had ever quoted. The gap between the second and third bars is the other finding: a maximum over 125 sampled points understates the region's own maximum by up to 34 per cent.

An extremum is still not a sample

Two rounds ago three measurements turned up that took a maximum over a sample of a set and were short by up to a factor of two. The same error was live in a fourth place the whole time, on the set of surfaces every adaptation number is averaged over, and it is short by up to a third.

matching · Gamut
Every worst surface sits on a number somebody typed. The region the test surfaces are drawn from, in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the requirement that the two depths sum to no more than 0.7. The 14 marked points are the worst surface for each change of light in the adaptation census, found by search over the whole region. Every one of them lies exactly on the diamond, and every one is also at the brightest level the region allows — both declared constraints active, on all 14 rows, with no interior maximum anywhere. That is the opposite of what bounding the wall gave: there the worst case turned over at a band width of six nanometres because a narrow band returns too little light, which is physics. Here the worst case is a reading of two numbers. The one constraint that is about the world — a paint's excitation purity may not exceed 0.6 — is slack everywhere: the most saturated surface the region admits reaches 0.459.

Every worst surface sits on a declaration

Bounding the wall in a painted room produced a real worst case — the residual turns over at a band six nanometres wide because a narrower band returns too little light. Bounding the surfaces the residual is averaged over produces nothing of the kind, because all fourteen answers sit exactly on two numbers somebody typed and the one constraint that comes from the world never binds at all.

scene · Scene
How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one.

A chart decides what a camera scores

A camera profile's reported error changes by a factor of five when the test chart's saturation changes, with the camera and its matrix untouched. The elasticity is 0.67 — the same figure, to two per cent, that an entirely unrelated measurement over an entirely unrelated set of surfaces gives.

imaging · Capture
A gradient of one hue, mapped into a press's gamut 3 ways. Chroma asked for along the bottom, chroma delivered up the side, for a ramp at lightness 55 and hue angle 25°. A colorimetric intent follows the diagonal until the press runs out and is flat afterwards — the flat part is a gradient arriving as a single colour. The perceptual intent is under the diagonal from the start, which is the price of never going flat.

A budget drawn through one hue

The three-stage error budget this collection publishes for a colour-management chain is computed over twenty-four colours of a single hue at a single lightness. The quantity that actually varies with hue — how many distinguishable colours a rendering intent destroys — runs from nothing at all to more than a third, and the hue the budget uses is near the bottom of that range.

applied · Delivery
Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.

What the audit still cannot reach

Two rounds have now swept every declared width in this collection and one of its structural choices. Three structural choices remain, none of them has a multiplier to sweep, and the reason each resists is different — which makes the list a description of where this kind of audit ends rather than a queue of work.

limits · Limits
What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.

A choice with no magnitude

An audit can multiply a width by 1.25 and report an elasticity. It cannot multiply CIEDE2000 by anything. Auditing a structural choice needs a different instrument, and building one shows that six published numbers in this collection each carry a factor of about two of unit-choice — after the change of scale has been taken out.

limits · Limits
The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.

The census in six units

Recomputing every change of light in the adaptation census under six colour-difference formulae, with the scale factor divided out, leaves a table whose levels move by up to a factor of three point seven. The rows that move most are the mild ones, which is the opposite of what a reader would guess and is a property of where each formula was fitted.

light · Light
How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.

The weighting is the disagreement

Five colour-difference formulae, three decades and two committees, and the single property that predicts which of them agree is whether a chroma difference gets divided by the chroma it was measured at. It sorts the menu exactly, it cuts across the distinction between a matching difference and an appearance one, and it halves the census's largest sensitivity.

difference · Metric
The census along the line from ΔE76 to ΔE94, and past it. ΔE94 is ΔE76 with two weighting constants in it, and at zero those constants make every weight exactly one, so the two formulae are joined by a line rather than separated by a choice. The horizontal axis is how much of the published weighting is applied: 0 is exactly ΔE76, 1 is exactly ΔE94, and 3 is three times more weighting than anybody has proposed. The falling curve is the census's mean elasticity to how saturated its test set is, which drops from 1.11 to 0.74 — most of the fall happening before the published value is reached. The other curve is Kendall's τ against ΔE2000's ranking, and it peaks at w = 0.5, not at 1: the weighting that best reproduces the published ordering is about half the published weighting. There is no value of this dial that reaches ΔE2000, whose rotation term is not on this line at all.

A dial through a discrete menu

ΔE*94 is ΔE*ab with two weighting constants in it, and at zero those constants make every weight exactly one — so the two ends of the oldest disagreement in colour difference are joined by a line rather than separated by a choice. Walking it gives a derivative where a menu gives only a spread, and the derivative says the published weighting is on the far side of the interesting part.

difference · Metric
Which steps of the census ranking a change of unit reverses. Every adjacent pair in the published census ranking that at least one unit puts the other way round. The bar counts how many of the five other units reverse it. The marker on the left says whether the test set had already declared the pair unresolved — a gap smaller than twice its own paired standard error, which is a statement about sampling over 125 surfaces and shares no arithmetic with a change of ruler. The two pairs every unit reverses are both flagged, which is the agreement. The pair at the bottom is the disagreement: the test set resolves it at 9.1 standard errors and four of the five units reverse it anyway, because a sampling error cannot see a change of ruler and a change of ruler cannot see a sampling error.

Two instruments and one ranking

A sampling error over a hundred and twenty-five surfaces and a change of colour-difference formula share no arithmetic at all, and they were asked the same question of the same table. Every adjacency the whole menu reverses had already been flagged as unresolved. And one the test set settles at nine standard errors is reversed by four of the five formulae, which is what makes them two instruments rather than one.

limits · Limits
Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu.

The disagreement is at the near end

Every colour-difference formula on the menu was fitted to threshold data, so the expectation is that they agree about pairs an observer can only just tell apart and diverge on large differences. They do the opposite. Proportionally the disagreement is largest at the near end, by a factor of six for the appearance unit, and the cause is an exponent of 0.63.

difference · Metric
What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.

The observers differ by a unit's worth

The gap between the 1931 and 1964 standard observers is the one quantity in this collection's audit with no published number under it — nothing here reports it as a single figure over a stated set. It also has the second-largest dependence on which colour-difference formula is used, running from 1.60 to 3.55 across the menu.

matching · Gamut
Two sensitivities from two libraries, under every unit. Two quantities that share no code, no test set and no physical question: how much the adaptation census's residual depends on how saturated its surfaces are, and how much a camera profile's reported error depends on how saturated its test chart is. The first is a mean over fourteen changes of light built from cosine combinations; the second is one number about one silicon sensor scored on Gaussian bumps. Under the published unit they sit at 0.687 and 0.656. Across the whole menu they move together, from about 0.5 under the appearance unit to about 1.15 under plain CIELAB, staying within 12 per cent of each other at the worst point. Two numbers agreeing once is a coincidence; two curves agreeing at six points across a factor of two and a half is a shared mechanism, and the mechanism is the compression the unit applies to a chroma difference.

The coincidence was a mechanism

Two sensitivities from two libraries with no shared code came out two per cent apart, and the claim made about them was that they share a mechanism rather than a number. That claim has a colour-difference formula inside it, so it can be tested by changing the formula — and both curves move together across the whole menu, from 0.5 to 1.15.

imaging · Capture
A camera matrix refitted to minimise each unit, rather than solved in XYZ. Every camera profile here, and as far as can be told every camera profile anybody ships, is a linear least-squares solve in XYZ. That is an objective and it is on nobody's menu: it weights a difference by how large the tristimulus values are. Each row here refits the same 3×3 by direct search to minimise one of the six units instead. The upper bar is how much better the fit gets in that unit; the lower is how far the matrix itself moves, as a relative Frobenius norm. Both matter and they do not agree: CAM16-UCS moves the matrix least, at 0.34 per cent, for the largest improvement of the six, while ΔE*94 moves it 6.8 times as far for less. A score that changes is a report changing; a matrix that changes is the camera rendering different pixels.

The objective nobody chose

Every camera profile here, and as far as can be told everywhere, is a linear least-squares solve in tristimulus space. That is an objective and it is on nobody's menu — it weights an error by how bright the patch is. Refitting the same matrix to minimise a real colour-difference formula improves the fit in all six, and moves the matrix, which means different pixels rather than a different report.

imaging · Capture
What an observer is left with, by how much it is allowed to know about the room. Six ways of discounting a change of light, averaged over the fourteen changes in the adaptation census and 125 test surfaces each. The bar is what each leaves behind, on a logarithmic axis because the models span two orders of magnitude. The second line under each name is the count that matters: how many numbers about this room the model has to be given. Doing nothing leaves 15.7 ΔE₀₀. A single gain read off the two whites' luminances leaves 15.3. A matrix fitted across half the census and then applied everywhere, knowing nothing about the room at all, leaves 12.5. The published von Kries gain, which is told the white and nothing else, leaves 1.312 — and bolting a fixed correction onto it, at no cost in scene information, leaves 1.368, which is very slightly worse. The exact matrix leaves nothing and is not on the chart: its nine numbers are the change of light, which is the quantity being discounted.

Three numbers the scene supplies

An adaptation model's parameters are not all the same kind of thing. Some are numbers an observer must estimate from the room it is standing in; others could have been settled once by evolution. Counting them separately turns the diagonal gain from a crude approximation into the only model of the set that gets a large answer from information the observer can actually have.

scene · Scene
How much of the residual a partial correction removes. Between the diagonal gain and the exact matrix there is a line: apply the correction that would make a row exact, but only a fraction of it. The horizontal axis is that fraction and the vertical is the share of the row's residual it removes, for all fourteen census rows. The straight diagonal is where a correction worth exactly its fraction would fall, and in the published unit every curve lies on it to within 2.2 percentage points. The lower band of curves is the same interpolation measured in CAM16-UCS, which departs by up to 17 points — because its distance is a power of the Euclidean one and a power is not homogeneous along a ray, where every ordinary norm is. The straight line is therefore a property of the ruler rather than of the correction, and the exception is what says so.

A partial correction is worth its fraction

Between a diagonal gain and the exact matrix there is a line, and a bounded observer's natural hope is that the first part of it is worth a disproportionate share. It is not. On all fourteen changes of light, at every setting, the share of the residual removed matches the share of the correction applied to within 2.2 percentage points — which closes the last way the gap could have been cheap.

scene · Scene
What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.

Three choices reached

Two rounds ago this collection named three things it rested on and could not audit — a unit, a diagonal and a template. All three are now reached, and the interesting part is not the three answers but that four of the round's own predictions were refused by the arithmetic and one of its measurements was wrong in a way only a cross-check caught.

limits · Limits
Which of the collection's published quantities a departure can be pushed through. The six quantities the previous round recomputed under six different colour-difference units, and whether the same treatment works for a departure. Two do: the adaptation census and the metameric pair both take reflectances and a light, which is what a departure acts on. Four do not, and the reasons are different in each case rather than a single obstacle. A unit is a function applied to the answers, so it can be swapped at the end of any computation; a departure changes the object at the start, so it has to be accepted by every stage in between. That is the practical difference between auditing a convention and auditing a structure.

A departure is not a unit

The previous round audited six published quantities by swapping the unit they were quoted in — a function applied at the end of each computation. Nothing of that shape works here. A departure changes the object at the start, so every stage in between has to accept it, and only two of the same six quantities can take one at all. The fourth cannot even be expressed in the interface.

limits · Limits
The two tabulation choices over forty-two surfaces, under a tungsten lamp at 2856 K. Each column is one choice, measured over a family of forty-two analytic reflectances rather than on a single example: an absorption band of stated centre, width and depth. The four marks are the smallest, the median, the ninety-fifth percentile and the largest cost in ΔE₀₀, logarithmically. Under a smooth light the range is worth 6.3 times the step at the median, so a collection wanting one repair should widen its range rather than refine its step — and under a fluorescent tube the ranking reverses outright.

A neutral has no grid

A perfectly flat reflectance computes to exactly the same colour on every wavelength grid, through every slit, at every origin, and for every observer — not nearly, but to the last bit of a floating-point number. The condition is an identity rather than a limit, and what makes it one is the white point.

difference · Metric
How much the answer moves when the 5-nanometre grid is slid through one cell. Each bar is the spread of one light's colour across five grid origins, all at the same 5-nanometre step, in ΔE₀₀. A smooth light barely moves, and what movement it has is the end cells rather than the sampling. The fluorescent tube moves by 3.18 units and the laser projector by 35.0, because their emission lines are narrower than the step and whether a sample lands on one is a coincidence of arithmetic. This is the measurement that separates a quadrature error from an aliasing error, and no average over origins can substitute for it.

Where the grid starts

Holding the step at five nanometres and sliding the grid's origin through one cell moves a fluorescent tube's computed colour by 3.18 ΔE₀₀ and a laser projector's by 35.0. Refining the step does not fix it and averaging over origins hides it. It is the one tabulation fault with no smooth error to cancel against.

light · Light
What each end of the 380–780 nanometre range costs, by light. Two bars per light, on a logarithmic axis: the upper is what extending the range down to 300 nanometres moves the answer, the lower what extending it up to 830 does. The asymmetry is the whole figure. A thermal source has about a fifth of its power outside this collection's range and almost all of it at the long end, where the observer is already zero; what costs money is the short end, where the observer is small but not zero and daylight is still strong. A light with no ultraviolet — an LED lamp, a laser — pays nothing at either end, which is the pairing again: a range only costs what the light puts in it.

Two ends and one is empty

Extending this collection's wavelength range down to 300 nanometres moves a red pigment under daylight by 0.502 ΔE₀₀. Extending it up to 830 moves the same colour by 0.00015. A fifth of a thermal source's power lies outside the range and almost none of its colour does, and confusing those two shares is how a range gets argued about instead of measured.

light · Light
The two tabulation choices over forty-two surfaces, under a 6500 K thermal radiator. Each column is one choice, measured over a family of forty-two analytic reflectances rather than on a single example: an absorption band of stated centre, width and depth. The four marks are the smallest, the median, the ninety-fifth percentile and the largest cost in ΔE₀₀, logarithmically. Under a smooth light the range is worth 9.1 times the step at the median, so a collection wanting one repair should widen its range rather than refine its step — and under a fluorescent tube the ranking reverses outright.

Which end to buy

Refining a five-nanometre grid to one buys a daylight calculation 0.05 ΔE₀₀ and widening its range buys 0.54. Under a fluorescent tube the same two purchases are worth 0.83 and 0.0001. The ranking reverses completely, and what decides it is one length compared against one other length.

light · Light
What the normaliser cancels, per light. Two bars per light, logarithmic. The upper is the colour error a 5-nanometre sum makes when the white it is divided by is computed finely; the lower is the same sum divided by the white computed on the same coarse grid, which is what every colorimetric calculation actually does. The ratio is between 1.3 and 4.1. The grid appears twice in a tristimulus value and the two errors are the same error, so most of it divides out — which is why five nanometres has been good enough for a century without anybody having to be careful about it.

The normaliser carries the error too

A five-nanometre sum gets a red pigment's tristimulus value wrong by two hundredths of a per cent and its colour wrong by six hundredths of a unit. Those two numbers are not the same size because the grid appears twice in a colour — once in the sample and once in the white — and the two errors are largely the same error.

difference · Metric
What a tabulation step costs, by light, on paper. The horizontal axis is the tabulation step in nanometres, from one to twenty; the vertical is how far the resulting colour is from the same integral taken at a tenth of a nanometre over the same range, in ΔE₀₀, on a logarithmic scale. Each line is one light. The three with no feature narrower than the step fall smoothly and stay below a tenth of a unit at five nanometres, which is the grid used throughout. The fluorescent tube and the laser projector do not fall at all: their lines are narrower than any step drawn here, so the answer depends on where the samples land rather than on how many there are. The sample is held at paper throughout.

A grid is not a resolution

Ten essays into this collection there is one sentence about wavelength sampling, and it is that five nanometres is enough. Ten measurements later there are three decisions, three mechanisms, three repairs and two rankings, and the word resolution names none of them.

light · Light
Every departure under every light. Six departures across six lights, each cell the difference between two observers in ΔE₀₀, drawn as a bar whose length is the number. The rows are not multiples of one another: the lens is worst under tungsten and the pigment peaks are worst under a three-emitter LED, because a departure is a pairing and which light is being paired with decides it. The laser projector's row is empty, and that is not a fact about lasers — on this collection's five-nanometre grid a three-line spectrum is a one-line spectrum, and a single wavelength is a stimulus every observer agrees about exactly.

The lens is worst under tungsten

An ageing lens costs 4.38 ΔE₀₀ under a tungsten lamp and 2.20 under a fluorescent tube. The pigment peaks cost 2.84 under a three-emitter LED and 1.87 under the same tube. Six departures across six lights do not form a single ordering, which is what a pairing looks like and what a dominant term would not.

eye · Cones
Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.

The ranking is not stable

On a red pigment under daylight the six observer departures run from 2.38 down to 1.20 ΔE₀₀. Over forty-two surfaces two of them change places, the top two separate, and every one spans between a factor of ten and a factor of thirty-five. A chart of six bars is a chart of one example.

eye · Cones
The three cone absorptances at two settings of the macular pigment. Solid and dashed are the same construction at the two ends of two standard deviations of the reported spread. The curves are built from one pigment template through its ocular media, which is the same model its population of two hundred eyes is drawn from. The largest difference between the two sets is 31.5 per cent of the peak, and where it sits along the wavelength axis is what decides which stimuli the two observers disagree about — a departure concentrated in the blue is invisible on a sample with no blue in it.

The macular is a band, not a filter

The lens absorbs everything below 500 nanometres with a long tail; the macular pigment absorbs forty nanometres either side of 460 and nothing else. The two have similar sizes and completely different distributions, and the reason is that one is broad and one is narrow — which is what decides whether a sample is affected at all.

eye · Cones
Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.

A tolerance with an observer in it

A delivery tolerance is written in ΔE₀₀ against the 1931 observer, and six departures of that observer combine to about three of the same units on an ordinary saturated sample. A one-unit tolerance is being asked to contain a three-unit uncertainty that nothing in its budget mentions.

difference · Metric
Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.

The census under another observer

This collection's largest computed result is an adaptation census — fourteen changes of light judged over a hundred and twenty-five constructed surfaces. Every number in it was computed through one observer, and the observer's own departures are between one and two and a half units on the same surfaces, which is the size of the effects the census reports.

brain · Appearance
The two tabulation choices over forty-two surfaces, under a 6500 K thermal radiator. Each column is one choice, measured over a family of forty-two analytic reflectances rather than on a single example: an absorption band of stated centre, width and depth. The four marks are the smallest, the median, the ninety-fifth percentile and the largest cost in ΔE₀₀, logarithmically. Under a smooth light the range is worth 9.1 times the step at the median, so a collection wanting one repair should widen its range rather than refine its step — and under a fluorescent tube the ranking reverses outright.

The grid under the census

The adaptation census is computed on eighty-one wavelengths from 380 to 780 nanometres. Under the daylight and blackbody sources it uses, the range is worth about half a colour difference on ordinary surfaces and the step about a twentieth — so the census carries a tabulation term as well as an observer one, and they are not the same size.

brain · Appearance
The conditions under which an observer's departure is exactly zero. A departure of the observer is the pairing of something belonging to the observer with something belonging to the stimulus, so emptying either factor empties the product. The axis is logarithmic in what is left when the condition is imposed. Six rows empty the stimulus's factor — a perfectly neutral sample is the same colour for every observer, at any age and any field size — and two empty the observer's, since a gain on each cone and a change of basis are both absorbed exactly. All eight are identities rather than small numbers. The last two are the same two conditions imposed in a published cone space rather than in the observer's own, and they are worth eight and thirteen units: the identity is about the eye, and the arithmetic everybody uses is in somebody else's coordinates.

The conditions are the result

The round measured about twenty departures and established ten conditions. The departures are numbers that depend on a sample, a light and a construction; the conditions are exact, they hold under every substitution tried, and they are what a reader can act on. A size is a measurement and a condition is a mechanism.

limits · Limits
Which of the collection's published quantities a departure can be pushed through. The six quantities the previous round recomputed under six different colour-difference units, and whether the same treatment works for a departure. Two do: the adaptation census and the metameric pair both take reflectances and a light, which is what a departure acts on. Four do not, and the reasons are different in each case rather than a single obstacle. A unit is a function applied to the answers, so it can be swapped at the end of any computation; a departure changes the object at the start, so it has to be accepted by every stage in between. That is the practical difference between auditing a convention and auditing a structure.

What this round could not reach

Three audits, about twenty departures and ten conditions, and a longer list of things that were named and not measured. Every item on it is specific, most of them are an afternoon's work, and the reason none was done is the same in every case — the round ran out of round.

limits · Limits
The angle between the lens and the macular pigment, before and after adaptation. On each of 120 surfaces the angle, in the local metric, between what an older lens does to the reading and what a denser macular pigment does, binned in ten-degree steps. Read without adaptation, where both filters yellow the observer's white along with everything else, the two point nearly the same way: a median of 8 degrees. Read after each observer has adapted to its own white, the median is 156, and the two together cost less than the larger alone on 115 of the 120.

Two yellow filters cancel on a slope

An older lens and a denser macular pigment both take blue out of the light, and read before adaptation they move a colour in nearly the same direction, eight degrees apart. Once each eye has adapted to its own white they point a median 156 degrees apart on smooth reflectances and together cost less than the lens alone. On surfaces with a narrow absorption band they still sit 26 degrees apart and add. What decides it is the width of the surface's own features.

difference · Metric
How far the answer for the average is from the average answer, over eight spreads. For each spread of inputs, the distance between the mean of the model's answers and its answer for the mean input, as a share of the spread of the answers. Over a population of observers it is 1.4 per cent. Over surfaces it is 21 on smooth natural reflectances, 27 on a banded family with lightness in it, and 9 on a set of pale surfaces. Over the light one room sees in a day it is 59.

The average surface does not look average

Over a population of observers the appearance model is so nearly linear that the mean of its answers is its answer for the mean, to 1.4 per cent of the spread. Over the surfaces in a scene it is not. On 240 smooth reflectances the gap is 21 per cent of the spread, and the mean surface looks 4.2 units lighter than the surfaces look on average. The grey that matches the average light is a 47 per cent reflectance; the grey that matches the average look is a 43 per cent one.

brain · Appearance

Named alongside it

The objects these essays reach for when they reach for this one.

Chromatic adaptationColour differenceResidualReflectanceSensitivitySpecificationCalibrationMeanAuditElasticitySamplingStandard observer

All concepts