Where the model breaks

Three choices reached

Two rounds ago this collection named three things it rested on and could not audit — a unit, a diagonal and a template. All three are now reached, and the interesting part is not the three answers but that four of the round's own predictions were refused by the arithmetic and one of its measurements was wrong in a way only a cross-check caught.

Assumes What the audit still cannot reach, A choice with no magnitude and A model is a claim about what can be known.

Two rounds ago this collection audited every declared number it rests on and wrote down what it could not reach.

No structural choice was audited. That the population is a pigment template rather than the physiological fundamentals, that adaptation is a diagonal at all, that a colour difference is CIEDE2000 — none of those has a multiplier to sweep, and the audit reaches none of them.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 1 The first of the three, measured: six quantities from six files, recomputed under six colour-difference formulae with each formula’s scale divided out first. None of the six moves by less than a factor of 1.7.

The spread in that table is what a change of unit is worth. What it costs to reach one is a separate measure: how far each formula is from being a rescaling of the published one, which is the part no calibration removes.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 2 Each formula’s remaining disagreement with ΔE2000 once its own best scale factor has been divided out. A unit at zero here would be the published one in different clothes and would change no conclusion; the six quantities above move because these five are not.

The claim

All three are reached, none of them the way the sentence above expected, and the round’s own predictions were wrong four times out of four.

  • The unit is worth about a factor of two on every published difference in the collection, after the change of scale is removed.
  • The diagonal is a negative result, and a strong one: nothing fixed removes any of what it leaves, and nothing partial removes more than its share.
  • The template’s alternative did not exist. The choice was never template or fundamentals but a population or one observer, and among real templates the choice is worth five per cent.
  • Four predictions were refused, each in a different direction, and each is recorded where it was made rather than deleted.
  • And one measurement was wrong, caught by the cross-check that exists for exactly that.

The three, and what each cost

The unit. Six formulae, each calibrated onto ΔE2000’s scale over a common sample so that only real disagreement survives. Every published quantity tested moves by between ×1.71 and ×2.30. Rankings move less — between two and ten of the census’s ninety-one pairs reverse — and sensitivities move more, from 0.53 under an appearance unit to 1.11 under a plain Euclidean one. The single property that predicts which formulae agree is whether they divide a chroma difference by its own chroma, which cuts across every other way of classifying them.

The diagonal. Reached by a different move: not how far is it from exact, which was already published on fourteen rows, but where would the exact answer’s nine numbers come from. They are the change of light, so an observer with them would not need the model. Separating the parameters an observer must read off the room from the ones it could have been born with makes the diagonal the only entry on the ladder that gets a large answer from obtainable information: three scene numbers, 91.6 per cent removed.

The template. Reached by finding that the stated alternative was not one. A tabulated cone fundamental has no peak wavelength to move, so using it collapses two hundred observers into one. What is auditable is which nomogram, and the answer is 0.052 ΔE2000 between the two published ones — against 0.16 and 0.79 for two caricatures matched to them in width.

The property that predicts agreement, measured

That the menu divides by whether a formula weights a chroma difference, rather than by what kind of object it is, is a claim about a split — and a split can be scored.

unit weights chroma kind Spearman discordant pairs
ΔE′ (CAM16-UCS) yes appearance 0.9912 2
ΔE*94 yes matching 0.9780 4
ΔE*uv no matching 0.9516 7
ΔE*ab no matching 0.9341 8
ΔEok no matching 0.9121 10

The split is clean and it is the one named above. The two formulae that weight chroma are the two closest to the published ranking of the census’s rows; the three that do not are the three furthest; and the worse of the weighted pair, at 0.9780 and four discordant pairs, beats the better of the unweighted three, at 0.9516 and seven, with nothing in between them.

The other classification does not split the menu at all. Matching against appearance puts CAM16-UCS alone on one side, and its nearest neighbour by ranking is ΔE*94 — a plain matching formula from 1994 with a chroma weighting in it. Two formulae sharing nothing in their construction, one built out of an appearance model and one built by dividing CIELAB’s chroma term by the reference’s own chroma, agree with the published unit better than three formulae that share a great deal.

That is worth carrying as a rule for reading any difference number. What decides whether two colour-difference formulae rank a set of residuals the same way is not whether they are colorimetric or perceptual, modern or old, Euclidean or weighted in general. It is whether both of them discount a chroma difference by how chromatic the sample already is. Everything saturated on this census pulls a mean around under a formula that does not.

What the two groups disagree about

Splitting a menu is one thing and naming what the split is about is another, and the units’ behaviour across the six audited quantities says which.

Each quantity’s spread across the menu — its largest value over its smallest — runs from 1.705 for the settling residual to 2.302 for a camera profile’s error on saturated surfaces, with CAM16-UCS supplying the maximum on five of the six. So “about a factor of two” is a range with structure in it: the quantities computed over saturated samples spread most, and the one computed over a settling neutral spreads least.

The per-unit scatter about the menu’s own mean says the same thing from the other side. Across the six quantities ΔE00 varies by 7.6 per cent about its mean position, ΔE*94 by 10.2, ΔE*ab by 14.5, ΔE*uv by 17.5, CAM16-UCS by 19.5 and Oklab by 28.7. The two steadiest are again the two weighted ones.

So the ordering by consistency is very nearly the ordering by whether a formula weights chroma, and very nearly not the ordering by age. Oklab is the most recent formula on the menu and the least predictable of the six on this collection’s own quantities — which is a fact about what it was designed for, a smooth space to interpolate in, rather than a defect in it.

Four predictions the arithmetic refused

Each was written down before the number existed, and each is recorded where it was made. A prediction deleted after the fact leaves a file that looks as though it never made one.

That the appearance unit would be the outlier. CAM16-UCS is the only member of the menu that is not a distance between three numbers, so it should have been furthest from the published unit. It is second closest, at 24 per cent scatter against Oklab’s 35. The menu divides by whether a formula weights chroma, not by what kind of thing it is.

That the chroma weighting keeps the published unit ahead. On the collection’s own uniformity instrument, ΔE2000 scores 2.74 for anisotropy against CIELUV’s 2.37 — a plain Euclidean distance in the other 1976 space makes MacAdam’s ellipses rounder than the formula built to repair the first one does. And the weighting buys roundness by paying in evenness: over the three CIELAB-based formulae the anisotropy falls 3.42 → 2.89 → 2.74 while the spread rises 3.24 → 3.59 → 4.12. A trade with a sign that no published account mentions.

That a fixed correction would buy something. Nine numbers settled once, applied after the diagonal gain, costing the observer nothing: it improves the census rows it was fitted to by one per cent and makes the rows it was not worse by 1.2. There is nothing systematic left, which is a stronger result than the prediction would have been.

That the first tenth of a correction would be worth more than a tenth. It is worth exactly a tenth, on every row, to within 2.2 percentage points — and the explanation, that the residual is a norm along a ray, predicts its own exception: under the one unit whose distance is a power law the curve bends by 17.4 points and follows 1 − (1 − α)⁰·⁶³ to within 1.3.

Four for four, in four different directions, on four unrelated questions. The round’s rate of correct prediction about its own subject was zero, and the machinery caught all four.

MacAdam's twenty-five ellipses, measured in each unit. The uniformity instrument used here, applied to units rather than to spaces. The upper bar is anisotropy — the mean over the twenty-five of the largest radius divided by the smallest, where 1 would be a circle. The lower is spread — the largest mean radius divided by the smallest across all twenty-five, which asks whether a step of the same size means the same thing in different parts of the diagram. Reading down the three CIELAB-based formulae in the order they were published, the anisotropy falls 3.42 → 2.89 → 2.74 and the spread rises 3.24 → 3.59 → 4.12: the weighting divides a difference by the chroma it was measured at, which equalises directions at a point and unequalises magnitudes between points. Neither number is scaled, so no calibration is applied here. CAM16-UCS is ahead on both.
Fig. 3 The second refused prediction, drawn: the two uniformity measures moving in opposite directions as the chroma weighting strengthens. The upper bars fall and the lower rise.

What the machinery caught that the prose had wrong

One measurement in the round was simply wrong, and the way it was caught is the point of the arrangement.

The template audit scores each candidate by how far the median observer sits from the 1931 standard. The first implementation normalised both observers against the standard’s white, which reads as a perfectly plausible relative-colorimetric comparison and is a different one: it charges the template for a difference in white point that the collection’s own machinery discounts.

Under it the site’s median observer was 3.51 ΔE2000 from the standard. Under the correct construction — each observer judged against its own white, which is what the population machinery does — it is 0.95.

Nothing about the wrong number looked wrong. It was in the right units, the right order of magnitude for a residual this collection reports elsewhere, and it ranked the five templates in nearly the same order. What caught it was the requirement that the ΔE2000 column reproduce the number the source file publishes, to the last bit — a check that exists precisely because a re-implementation which drifted would give a plausible table and measure nothing.

That check is the round’s one methodological recommendation, and it is cheap. Every quantity in the audit’s inventory is re-implemented, and every one is required to reproduce its own source exactly. Five of six have a source; the sixth is marked as having none rather than given one that does not exist.

What did not change

Worth saying, because an audit that found everything exposed would be an audit of its own instrument.

Every ordering the collection actually argues survives. A blackbody at daylight’s temperature is the mildest change in the census under all six units; a green wall bounced twice is the harshest under all six; the room is harder than the lamp under all six. The census’s conclusions are not statements about the unit.

Ratios survive better than levels. A common-mode factor cancels, and most of what a unit does is common-mode — with one qualification this round adds: the disagreement is proportionally largest at the near end of the scale, so a ratio between a large quantity and a small one does not cancel.

And the two published nomograms are interchangeable. Nothing in the population work turns on Govardovskii rather than Lamb.

The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.
Fig. 4 The census in six units. The lines cross in the middle and not at the ends, which is the whole of what survives and what does not.

What is left, measured

Four things, and the round’s standard is that a deferral must rest on a measured number rather than an assumed one.

The caption convention, partly. This round’s three new figure families name the set or the unit or the template in the caption strip’s right-hand slot — the thing that was chosen rather than the thing held fixed. The figures that predate the convention did not, and the largest offender is now fixed: the two generators that draw census residuals across 114 placements carry the set in their strip, which is one constant and moves them into the figure index’s set band where they belong. The deferral it was under was a width nobody had measured, and the measurement is 567 px of caption against the narrowest affected figure’s 680 — 113 pixels of slack on the tightest of fifteen views. What is not done is every other family that quotes a number over a set, which is a longer list and a per-family judgement rather than one constant.

The corresponding-colour data, absent for the sixth round running. Every adaptation number here is computed on constructed surfaces under constructed changes of light, because nothing measured is held. This round makes the absence sharper again: the residual is now known to be per-scene information rather than a modelling shortfall, which is a claim that measured corresponding pairs would test directly.

The wavelength grid’s lower bound. It begins at 380 nanometres, decided in the first weeks and never revisited, and the pigment template’s secondary band peaks at 367. The collection is seeing that band’s upper flank only. Extending the grid changes every spectral integral here by a small amount, which would have to be measured before it could be believed.

And the reflectance family still has no measured alternative, which is not work this collection can schedule because it needs data. What it can do, and now has, is print the cost of the absence on the numbers that depend on it.

What the round is worth as a method

Three audits, three shapes of answer, and the shapes are more transferable than the numbers.

A menu gives a spread. Where a choice has published alternatives, recomputing under each of them gives a distribution and a list of what the distribution moves. That is what the unit audit is, it costs a calibration step to make the columns comparable, and the calibration is the part most likely to be got wrong — a comparison of rescaled quantities has a trap in it that only shows up as a turning point where none should be.

A dial gives a derivative. Where two members of a menu are joined by a continuous parameter — and one pair here is, exactly rather than approximately — the sweep says whether the published choice sits on a flat part of the curve. It does not.

And a negative result needs three instruments. The diagonal’s audit is not a spread at all; it is three separate arguments closing three separate escape routes, and any one of them alone would have left the question open. Nothing fixed helps, nothing partial helps disproportionately, and what remains is per-scene by construction.

The fourth shape is the one the template audit produced and is the least expected: the alternative did not exist, and finding that out was the whole of the work. A choice named as A rather than B where B is unavailable was never a choice between A and B; restating it correctly took an afternoon and was worth more than sweeping anything.

What each audit cost to build

The three took very different amounts of work and the ratio is not what it looks like from the outside.

The unit audit was the largest and the most mechanical. A registry of six formulae at a common interface, a calibration, and then a re-implementation of six published quantities so that each could be recomputed. Most of the effort was in the re-implementations and in the check that each reproduces its own source, which is where the round’s one real error was caught.

The diagonality audit was the smallest and the hardest to see. Its arithmetic is a ladder of six matrices scored over the census, which is an afternoon; the work was in noticing that a parameter count has two kinds of parameter in it. Nothing was built that could not have been built two rounds ago.

And the template audit was almost all reading. Four candidate curves, a refitted matrix each, one quantity scored — an hour of code — preceded by working out that the alternative named in the original sentence does not exist. That restatement is the finding and it took longer than everything after it.

The pattern is that the cheap part is the computation and the expensive part is deciding what to vary. Which is the reverse of the intuition, and is the practical reason a multiverse analysis is worth doing: the multiverse is cheap, and the framing that makes it meaningful is where the judgement goes.

Where the model stops

Three structural choices are three, and there are more. The basis a diagonal is taken in has been audited at length and is a choice; the set a mean is over was the previous round’s subject; the wavelength grid is named above. Beyond those there is a longer tail: the decision to work in tristimulus space at all, the decision that a surface is a reflectance rather than a bidirectional distribution, the decision that an observer is a set of three curves.

Each of those is a choice and each is less auditable than the last, because at some point varying a choice means writing a different collection. The three reached here were reachable because a menu of alternatives existed and somebody had published it. A choice with no alternatives in the literature is not a choice this method can reach, and the honest name for that boundary is the discipline’s own vocabulary rather than any property of the arithmetic.

Four pigment templates at 566 nm, each normalised to its own peakThe same peak wavelength through four templates: the Govardovskii nomogram this site uses, the same nomogram with its secondary band removed, Lamb's 1995 nomogram which Govardovskii's is a refinement of, and two Gaussians of the same width — one in wavelength, which is symmetric, and one in wavenumber, which is what an absorption band is usually approximated by and which is asymmetric the wrong way in wavelength. The two published nomograms are almost on top of each other. What separates them from the caricatures is the long tail towards the short wavelengths, and what separates the site's from Lamb's is the secondary band rising to 0.25 of the peak below 400 nm — a feature that sits where the lens has almost stopped transmitting and that no treatment of colour vision this collection has read mentions at all.0.000.250.500.751.00400450500550600650700750GovardovskiiGovardovskii, α onlyLamba Gaussian in wavelengtha Gaussian in wavenumberλmax 566 nmwavelength / nmabsorbanceL cone, 566 nma population of eyes · the pigment template, varied
Fig. 5 The third choice, drawn: four candidate pigment templates at one peak. Two are published, two are caricatures matched to them in width, and the gap between the groups is eighty per cent where the gap inside the published pair is five.
What an observer is left with, by how much it is allowed to know about the room. Six ways of discounting a change of light, averaged over the fourteen changes in the adaptation census and 125 test surfaces each. The bar is what each leaves behind, on a logarithmic axis because the models span two orders of magnitude. The second line under each name is the count that matters: how many numbers about this room the model has to be given. Doing nothing leaves 15.7 ΔE₀₀. A single gain read off the two whites' luminances leaves 15.3. A matrix fitted across half the census and then applied everywhere, knowing nothing about the room at all, leaves 12.5. The published von Kries gain, which is told the white and nothing else, leaves 1.312 — and bolting a fixed correction onto it, at no cost in scene information, leaves 1.368, which is very slightly worse. The exact matrix leaves nothing and is not on the chart: its nine numbers are the change of light, which is the quantity being discounted.
Fig. 6 The second choice, drawn: six models by how many numbers about the room each is told. The exact answer needs nine and they are the change of light, which is what makes the diagonal the last rung rather than a rough one.

Those are the three choices themselves. The last picture is the one place two of the round’s instruments were turned on the same object, which makes it the only agreement in the set and the only entry that could have been a disagreement.

Which steps of the census ranking a change of unit reverses. Every adjacent pair in the published census ranking that at least one unit puts the other way round. The bar counts how many of the five other units reverse it. The marker on the left says whether the test set had already declared the pair unresolved — a gap smaller than twice its own paired standard error, which is a statement about sampling over 125 surfaces and shares no arithmetic with a change of ruler. The two pairs every unit reverses are both flagged, which is the agreement. The pair at the bottom is the disagreement: the test set resolves it at 9.1 standard errors and four of the five units reverse it anyway, because a sampling error cannot see a change of ruler and a change of ruler cannot see a sampling error.
Fig. 7 The one place two instruments were pointed at the same object and agreed: every adjacency a change of unit universally reverses had already been flagged as unresolved by a sampling error that shares no arithmetic with it. Agreement between independent instruments is the round’s only positive confirmation of anything.

Who found it, and when

The three choices were named by this collection two rounds ago and the naming was the useful part. Nothing here required an idea that was not available then; it required treating three fixed things as variables, which needed a menu in each case and got one.

The general form has a name outside this subject. Multiverse analysis is the practice of computing a result under every defensible combination of the arbitrary choices inside it and reporting the distribution rather than a point, and it has had about a decade in the social sciences. Its standard finding is the one reached here: the construction moves the answer more than the sampling does, and a paper reporting a confidence interval is reporting the smaller of the two exposures.

What a computational collection adds is that the multiverse is cheap. Six units over six quantities is under a second; five templates is forty milliseconds; the whole round’s arithmetic runs in the time a single figure takes to draw. The cost of not doing it was never the computation.

Where the ladder goes next

The audit has been pointed at the collection’s inputs, its sets and now its conventions, and each round has found the newest instrument the largest exposure. The obvious next target is the one thing all three rounds have held fixed: that a surface is a reflectance and a light is a spectral power distribution, so that a colour is an integral. Where that stops being true — fluorescence, gloss, translucency, anything the geometry decides — the collection has essays and no audit.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationChromatic adaptationColour differenceElasticityHeld-outPigment templateSensitivityStructural choiceTest setUncertainty