Three choices reached
Assumes What the audit still cannot reach, A choice with no magnitude and A model is a claim about what can be known.
Two rounds ago this collection audited every declared number it rests on and wrote down what it could not reach.
No structural choice was audited. That the population is a pigment template rather than the physiological fundamentals, that adaptation is a diagonal at all, that a colour difference is CIEDE2000 — none of those has a multiplier to sweep, and the audit reaches none of them.
The spread in that table is what a change of unit is worth. What it costs to reach one is a separate measure: how far each formula is from being a rescaling of the published one, which is the part no calibration removes.
The claim
All three are reached, none of them the way the sentence above expected, and the round’s own predictions were wrong four times out of four.
- The unit is worth about a factor of two on every published difference in the collection, after the change of scale is removed.
- The diagonal is a negative result, and a strong one: nothing fixed removes any of what it leaves, and nothing partial removes more than its share.
- The template’s alternative did not exist. The choice was never template or fundamentals but a population or one observer, and among real templates the choice is worth five per cent.
- Four predictions were refused, each in a different direction, and each is recorded where it was made rather than deleted.
- And one measurement was wrong, caught by the cross-check that exists for exactly that.
The three, and what each cost
The unit. Six formulae, each calibrated onto ΔE2000’s scale over a common sample so that only real disagreement survives. Every published quantity tested moves by between ×1.71 and ×2.30. Rankings move less — between two and ten of the census’s ninety-one pairs reverse — and sensitivities move more, from 0.53 under an appearance unit to 1.11 under a plain Euclidean one. The single property that predicts which formulae agree is whether they divide a chroma difference by its own chroma, which cuts across every other way of classifying them.
The diagonal. Reached by a different move: not how far is it from exact, which was already published on fourteen rows, but where would the exact answer’s nine numbers come from. They are the change of light, so an observer with them would not need the model. Separating the parameters an observer must read off the room from the ones it could have been born with makes the diagonal the only entry on the ladder that gets a large answer from obtainable information: three scene numbers, 91.6 per cent removed.
The template. Reached by finding that the stated alternative was not one. A tabulated cone fundamental has no peak wavelength to move, so using it collapses two hundred observers into one. What is auditable is which nomogram, and the answer is 0.052 ΔE2000 between the two published ones — against 0.16 and 0.79 for two caricatures matched to them in width.
The property that predicts agreement, measured
That the menu divides by whether a formula weights a chroma difference, rather than by what kind of object it is, is a claim about a split — and a split can be scored.
| unit | weights chroma | kind | Spearman | discordant pairs |
|---|---|---|---|---|
| ΔE′ (CAM16-UCS) | yes | appearance | 0.9912 | 2 |
| ΔE*94 | yes | matching | 0.9780 | 4 |
| ΔE*uv | no | matching | 0.9516 | 7 |
| ΔE*ab | no | matching | 0.9341 | 8 |
| ΔEok | no | matching | 0.9121 | 10 |
The split is clean and it is the one named above. The two formulae that weight chroma are the two closest to the published ranking of the census’s rows; the three that do not are the three furthest; and the worse of the weighted pair, at 0.9780 and four discordant pairs, beats the better of the unweighted three, at 0.9516 and seven, with nothing in between them.
The other classification does not split the menu at all. Matching against appearance puts CAM16-UCS alone on one side, and its nearest neighbour by ranking is ΔE*94 — a plain matching formula from 1994 with a chroma weighting in it. Two formulae sharing nothing in their construction, one built out of an appearance model and one built by dividing CIELAB’s chroma term by the reference’s own chroma, agree with the published unit better than three formulae that share a great deal.
That is worth carrying as a rule for reading any difference number. What decides whether two colour-difference formulae rank a set of residuals the same way is not whether they are colorimetric or perceptual, modern or old, Euclidean or weighted in general. It is whether both of them discount a chroma difference by how chromatic the sample already is. Everything saturated on this census pulls a mean around under a formula that does not.
What the two groups disagree about
Splitting a menu is one thing and naming what the split is about is another, and the units’ behaviour across the six audited quantities says which.
Each quantity’s spread across the menu — its largest value over its smallest — runs from 1.705 for the settling residual to 2.302 for a camera profile’s error on saturated surfaces, with CAM16-UCS supplying the maximum on five of the six. So “about a factor of two” is a range with structure in it: the quantities computed over saturated samples spread most, and the one computed over a settling neutral spreads least.
The per-unit scatter about the menu’s own mean says the same thing from the other side. Across the six quantities ΔE00 varies by 7.6 per cent about its mean position, ΔE*94 by 10.2, ΔE*ab by 14.5, ΔE*uv by 17.5, CAM16-UCS by 19.5 and Oklab by 28.7. The two steadiest are again the two weighted ones.
So the ordering by consistency is very nearly the ordering by whether a formula weights chroma, and very nearly not the ordering by age. Oklab is the most recent formula on the menu and the least predictable of the six on this collection’s own quantities — which is a fact about what it was designed for, a smooth space to interpolate in, rather than a defect in it.
Four predictions the arithmetic refused
Each was written down before the number existed, and each is recorded where it was made. A prediction deleted after the fact leaves a file that looks as though it never made one.
That the appearance unit would be the outlier. CAM16-UCS is the only member of the menu that is not a distance between three numbers, so it should have been furthest from the published unit. It is second closest, at 24 per cent scatter against Oklab’s 35. The menu divides by whether a formula weights chroma, not by what kind of thing it is.
That the chroma weighting keeps the published unit ahead. On the collection’s own uniformity instrument, ΔE2000 scores 2.74 for anisotropy against CIELUV’s 2.37 — a plain Euclidean distance in the other 1976 space makes MacAdam’s ellipses rounder than the formula built to repair the first one does. And the weighting buys roundness by paying in evenness: over the three CIELAB-based formulae the anisotropy falls 3.42 → 2.89 → 2.74 while the spread rises 3.24 → 3.59 → 4.12. A trade with a sign that no published account mentions.
That a fixed correction would buy something. Nine numbers settled once, applied after the diagonal gain, costing the observer nothing: it improves the census rows it was fitted to by one per cent and makes the rows it was not worse by 1.2. There is nothing systematic left, which is a stronger result than the prediction would have been.
That the first tenth of a correction would be worth more than a tenth. It is worth exactly a tenth, on every row, to within 2.2 percentage points — and the explanation, that the residual is a norm along a ray, predicts its own exception: under the one unit whose distance is a power law the curve bends by 17.4 points and follows 1 − (1 − α)⁰·⁶³ to within 1.3.
Four for four, in four different directions, on four unrelated questions. The round’s rate of correct prediction about its own subject was zero, and the machinery caught all four.
What the machinery caught that the prose had wrong
One measurement in the round was simply wrong, and the way it was caught is the point of the arrangement.
The template audit scores each candidate by how far the median observer sits from the 1931 standard. The first implementation normalised both observers against the standard’s white, which reads as a perfectly plausible relative-colorimetric comparison and is a different one: it charges the template for a difference in white point that the collection’s own machinery discounts.
Under it the site’s median observer was 3.51 ΔE2000 from the standard. Under the correct construction — each observer judged against its own white, which is what the population machinery does — it is 0.95.
Nothing about the wrong number looked wrong. It was in the right units, the right order of magnitude for a residual this collection reports elsewhere, and it ranked the five templates in nearly the same order. What caught it was the requirement that the ΔE2000 column reproduce the number the source file publishes, to the last bit — a check that exists precisely because a re-implementation which drifted would give a plausible table and measure nothing.
That check is the round’s one methodological recommendation, and it is cheap. Every quantity in the audit’s inventory is re-implemented, and every one is required to reproduce its own source exactly. Five of six have a source; the sixth is marked as having none rather than given one that does not exist.
What did not change
Worth saying, because an audit that found everything exposed would be an audit of its own instrument.
Every ordering the collection actually argues survives. A blackbody at daylight’s temperature is the mildest change in the census under all six units; a green wall bounced twice is the harshest under all six; the room is harder than the lamp under all six. The census’s conclusions are not statements about the unit.
Ratios survive better than levels. A common-mode factor cancels, and most of what a unit does is common-mode — with one qualification this round adds: the disagreement is proportionally largest at the near end of the scale, so a ratio between a large quantity and a small one does not cancel.
And the two published nomograms are interchangeable. Nothing in the population work turns on Govardovskii rather than Lamb.
What is left, measured
Four things, and the round’s standard is that a deferral must rest on a measured number rather than an assumed one.
The caption convention, partly. This round’s three new figure families name the set or the unit or the template in the caption strip’s right-hand slot — the thing that was chosen rather than the thing held fixed. The figures that predate the convention did not, and the largest offender is now fixed: the two generators that draw census residuals across 114 placements carry the set in their strip, which is one constant and moves them into the figure index’s set band where they belong. The deferral it was under was a width nobody had measured, and the measurement is 567 px of caption against the narrowest affected figure’s 680 — 113 pixels of slack on the tightest of fifteen views. What is not done is every other family that quotes a number over a set, which is a longer list and a per-family judgement rather than one constant.
The corresponding-colour data, absent for the sixth round running. Every adaptation number here is computed on constructed surfaces under constructed changes of light, because nothing measured is held. This round makes the absence sharper again: the residual is now known to be per-scene information rather than a modelling shortfall, which is a claim that measured corresponding pairs would test directly.
The wavelength grid’s lower bound. It begins at 380 nanometres, decided in the first weeks and never revisited, and the pigment template’s secondary band peaks at 367. The collection is seeing that band’s upper flank only. Extending the grid changes every spectral integral here by a small amount, which would have to be measured before it could be believed.
And the reflectance family still has no measured alternative, which is not work this collection can schedule because it needs data. What it can do, and now has, is print the cost of the absence on the numbers that depend on it.
What the round is worth as a method
Three audits, three shapes of answer, and the shapes are more transferable than the numbers.
A menu gives a spread. Where a choice has published alternatives, recomputing under each of them gives a distribution and a list of what the distribution moves. That is what the unit audit is, it costs a calibration step to make the columns comparable, and the calibration is the part most likely to be got wrong — a comparison of rescaled quantities has a trap in it that only shows up as a turning point where none should be.
A dial gives a derivative. Where two members of a menu are joined by a continuous parameter — and one pair here is, exactly rather than approximately — the sweep says whether the published choice sits on a flat part of the curve. It does not.
And a negative result needs three instruments. The diagonal’s audit is not a spread at all; it is three separate arguments closing three separate escape routes, and any one of them alone would have left the question open. Nothing fixed helps, nothing partial helps disproportionately, and what remains is per-scene by construction.
The fourth shape is the one the template audit produced and is the least expected: the alternative did not exist, and finding that out was the whole of the work. A choice named as A rather than B where B is unavailable was never a choice between A and B; restating it correctly took an afternoon and was worth more than sweeping anything.
What each audit cost to build
The three took very different amounts of work and the ratio is not what it looks like from the outside.
The unit audit was the largest and the most mechanical. A registry of six formulae at a common interface, a calibration, and then a re-implementation of six published quantities so that each could be recomputed. Most of the effort was in the re-implementations and in the check that each reproduces its own source, which is where the round’s one real error was caught.
The diagonality audit was the smallest and the hardest to see. Its arithmetic is a ladder of six matrices scored over the census, which is an afternoon; the work was in noticing that a parameter count has two kinds of parameter in it. Nothing was built that could not have been built two rounds ago.
And the template audit was almost all reading. Four candidate curves, a refitted matrix each, one quantity scored — an hour of code — preceded by working out that the alternative named in the original sentence does not exist. That restatement is the finding and it took longer than everything after it.
The pattern is that the cheap part is the computation and the expensive part is deciding what to vary. Which is the reverse of the intuition, and is the practical reason a multiverse analysis is worth doing: the multiverse is cheap, and the framing that makes it meaningful is where the judgement goes.
Where the model stops
Three structural choices are three, and there are more. The basis a diagonal is taken in has been audited at length and is a choice; the set a mean is over was the previous round’s subject; the wavelength grid is named above. Beyond those there is a longer tail: the decision to work in tristimulus space at all, the decision that a surface is a reflectance rather than a bidirectional distribution, the decision that an observer is a set of three curves.
Each of those is a choice and each is less auditable than the last, because at some point varying a choice means writing a different collection. The three reached here were reachable because a menu of alternatives existed and somebody had published it. A choice with no alternatives in the literature is not a choice this method can reach, and the honest name for that boundary is the discipline’s own vocabulary rather than any property of the arithmetic.
Those are the three choices themselves. The last picture is the one place two of the round’s instruments were turned on the same object, which makes it the only agreement in the set and the only entry that could have been a disagreement.
Who found it, and when
The three choices were named by this collection two rounds ago and the naming was the useful part. Nothing here required an idea that was not available then; it required treating three fixed things as variables, which needed a menu in each case and got one.
The general form has a name outside this subject. Multiverse analysis is the practice of computing a result under every defensible combination of the arbitrary choices inside it and reporting the distribution rather than a point, and it has had about a decade in the social sciences. Its standard finding is the one reached here: the construction moves the answer more than the sampling does, and a paper reporting a confidence interval is reporting the smaller of the two exposures.
What a computational collection adds is that the multiverse is cheap. Six units over six quantities is under a second; five templates is forty milliseconds; the whole round’s arithmetic runs in the time a single figure takes to draw. The cost of not doing it was never the computation.
Where the ladder goes next
The audit has been pointed at the collection’s inputs, its sets and now its conventions, and each round has found the newest instrument the largest exposure. The obvious next target is the one thing all three rounds have held fixed: that a surface is a reflectance and a light is a spectral power distribution, so that a colour is an integral. Where that stops being true — fluorescence, gloss, translucency, anything the geometry decides — the collection has essays and no audit.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The coincidence was a mechanism calibration · chromatic adaptation · colour difference · elasticity · sensitivity · test set
- Saturation is nearly everything chromatic adaptation · colour difference · elasticity · sensitivity · test set
- Two instruments and one ranking calibration · chromatic adaptation · colour difference · test set · uncertainty
- A chart decides what a camera scores colour difference · elasticity · sensitivity · test set
- The census in six units calibration · chromatic adaptation · colour difference · test set
- The input nobody declared chromatic adaptation · elasticity · sensitivity · test set
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CalibrationChromatic adaptationColour differenceElasticityHeld-outPigment templateSensitivityStructural choiceTest setUncertainty