Where the model breaks

What the audit still cannot reach

Two rounds have now swept every declared width in this collection and one of its structural choices. Three structural choices remain, none of them has a multiplier to sweep, and the reason each resists is different — which makes the list a description of where this kind of audit ends rather than a queue of work.

Assumes The input nobody declared, What would have to be wrong and A theorem about a family.

The round before this one ended by naming three things its audit could not reach, and missed a fourth that was in the same file as the numbers it was sweeping. This round reached the fourth. The three are still there, and each is unreachable for a different reason.

Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.
Fig. 1 What reaching one of them looks like: the same census under five constructions of its test set. The instrument is a second construction rather than a multiplier, and each of the remaining three needs a different one.

The claim

A sensitivity audit reaches quantities that have a magnitude. Three of this collection’s load-bearing choices do not, they resist for three different reasons, and naming the reason is what says what would answer each.

  • That the population is a pigment template rather than the physiological fundamentals. Resists because the alternative is a different model, not a different value.
  • That adaptation is a diagonal at all. Resists because it is a discrete choice with two options, and the collection has measured the alternative without being able to sweep it.
  • That a colour difference is CIEDE2000. Resists because it is the metric every other number is reported in, so changing it changes the units of the whole collection at once.
  • And two things this round reached are worth counting, because they establish that the boundary moves: the test set’s shape, and the family’s dimension.

What was reached, and how

Two structural choices fell this round and neither fell to a multiplier.

The test set’s shape turned out to have three scalars inside it — how bright the surfaces are, how saturated, and how far two modulations may go together. Those are sweepable, and sweeping them gave the largest elasticity in the collection. The lesson is that a structural choice sometimes contains declarable numbers, and the way to find them is to sweep the default arguments rather than the constant tables.

The family’s dimension has no multiplier at all — a set is three-dimensional or it is not — so the instrument was a construction rather than a sweep. Adding a fourth basis function at an amplitude and varying the amplitude turns a discrete assumption into a continuous dial, and the answer was that the idealisation is worth as much as the effect it isolates.

The objects the average is over, and the region they come from. Two panels. On the left, eight of the 125 reflectance spectra in the test set, drawn as reflectance against wavelength from 380 to 780 nanometres — smooth, broad curves between about 0.02 and 0.9, with at most two gentle undulations each, because each is a level times a combination of two cosines. None of them has a narrow feature, because the family has no basis function that could make one. On the right, the region those surfaces come from, drawn in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the constraint that the two depths may not exceed 0.7 in sum, and 25 lattice points inside it. Five levels of each of those pairs is the whole test set. The square's four corners — the most saturated surfaces the two cosines could make — are outside the diamond and are not in the set at all.
Fig. 2 The choice that turned out to have numbers in it: the region the census’s test surfaces are drawn from, in its own two modulation coordinates, with the brightness levels and the L¹ bound that decides how far two modulations may run together. Three scalars, none of them declared until somebody swept them.

So the boundary is not fixed and the method for moving it is the same both times: find the continuous parameter the discrete assumption is a limit of. That is the test to apply to the three below.

The population is a pigment template

Every individual observer this collection models is built by shifting the peaks of a pigment absorption template — a single curve shape, translated in log-wavenumber, that stands in for all three cone photopigments. The alternative is to perturb the physiological fundamentals directly, which is what the CIE’s individual-observer model does.

Why it resists. There is no parameter that interpolates between the two. A template is a functional form with a peak wavelength; a set of fundamentals is a table of measured values. Moving from one to the other replaces the model rather than moving a number in it, and the two do not agree even at their reference points — the template places the protan confusion point at (0.99, 0.20) against a measured (0.7465, 0.2535), which is why this collection uses the template only for differences and takes the location from measurement.

What would answer it. Implementing the second model and comparing. That is a substantial piece of work, it is well-defined, and it would produce a comparison rather than a sensitivity — two numbers rather than a curve. The partial answer already here is the transfer construction itself, which exists precisely because the template cannot be trusted for absolute positions, and whose existence is an admission of the size of the problem.

A template that cannot place a point, fitted three ways. Three rows, one per set of stimuli the pigment template's cone matrix can be fitted over, each listing the three confusion points that matrix implies. The protanope's point wanders from (0.99, 0.20) to (0.76, 0.13) against a measured (0.75, 0.25), and the deuteranope's moves by 22.5 in chromaticity — further than the whole diagram is wide. A copunctal point is where two nearly parallel planes meet, so a template good to a few per cent, which is far more than enough to place a spectrum, is nowhere near enough to place this. It is why the population is built by moving the measured points rather than by deriving them.
Fig. 3 The construction that is not used, and the reason: confusion points read straight off the pigment template, which is what anybody would write first and is a long way from the measured points.

Adaptation is a diagonal

A von Kries transform is a diagonal matrix in some basis: three independent gains, one per channel. Every adaptation number in this collection is a residual after such a gain, and the whole census is a measurement of what a diagonal cannot do.

Why it resists. It is a discrete choice with a small number of alternatives — a diagonal, a full 3×3, a non-linear model — and there is no continuous family between “three gains” and “nine coefficients” that anybody uses. A dial from one to the other could be built (interpolate a fitted full matrix towards its own diagonal) and would be a construction nobody proposes, so its intermediate points would answer no question anybody has.

What has been answered. More than the earlier audit’s own summary suggests. The census reports three numbers per row and one of them is the full-matrix residual, which is exactly zero on this family — so the collection already knows what a full 3×3 buys: everything, and it is a different matrix for every change of light, and an observer has only one mechanism. That is the answer to why not a full matrix, and it is a good one.

What is not answered is the degree. Adaptation is not complete, the appearance model’s own formula puts the degree at 0.941 in an ordinary room, and correcting for it is a factor of 1.74 on the mean residual — which this collection found a round ago and which is a continuous parameter that now has a dial on it. So the diagonal’s form is a discrete assumption and its completeness is a swept one, and separating the two is most of what can be said.

A colour difference is CIEDE2000

Every number in the collection that is a difference is a ΔE*₀₀: every census residual, every camera profile error, every delivery tolerance, every worst case in this round.

Why it resists, and it is the hardest of the three. The metric is not an input to a calculation; it is the unit the calculation reports in. Changing it does not perturb an answer, it re-denominates every answer at once — and the collection’s conclusions are comparisons, so a metric change that scaled everything equally would change nothing and one that did not would change the comparisons in a way that is not separable from what is being compared.

And it is partly self-referential here. The elasticity of a residual to the test set’s saturation is 0.7 rather than 1.0 precisely because CIEDE2000’s chroma weighting compresses at high chroma. So the metric is inside the measurement of the sensitivity to the set, and swapping the metric would change the sensitivity as well as the answer.

What is answered. The collection has measured how differently the metrics rank things — no diagram makes the ellipses circles is that measurement for colour spaces, and the uniformity table is computed under several. What has not been done is recomputing the adaptation census under ΔE*₇₆ and ΔE*₉₄, which is cheap, would produce three tables, and would answer the narrow question does the census’s ordering survive its metric even though it cannot answer the broad one.

That was worth doing, it is done below, and naming a thing as cheap is worth less than doing it.

Three colour-difference formulae, disagreeing. ΔE76, ΔE94 and ΔE2000 for the same 9 pairs of colours. The largest disagreement between ΔE76 and ΔE2000 here is 26.6 units — larger than the threshold usually quoted for a just-noticeable difference, so the choice of formula can decide whether two colours count as matching.
Fig. 4 The metric everything is reported in. Its chroma weighting is why a fixed physical mismatch on a saturated sample earns a smaller reported difference, which is inside every elasticity this round measured.

The absence that is not a choice

One more, and it is different in kind from the three above because it is a missing thing rather than a decision.

There are no measured reflectance data here. Every set of surfaces in the collection is constructed, is stated to be constructed, and is stated wherever the question arises. This round has now measured what that construction is worth — the region’s shape carries an elasticity of 0.7 to 0.99, its dimension carries as much as the effect it isolates, and its rule change reverses a pair at nearly nine standard errors.

So the absence is now quantified rather than merely declared, which is the most this round can do about it. What it cannot do is fill the gap: a measured collection is a measurement somebody else makes, and the honest position is that every set-dependent number here is conditional and the size of the condition is printed.

The corresponding-colour data are the same absence in the adaptation half, for the fourth phase running, and this round makes their absence sharper rather than smaller. The census whose constructed rows have now been audited from four directions is a census of constructed spectra precisely because nothing measured is held.

What the pattern of the three says

They resist in three different ways and the ways are worth separating, because they say what each would need.

choice why it resists what would answer it
the pigment template the alternative is a different model implement the alternative; compare
the diagonal discrete, and the alternative is already known to be perfect separate form from degree; the degree is swept
CIEDE2000 it is the unit, not an input recompute the census under other metrics — done below, and the ordering survives

None of the three needs an instrument that does not exist. Two need a second implementation and one needs three runs of an existing one. That is a different situation from the one the round before this one described, which read as though the three were beyond the method — and the difference is that this round learned, by reaching two of them, that the method for a structural choice is build the alternative rather than find the multiplier.

What a fourth reflectance dimension costs the theorem that a change of light is a matrix. Four rising curves on axes of the fourth dimension's amplitude, left to right, against what is left of daylight to tungsten after the exact 3×3 change-of-light matrix has been applied, in ΔE₀₀. All four begin at exactly zero: on the three-dimensional family the matrix is solved rather than fitted and there is no remainder at all, which is the theorem this collection's adaptation argument is built on. Adding a fourth reflectance dimension breaks it, and how badly depends far more on the fourth function's shape than on its size — at five per cent amplitude the four shapes cost 0.329, 0.572, 0.063, 0.124 ΔE₀₀ respectively, a factor of 9.1 between the dearest and the cheapest. For scale, the smallest von Kries residual anywhere in the census is 0.26 ΔE₀₀, so the cheapest of the four is a quarter of it and the dearest is twice it.
Fig. 5 What building the alternative looks like: a discrete assumption — that the surfaces span three dimensions — turned into a continuous dial by constructing the fourth dimension it excludes.
How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one.
Fig. 6 And what finding the multiplier looks like: a structural choice that turned out to contain three declarable scalars, once anybody looked at the default arguments rather than at the constant tables.

What this round leaves exposed

Two sentences in the collection that are now known to be conditional and are stated flatly.

Every adaptation residual is a level for one region of surfaces. The orderings survive everything tested; the levels carry a few per cent from the quadrature rule, tens of per cent from the region’s shape, and up to half from a change of construction rule. Figures in the new family name the set; older figures do not, and bringing them into line is work this round did not do.

And the worst cases are readings of a declaration. All fourteen sit on the region’s boundary, so the ratio is the robust quantity and the level is not. That is stated in the essays that report them and is not yet stated in the machinery’s own docstrings everywhere it should be.

Where the model stops

This essay is a list and a list is not a result. Its value is that each entry now carries a reason and a remedy rather than only a name, and that two entries were removed from the round before this one’s version of it by doing the remedy rather than by deciding the entry did not matter.

What it cannot do is bound what is not on the list. The round before this one’s list was three items and missed a fourth that was more elastic than everything it had swept; this one is three items and the same argument applies to it. The reliable finding across two rounds is not the content of either list — it is that a list written from inside the codebase misses the assumptions the codebase shares with its own checks.

Why the list gets shorter and never empties

Two rounds have now produced two versions of this list, and the pattern between them is the finding that will still be true after a third.

The round before this one’s list had three entries and missed a fourth that was more elastic than everything it had swept. This round reached that fourth, kept the three, and has no reason to think its own list is complete — the same argument that explains the earlier miss applies here unchanged.

The mechanism is that a list of assumptions is written from inside the model that makes them. An assumption is visible when something in the code marks it as a choice: a name, a comment about a value, a range, a citation. An assumption that has been absorbed into the apparatus — a grid, a domain, a basis, a metric — has no such marker, and the person writing the list is the person for whom it is invisible.

The reliable consequence is that the boundary is where nobody is looking rather than where somebody stopped. So the useful practice is not to extend the list by thinking harder about it; it is to run a mechanical sweep over the things a list would not contain — the default arguments, the constructions, the checks that share an assumption with what they check — and see what comes out. That sweep is what found this round’s entry, and it took an afternoon.

Who found it, and when

The categories are standard in uncertainty quantification and the vocabulary is parametric against structural uncertainty, with the standard observation that the second is larger and the first is what gets reported. Every review of model uncertainty in any field says so; the reason it keeps needing saying is the one this collection has now measured twice, which is that a structural assumption does not look like an assumption from inside the model that makes it.

The specific improvement over the round before this one’s version of this list is the method rather than the entries: a structural choice is reachable when the discrete assumption is exhibited as the limit of a continuous family, and building that family is the work. That was learned by doing it twice, on the test set’s dimension and on its shape, and it is the sentence the next audit should start from.

Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.
Fig. 7 And what survived all of it: the census’s ranking, whose established steps are established under every construction tried. The conclusions the collection actually argues from are comparative, and comparative quantities are what a structural audit leaves alone.

The one that is cheap

Of the three, recomputing the census under other colour-difference metrics is the one that could be done in an afternoon, and it is worth setting out exactly what it would and would not answer.

What it would answer. Whether the census’s ordering survives its metric. Compute every row under ΔE*₇₆, ΔE*₉₄ and ΔE*₀₀, rank the fourteen three times, and compare. That is thirteen adjacencies times three metrics, with the paired errors already available for each, and the output is a three-tier classification exactly like the one the construction audit produced.

What it would predict. ΔE*₇₆ has no chroma weighting at all, so the saturated surfaces would carry more of every mean and the saturation elasticity would rise towards one. Whether that reorders anything is the question; the extremes are separated by factors of ten and should not move, and the middle is already unordered.

What it would not answer. Whether CIEDE2000 is the right metric. Three metrics agreeing is evidence that the ordering is not an artefact of one weighting function; it is not evidence that any of the three corresponds to what a person would report, which is a question about visual data this collection does not hold.

That distinction is the same one the whole essay is about. A structural choice can be shown to matter or not to matter for a particular conclusion, and cannot be shown to be right, and the first is worth doing precisely because the second is not available.

The cheap one, done

Recomputing the census under all three metrics takes seconds rather than an afternoon, and the prediction made for it survives.

row ΔE*₇₆ ΔE*₉₄ ΔE*₀₀ 76/00
a blackbody at the same temperature 0.404 0.242 0.263 1.539
the macular filter 0.456 0.390 0.368 1.241
D65 to D50 0.548 0.414 0.458 1.196
D65 to a three-emitter source 1.189 0.860 0.980 1.214
the ageing lens 1.825 1.077 1.117 1.634
a red wall 1.775 1.347 1.357 1.308
D65 to a white LED 1.725 1.382 1.697 1.016
D65 to tungsten 2.027 1.540 1.635 1.240
a green wall 2.378 1.457 1.722 1.381
D65 to a triphosphor tube 3.231 2.081 2.323 1.391
two bounces off a green wall 4.822 2.832 3.375 1.429

The ordering is very nearly not the metric’s to decide. Spearman between the ΔE*₀₀ ranking of all fourteen rows and the ΔE*₉₄ one is 0.978, and between ΔE*₀₀ and ΔE*₇₆ it is 0.934. Four inversions in the worst case, out of thirteen adjacencies.

And the prediction about where those inversions would sit is exactly right. The three easiest rows — the blackbody at the same temperature, the macular filter, and D65 to D50 — hold the same three places under all three metrics, and so do the three hardest: the two-bounce green wall, the triphosphor tube and the single green wall. Every inversion is in the middle. Under both alternative metrics, D65-to-D40 crosses the three-emitter source, the red wall crosses the halophosphate tube, and tungsten crosses the white LED; ΔE*₇₆ adds the ageing lens crossing the red wall. Those are the adjacencies this collection had already declined to order.

So the narrow question has an answer: the census’s ordering does not depend on its metric, except among rows nothing here claims to order.

What the metric does change

The level, and not by a constant — which is the part worth recording, because it decides whether a number quoted in one metric can be read in another.

ΔE*₇₆ is larger than ΔE*₀₀ on every row, by between 1.016 and 1.634: a spread of sixty per cent across fourteen rows. ΔE*₉₄ is smaller on twelve of the fourteen and larger on two, running from 0.814 to 1.061.

Neither is a rescaling. If either were, that last column would be flat, and a tolerance in one metric could be carried into the other by multiplication. It is not, so it cannot be, which is where the formulae themselves arrive from the other direction.

Which rows move most is informative in its own right. The largest ratios between the 1976 and 2000 formulae are the ageing lens at 1.634 and the blackbody at 1.539 — small, smooth, low-chroma changes, where the older formula has no weighting to apply and the newer one’s lightness and chroma terms are both discounting. The smallest is the white LED at 1.016. So the two metrics agree most closely on the rows that are hardest for every basis, and diverge on the easy ones, which is the reverse of what a reader would guess.

What the run still does not answer

The distinction drawn above holds, and it is worth restating now that a number exists to attach it to. Three metrics agreeing on an ordering is evidence that the ordering is not an artefact of one weighting function. It is not evidence that any of the three corresponds to what a person would report — all three are fits to visual data this collection does not hold, and a shared ancestry is exactly what would make three metrics agree while all three were wrong together.

What the run does establish is a boundary on the third entry in the table above. The metric is the unit; changing it re-denominates every number in the collection at once; and the re-denomination is a factor between 1.0 and 1.6, and reorders nothing the collection argues from. That is less than an answer and considerably more than a name, which is the most this essay claims a structural audit can produce.

Where the ladder goes next

Three named choices with three named remedies, one of them cheap. And a fourth item that is not a choice — the measured data nobody here has — whose cost is now printed on every number that depends on it.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 12 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AuditChromatic adaptationColour differenceDegrees of freedomMeasurement errorModelling assumptionRobustnessSensitivitySpecificationTest set