Where the model breaks

The input nobody declared

An audit that swept every declared width in this collection found the largest elasticity anywhere to be about a half. The most elastic input turns out to be one that was never declared, never quoted with a range and never varied — and being undeclared is exactly why it escaped the audit that was looking for it.

Assumes Saturation is nearly everything, What would have to be wrong and Which measurement is worth making.

An audit sweeps what has a number attached. That is not a limitation anybody chose; it is what a sweep is, and it is why the largest sensitivity in this collection went four rounds without being found.

How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one.
Fig. 1 The elasticity of every adaptation result to the three numbers that describe the test set. The longest bar is 0.99. The longest bar in the equivalent table for the five declared population widths is about a half.

The claim

A sensitivity audit is a survey of the inputs that were written down as inputs, and an input that was written down as a fact escapes it — however elastic it is.

  • The declared widths’ largest elasticity is about 0.5. Five quantities, each quoted with a range from the individual-observer literature, each swept.
  • The undeclared set’s largest elasticity is 0.99, and its smallest is 0.49 — so its floor is at the declared table’s ceiling.
  • The undeclared input also has the wider span. The declared widths differ between studies by a factor of about 1.5 to 2. The test set’s saturation has no reported range at all, because nobody reports it.
  • So it dominates on both terms of the product that ranks how much doubt an input carries.
  • And the previous audit said, correctly, that it could not reach this. It listed three structural choices it had missed and did not list the fourth, which was in the same file as the numbers it was sweeping.
What one change of light costs, surface by surface — daylight to tungsten. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to tungsten is 1.635 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 3.058, which is 1.87 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 2 What the undeclared input decides: how far this distribution reaches. Scaling the set’s saturation by a quarter moves the whole curve by about a fifth, and nothing in the collection had ever moved it.

What an audit actually surveys

The round before this one built an instrument and applied it to everything it could see. For each published conclusion C and each declared input width w it computed the elasticity E = d ln C / d ln w as a central difference over a stated finite span, and then ranked the inputs by E times how badly each is known — because a term can dominate an answer while being known precisely, so that measuring it better buys nothing, while a small term whose width is a guess carries most of the doubt.

That is the right instrument and the ranking it produced was useful. It found that the lens carries most of the population’s spread and that a different variate carries most of the answer’s uncertainty, which is a distinction worth a round on its own.

What it surveyed was a list. Five widths, five census constants, three bounds, twenty-five ellipse axes. Every one of them existed in the code as a named quantity with a value and, usually, a comment about where the value came from — which is exactly the property that made them findable.

What that misses, structurally

An input has to be recognised as an input before it can be swept, and recognition happens at the moment somebody writes the value down with a reason beside it. A quantity written down without a reason does not look like an input; it looks like a definition.

The three numbers describing the test set are the clearest possible case. They are levels, depths and maxDepth, they sit in a function signature as default arguments, and the docstring above them explains at length why the set is a lattice rather than a random sample and says nothing whatever about the values. Read as code they are configuration. Read as modelling they are three of the most consequential declarations in the collection.

The audit’s own closing paragraph is worth quoting because it is nearly right:

No structural choice was audited. That the population is a pigment template rather than the physiological fundamentals, that adaptation is a diagonal at all, that a colour difference is CIEDE2000 — none of those has a multiplier to sweep, and the audit reaches none of them.

Three examples, correctly identified, and each genuinely without a multiplier. The fourth structural choice does have multipliers — three of them, sitting in the same file as several of the constants that were swept — and it was not on the list, because it did not look like a choice.

No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for the macular pigment. The tallest bar is 2.40 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 94.4 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 3 The row with the largest elasticity to the undeclared input, 0.99. Its answer is concentrated at the deep end of the set, which is exactly the part the L¹ reach decides the extent of.

The two tables

what largest elasticity declared range swept before?
the lens, as an age range ~0.5 20–70 years, quoted yes
macular pigment density ~0.3 mean 0.35, sd 0.13, quoted yes
cone outer-segment density ~0.2 0.4 ± 0.036, quoted yes
pigment peak wavelengths ~0.3 ±1.5 nm, quoted yes
the census’s wall centre ~0.2 none yes
the test set’s L¹ reach 0.99 none no
the test set’s saturation 0.91 none no
the test set’s brightness 0.15 none no

The comparison is deliberately not made from one table’s numbers against the other’s inside a single assertion, because a gate whose verdict depends on two files’ machinery is a gate that goes red for the wrong reasons. It is made against a stated floor: the set’s elasticity exceeds 0.6, which is above anything the previous audit found anywhere.

And the second term compounds it. The ranking that matters is elasticity times how badly the input is known. The declared widths at least have literature ranges — a factor of one and a half to two between studies, which is what makes their contribution to the doubt boundable. The test set’s saturation has no range at all: no study reports it, because it is not a quantity anybody outside this collection has an opinion about. An input with a large elasticity and an unbounded span carries unbounded doubt, and the honest thing to say about it is that its contribution cannot be quoted rather than that it is small.

Which width carries the answer, and which carries the doubt. Two columns of bars over the four things that differ between two pairs of eyes. On the left, the share of the population's disagreement each one accounts for — the attribution quoted here since the population was built, which puts the lens first at 81%. On the right, how much of the doubt each one puts on everything published here, which is its elasticity multiplied by how badly the width itself is known. The macular pigment comes first there, at 0.65 against the lens's 0.60 — a lead of 8%. The two lists agree exactly below the top.
Fig. 4 The earlier audit’s ranking: each declared width by how much of the collection’s doubt it carries, which is its elasticity times how badly it is known. Every bar in it is shorter than the undeclared input’s would be, and the undeclared input’s cannot be drawn because its span is not a quantity anybody reports.

Why the shape of the mistake is general

Not a fact about colour, and worth stating so it transfers.

A modelling exercise has three kinds of number in it. Measurements, which come with an uncertainty and are usually treated correctly. Declared inputs, which come with a reason and a range and are what a sensitivity analysis sweeps. And constructions — the grid, the sample, the discretisation, the domain, the basis — which are chosen for tractability, written once, and thereafter read as part of the apparatus rather than as part of the model.

The third kind is systematically the least examined and is not systematically the least important. It has three properties that keep it out of an audit:

  • It has no literature range, so there is nothing obvious to sweep it over. A width with a quoted range invites a sweep; a lattice size does not.
  • Its author had a good reason for the structure and no reason for the values. The docstring explains why the set is a lattice — which is the interesting decision — and the values were then whatever made the lattice a convenient size.
  • It is shared between the result and the check. The assertion that a change of light is exactly a matrix is asserted on the set constructed to have that property, so the strongest-looking check in the neighbourhood cannot see the assumption it shares.

Any one of those would be enough. Together they make a construction nearly invisible from inside the codebase that contains it.

The objects the average is over, and the region they come from. Two panels. On the left, eight of the 125 reflectance spectra in the test set, drawn as reflectance against wavelength from 380 to 780 nanometres — smooth, broad curves between about 0.02 and 0.9, with at most two gentle undulations each, because each is a level times a combination of two cosines. None of them has a narrow feature, because the family has no basis function that could make one. On the right, the region those surfaces come from, drawn in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the constraint that the two depths may not exceed 0.7 in sum, and 25 lattice points inside it. Five levels of each of those pairs is the whole test set. The square's four corners — the most saturated surfaces the two cosines could make — are outside the diamond and are not in the set at all.
Fig. 5 The construction: three basis functions, a lattice of five levels and two modulation depths, and an L¹ bound. Everything visible in this picture was chosen; none of it was declared as an input.
What a fourth reflectance dimension costs the theorem that a change of light is a matrix. Four rising curves on axes of the fourth dimension's amplitude, left to right, against what is left of daylight to tungsten after the exact 3×3 change-of-light matrix has been applied, in ΔE₀₀. All four begin at exactly zero: on the three-dimensional family the matrix is solved rather than fitted and there is no remainder at all, which is the theorem this collection's adaptation argument is built on. Adding a fourth reflectance dimension breaks it, and how badly depends far more on the fourth function's shape than on its size — at five per cent amplitude the four shapes cost 0.329, 0.572, 0.063, 0.124 ΔE₀₀ respectively, a factor of 9.1 between the dearest and the cheapest. For scale, the smallest von Kries residual anywhere in the census is 0.26 ΔE₀₀, so the cheapest of the four is a quarter of it and the dearest is twice it.
Fig. 6 A second undeclared construction in the same file, given a dial: the family’s dimension. It has no multiplier in the ordinary sense, which is why it needed a different instrument rather than a wider sweep.

What to do instead

Three things, and the first is nearly free.

Sweep the defaults, which is where a construction’s own scalars hide. Every default argument in a function that produces a published number is a declaration. Multiplying each of them by 1.25 and 0.8 and recomputing the headline results is an afternoon’s work and would have found this in one run. It is the same instrument the audit already had, pointed at the function signatures rather than at the constant tables.

Say which numbers are constructions, in the file. The three test-set numbers now have a section that names them, states that none of them is quoted from anything, and carries their elasticities. That is not documentation; it is what makes them appear on the next audit’s list.

And be suspicious of a check that shares its assumption. An assertion that a property holds on an object constructed to have it is a consistency check, and this collection has an explicit rule about assertions that have never rejected anything. The rule needs a second half: an assertion that cannot reject a particular class of error should say which class.

What survives, and what has to be restated

Not much has to change, which is worth saying plainly.

Every ordering survives, including the ones the set’s construction was tested against. Scaling the test set’s saturation multiplies every census row by nearly the same factor, so the ranking of changes of light — which is what the collection’s arguments actually rest on — is untouched. The extremes hold under every construction and the middle was never ordered anyway.

Every published level has to be read as a level for this set. A residual of 1.635 ΔE*₀₀ for daylight to tungsten becomes, under a set a quarter more saturated, about 1.95, and under one a fifth less saturated, about 1.40. Those are not corrections to an error; they are the number’s actual precision, and it was being printed to four figures.

And the comparison with the declared widths is the finding rather than the numbers. The specific elasticities will move as the machinery does. What will not move is that the input nobody wrote down as an input was more elastic than every input somebody did.

Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.
Fig. 7 The other undeclared construction in the same file: the rule that samples the region. Its effect is one to five per cent — an order of magnitude smaller than the region’s shape, and it was equally undeclared.
Which fourth dimensions are expensive, and how fine is too fine to matter. Two curves on axes of how many half-cycles a cosine fourth basis function makes across the visible band, against what it costs the matrix theorem in ΔE₀₀, at a fixed ten per cent amplitude. Both curves touch zero at exactly one and two half-cycles: those are the family's own second and third basis functions, so a fourth coefficient along them adds no dimension and a change of light stays exactly a matrix. Between them the cost climbs, reaches a maximum, and — for the smooth source — falls away again, because structure finer than the scale on which three broad cone sensitivities differ integrates to nearly nothing. The two curves part company at the fine end. Under daylight-to-tungsten the cost has fallen by a factor of 2.2 from its peak; under daylight-to-a-triphosphor-tube it has barely fallen at all, because a source with three narrow emission lines has structure of its own at that scale for the surface's structure to beat against. The observer is identical in both curves.
Fig. 8 And what a construction looks like when it is swept properly: two exact zeros where the perturbation coincides with the construction’s own basis, which is the check that the sweep is measuring the construction rather than the arithmetic.

The three properties, as a checklist

The essay’s general claim is that constructions escape audits, and the three reasons are specific enough to be turned into questions somebody can ask of their own code.

Does the quantity have a comment about its value, or only about its structure? The test set’s docstring explains at length why the set is a lattice rather than a random sample and says nothing about how many levels or how deep. A comment that justifies the form and not the numbers is the signature of a construction being treated as apparatus.

Is the quantity a default argument rather than a named constant? Every number in this collection that had been swept lived in an exported table with a name. The three that had not lived in a function signature. That is not a coincidence: a named constant announces itself as a decision and a default argument announces itself as a convenience.

And does any check share the assumption? The assertion that a change of light is exactly a matrix is asserted on the family constructed to make it exact. A check whose truth is guaranteed by the construction it checks is a consistency test, and it is the strongest-looking thing in the neighbourhood.

All three questions are answerable by reading, need no instrument, and would have found this one in an afternoon. The instrument was never the problem; the list it was pointed at was.

Where the model stops

This essay establishes that one structural choice, given multipliers, is more elastic than every declared input. It does not establish that structural choices are in general more elastic than declared ones, and the sample is one.

The three the previous audit named remain unswept, and two of them have no multiplier to sweep in any obvious way. That adaptation is a diagonal at all is a discrete choice, and the natural instrument for it is not an elasticity but a comparison with the alternative — which this collection has run and reports as what no adaptation can remove. That a colour difference is CIEDE2000 is likewise discrete; a partial answer exists in the uniformity table’s own dependence on its metric, and it is partial.

So the honest closing position is that the audit’s reach has been extended by one and the boundary is still there. What the audit still cannot reach is a list rather than a footnote, and it is shorter than it was.

What the two reachable cases had in common

Two structural choices fell this round and the method was different each time, which is worth separating because the difference is the transferable part.

The set’s shape fell to a sweep, because it turned out to contain scalars. The move was to look at default arguments rather than at constant tables, and the instrument was the one the round before this one had already built. That is the cheap case and it is worth trying first on anything that looks structural: a construction often has numbers in it, and the numbers are usually in the signature.

The set’s dimension fell to a construction, because it contains no scalar at all — a family is three-dimensional or it is not. The move was to build the fourth dimension the assumption excludes and give it an amplitude, which turns a discrete assumption into a continuous one and lets the ordinary instrument work on it.

Both are the same manoeuvre at different depths: find the continuous parameter the discrete assumption is a limit of. For the shape it was already there and unnoticed; for the dimension it had to be built. The three remaining structural choices are all in the second category, and naming what would have to be built for each is most of the work of deciding whether to.

Who found it, and when

The distinction between measured, declared and constructed inputs is not novel and turns up wherever numerical modelling meets uncertainty quantification — the standard vocabulary is aleatory, epistemic and numerical uncertainty, and the standard observation is that the third is the one practitioners quantify last and that reviewers ask about least.

Its arrival here has a specific and slightly embarrassing shape. The audit that preceded this one was written to catch exactly this class of failure, its own docstring names the failure precisely — a published number computed at a point, a width, a box or an objective that somebody chose, with nothing having ever asked how far the number moves when the choice does — and it swept everything in the collection that had a number and a comment. The set had a number and a comment about something else.

How far each published number moves when a declared width does. A grid of bars, one row per published conclusion and one bar in each row per declared width of the population model. A bar's length is the elasticity — the proportional change in the conclusion for a proportional change in that width — so a bar of length one means an answer that doubles when the width doubles. The largest here is 0.92 and most are between a tenth and a half. One row is empty: the median member's distance from the standard observer does not respond to any width at all, exactly, because scaling a width leaves every median where it was. That row is the control and its being exactly flat is what says the rest are measuring spread.
Fig. 9 The audit as it stood: five declared widths against the conclusions that rest on them. The instrument is right, the arithmetic is right, and the list it was pointed at was the list of things that looked like inputs.

The floor is not where the claim puts it

The claim section says the undeclared set’s smallest elasticity is 0.49, so that its floor is at the declared table’s ceiling. The table three sections below prints the set’s third number, its brightness, at 0.15 — which is not merely below the declared ceiling of about 0.5, it is below every one of the five declared entries, the lowest of which is 0.2.

The 0.49 is a real figure and it belongs to something narrower: the smallest elasticity across the census’s rows for the two saturation-like parameters. Stated as a property of the undeclared set it is wrong, and the wrongness runs in the direction that flatters the essay’s headline.

The correction matters because it changes what the essay has shown. Two of the three undeclared numbers are more elastic than anything declared; the third is less elastic than anything declared. So being undeclared does not predict being elastic, and the finding is about a maximum rather than about a class — which the closing section already says carefully, and which the opening bullets do not. An audit that had swept the defaults, as this essay recommends, would have found one number at 0.99, one at 0.91 and one it could safely have ignored, and it would have had no way to tell in advance which was which.

The product the ranking is supposed to use

The essay’s own instrument is elasticity multiplied by how badly the input is known, and it argues that the undeclared input dominates on both terms. Worked through, the second term does not support that, and the reason is that a large elasticity applied to a narrow span is a small effect.

Raising each input’s span factor to the power of its elasticity gives how far the published answer moves across that input’s range:

input elasticity span factor answer moves by
the lens, 20 to 70 years 0.5 1.5 – 2.0 ×1.22 – ×1.41
the test set’s L¹ reach 0.99 1.25, as swept ×1.25
the test set’s saturation 0.91 1.25, as swept ×1.22
macular pigment density 0.3 1.37 ×1.10
cone outer-segment density 0.2 1.09 ×1.02
the test set’s brightness 0.15 1.25, as swept ×1.03

At the span the collection actually sweeps, the most elastic undeclared input carries less doubt than the lens does. A factor of two in age raised to a half beats a quarter in saturation raised to one, and it is not close: 1.41 against 1.25.

That does not rescue the declared table, and it is not an argument that the audit was fine. It sharpens what the undeclared input’s problem is. The problem is not that it dominates the doubt; it is that nobody can say whether it does. The ±25 per cent used above is the collection’s own sweep convention, chosen because it is the convention, not because anybody has evidence that the test set’s saturation is known to a quarter. Solving for the span at which the L¹ reach would match the lens’s worst case gives a factor of 1.42 — so the two are level if the set’s saturation is uncertain by about forty per cent, the undeclared input wins above that and loses below it, and there is no measurement anywhere that says which side of it the truth is on.

The honest ranking is therefore not a longer bar than the lens’s. It is a bar with no end drawn on it, next to five bars that have ends, and the reader cannot be told which is longest. That is a worse position to be in than the essay’s version, because the essay’s version at least states an ordering; this one says the ordering is unavailable. An unbounded contribution does not top a ranking — it makes the ranking undefined, and the practical consequence is the one the closing sections already draw: every published level has to be read as a level for this set, and no interval can be put around it.

Where the ladder goes next

The set’s members are not all doing the same work — five of them contribute nothing at all, exactly, and the reason is worth understanding before the set is trusted with anything else.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 14 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AuditChromatic adaptationDegrees of freedomElasticityMeasurement errorModelling assumptionRobustnessSensitivitySpecificationTest set