Where the model breaks

What would have to be wrong

A great many statements here have thresholds written into them, which turns out to make an audit possible — for each one, the smallest change in a declared input that would stop it holding. Most are unreachable. One is inside a factor of one and a third.

Assumes Which measurement is worth making, Twenty-five is a sample of the diagram and An extremum is not a sample.

A collection that writes its claims as thresholds can be asked a question a collection of numbers cannot: which of them is nearest to being false.

Room is not safety: two orderings of the same three claims. Three pairs of bars, one pair per published statement about the confusion points. The upper bar in each pair is the margin — how far the measured number is from the threshold that makes the statement true, as a ratio. The lower bar is the headroom — the factor by which one declared width of the population model would have to be wrong for the statement to fail. Both start at one, which is the line. Ordered by margin the three read the protan margin, the tritan margin, the deutan margin; ordered by headroom they read the protan margin, the deutan margin, the tritan margin, and the middle two change places. Every one of the three is inside a factor of two of failing, which the margins do not say.
Fig. 1 Three statements this collection makes about the confusion points, each with the margin between its measured value and its own threshold, and the factor by which one declared input would have to be wrong for it to fail.

The claim

Every claim here that has a threshold has been audited for the smallest change in a declared input that would break it. Most cannot be broken at all. One goes at a factor of 1.34, and it is not the one the numbers would have nominated.

  • The audit is possible because the claims are already thresholds. This site writes its assertions with a line in them — outside two standard deviations, inside four, within two units — so each one has an edge to measure a distance to.
  • Six claims about the population were audited against four declared widths. Three are statements with thresholds and three are quantities without them; one of the three statements fails at 1.34 times its most dangerous input.
  • One claim cannot be broken by any width inside a factor of eight, and it is the control: a statement about the median member, which no width can move by construction.
  • The claim nearest to failing is not the claim that responds most. The largest elasticity in the table belongs to a different statement, which starts further from its own line.
  • And three other audits ran this round on entirely different machinery — the uniformity ranking, the adaptation census, the searched worst case — with the same shape of answer each time: the extremes hold and the middles do not.

Why this collection can be audited at all

The audit needs something most bodies of work do not have, and it is worth saying what.

A claim has to have an edge. The median disagreement is 9.34 ΔE00 is a number, and asking what would have to be wrong for it to be false is meaningless — any change makes it false and no change makes it importantly false. The nearest published transform’s protan point is at least two standard deviations outside the population’s cloud is a statement, with a line that can be crossed, and the distance to that line is a quantity.

This site has such lines everywhere, and not for this reason. They are there because every claim that two things are identical is an assertion in code and an assertion needs a criterion. Writing at least two σ rather than comfortably outside was a decision about testability made a dozen phases ago, and it turns out to be exactly what a sensitivity audit needs to attach itself to.

A collection that reports numbers cannot be audited this way. That is worth stating because it is the strongest argument this site has yet produced for its own habit, and it arrived as a side effect.

The audit, claim by claim

Four declared inputs — the lens as an age range, the macular pigment’s spread, the cone density’s spread, and the pigment peaks’ spread. Six published quantities. Each input is multiplied by a factor, each claim recomputed, and the factor bisected until the claim crosses its line or the search runs out at a factor of eight.

  • The protanope’s point at two σ — measured 2.24 — fails at 1.34 times the pigment-peak spread.
  • The deuteranope’s point at two σ — measured 3.74 — fails at 1.55 times the same spread.
  • The tritanope’s point inside four σ — measured 2.60 — fails at 1.79 times the lens range.
  • The median member within two units of the standard observer — measured 0.97 — fails at nothing. No width inside a factor of eight moves it, exactly.
  • The three quantities without thresholds move by a tenth at factors of 1.11, 1.24 and 1.64 on the lens range, which is the same information in a form with no consequence attached.
How far each published number moves when a declared width does. A grid of bars, one row per published conclusion and one bar in each row per declared width of the population model. A bar's length is the elasticity — the proportional change in the conclusion for a proportional change in that width — so a bar of length one means an answer that doubles when the width doubles. The largest here is 0.92 and most are between a tenth and a half. One row is empty: the median member's distance from the standard observer does not respond to any width at all, exactly, because scaling a width leaves every median where it was. That row is the control and its being exactly flat is what says the rest are measuring spread.
Fig. 2 Every published conclusion against every declared width, as an elasticity. The row with no bars is the control; the largest bar in the table is not on the claim with the least headroom.

All three thresholds are inside a factor of two. That is the sentence to carry away. The margins — 1.12, 1.87 and 1.54 times their lines — read as though one claim were tight and two were comfortable, and the audit says all three are within a plausible revision of one number.

One row of that list deserves a second look because it is the audit’s own control. The median member’s distance from the standard observer is immovable — not merely far from its line, but unreachable by any width, in every digit. That is not luck: the residual on this collection’s fitted cone matrix is a statement about the median member, and scaling a width leaves every median exactly where it was. A claim that cannot be moved by the inputs being audited is the audit’s own evidence that it is measuring spread rather than accidentally re-fitting something.

The most exposed claim on this site

The protanope’s row deserves its own paragraph, because it is the answer to the question in the title.

The claim that a fitted adaptation transform is not a set of cone responses rests on two of the three confusion points being far outside the population’s own clouds, and on the third not being — which is why the essay is careful to name the points it rests on. The protan point is the thinnest of the two supports, and a pigment-peak spread a third larger than the declared one takes it under the line.

Whether a third larger is plausible is a question about microspectrophotometry and genetics that this collection cannot settle by computation, which is exactly why the number is reported as a factor rather than as a probability. A reader who thinks the peaks are known to a nanometre should read the claim as safe. One who thinks two nanometres is defensible should read it as marginal.

Every width at the wide end of its span, and at the narrow end. One line per published quantity, each spanning the value it takes when all four declared widths are read at the narrow end of their reported ranges to the value at the wide end, with a marker at the value as declared. The largest span is the deutan margin at a factor of 2.62; the smallest is 1.22. This is the reading the population model's own documentation promised for four phases and nothing ever took. It is not a confidence interval — the four ends are not quantiles and the widths are not independent draws — it is what a reader who distrusts all four at once sees.
Fig. 3 Every published population quantity read with all four widths at the wide ends of their spans and again at the narrow ends. The three σ-distances are the rows that move most.

Two of the collection’s other constants have the same shape and are worth reading beside these, because a headroom is only interesting against a comparable one.

Which width carries the answer, and which carries the doubt. Two columns of bars over the four things that differ between two pairs of eyes. On the left, the share of the population's disagreement each one accounts for — the attribution quoted here since the population was built, which puts the lens first at 81%. On the right, how much of the doubt each one puts on everything published here, which is its elasticity multiplied by how badly the width itself is known. The macular pigment comes first there, at 0.65 against the lens's 0.60 — a lead of 8%. The two lists agree exactly below the top.
Fig. 4 Each declared width by how much of the collection’s doubt it carries. The ranking is not the ranking of elasticities, because a small elasticity on a wide declaration outweighs a large one on a narrow declaration.
A mid-grey's lightness across a continuum of rooms. The lightness a mid-grey is predicted to have, plotted along the continuous surround parameter running from an average room to a dark one. The three rooms the standard tabulates are marked on it: average at the left, dark at the right, and dim 61% of the way between them rather than halfway. The whole span is 9.71 units of lightness and the step from average to dim is 5.64 of it — 58% — so choosing one of the three rows is a decision worth most of the range.
Fig. 5 And the surround, swept along the continuum the standard tabulates three points of. It has no declared width at all, which is what puts it outside every column of the table above.

The right response to a small headroom is not to stop believing the claim. It is to find a form of it that does not depend on the width. The underlying observation — that the published transforms are nowhere near anybody’s receptors, on two points of three — does not evaporate at 1.99 σ; what fails is one way of saying it. That repair is available and is not attempted here, because saying clearly which sentence is exposed is the more useful thing to publish first.

Which inputs can break which claim

Its most dangerous input suggests one width matters and the others do not, and for the claim that breaks first the table says otherwise.

The factor by which each width would have to be enlarged to break each statement:

statement margin lens macular density peaks
standard 2.065
protanNear 1.122 1.38 1.34 4.63 1.34
deutanNear 1.872 1.66 4.12 1.55
tritanFar 1.539 1.79

The claim that breaks first is close to its line in three directions at once. The macular pigment and the pigment peaks both break protanNear at 1.34 and the lens breaks it at 1.38 — three of the four widths inside a factor of 1.4, with only the cone density far off at 4.63. That is a different situation from one fragile dependency: the statement sits near its edge whichever way it is pushed.

The other two statements have the opposite shape. tritanFar can be broken only by the lens, at 1.79, and is immune to the other three inside a factor of eight. deutanNear is immune to the lens altogether and goes on the pigment peaks at 1.55.

So each statement carries its own profile of immunities, and the profiles do not resemble one another. A claim about the protan point is sensitive to nearly everything and sits near its line; a claim about the tritan point is sensitive to one thing and is comfortably clear of it; a claim about the median observer is immune by construction to all four.

That is worth more than the headline number on its own. 1.34 is not the distance to a weak point. It is the distance to the nearest of three, on the one claim here with no strong direction to lean on.

The other three audits

The same question was asked this round of three other pieces of machinery, each with its own declared inputs, and the answers rhyme.

The uniformity ranking. Eight colour spaces ordered by how nearly they make MacAdam’s ellipses circles. Three of the seven adjacent pairs survive the sampling error of the twenty-five ellipses; the other four do not, and one changes places in a third of resamples. The ends of the table are ordered and its middle is not a ranking.

The adaptation census. Five transforms ranked by their mean residual over fourteen changes of light, five of which are constructed rather than measured. Perturbing those five by amounts plausible in their own units moves the mean residual by up to forty per cent and never changes which transform wins; it does swap the second and third places, which the table separates by six parts in a thousand.

The searched worst case. A worst change of light is a number about the bound it was searched under, and three nested bounds give 28.4, 21.9 and 21.9 ΔE00. The third does not bite, because the worst wall turns out to be dark rather than saturated.

The shape all four share

Read together the four audits say one thing, and it is more useful than any of them alone.

The extremes of every result here are robust and the middles are not. The best and worst colour spaces are separated by five standard errors and the four in between are not ordered. The best adaptation transform wins under every perturbation and the second and third are interchangeable. The two ends of the confusion-point argument hold and the marginal one is marginal.

That is not a coincidence and it is not a fact about this collection. It is what happens when a ranking is produced by a continuous measurement over a set of candidates that were not designed to be far apart. The extremes are extreme because something real puts them there; the middle is a cloud because nothing does.

The winner survives the census's own construction; the middle of it does not. One row per perturbation of a constant the adaptation census is built from — the imaginary wall's centre wavelength, its width, its depth, its base, the macular filter's density and the two lens ages — each moved by an amount plausible for that quantity in its own units, up and down, and then all of them together. Each row shows where the five published transforms rank under it. Bradford holds the first column in all 14 rows. The second and third columns, which the table as built separates by six parts in a thousand, change places in 2 of them — so that ordering was never a fact about the transforms.
Fig. 6 The census’s constructed constants perturbed one at a time and together, and the ranking of five transforms under each. Everything left of the marked line is the same transform in every row.

The practical form of the lesson is a reading habit rather than a method. When a result is a position in a ranking, ask how far the neighbours are. When it is a comparison between the extremes of a set, it is probably safe. When it is this one beat that one, look for the error bar before repeating it.

What an audit is not

Three misreadings are available here and all three are worth blocking, because an audit is the kind of document that gets summarised badly.

It is not a list of errors. Nothing in it is wrong. Every claim examined holds at the inputs as declared, and the collection came out of the exercise with one sentence marked as exposed and everything else confirmed.

It is not a ranking of importance. A claim with a large headroom is not more important than one with a small one, and the three σ-distances that dominate this essay are not the three most interesting things on the site. They are the three that have thresholds and declared inputs on the same page, which is a fact about how they happen to be written.

And it is not complete. The audit covers what can be swept. Which of these is a convention is the essay about the choices that have no multiplier — a wavelength grid, a white point, a difference formula — and those are structural in exactly the way the four widths are not.

What was computed, and how

Every headroom is a bisection on the logarithm of a multiplier, twenty-two halvings over a factor of eight either way, with the claim recomputed at each trial. A claim that does not cross inside that range is reported as unreachable rather than as a large number, because an extrapolated factor of fourteen is not a statement about anything.

The search does not assume monotonicity beyond checking that the far end has crossed. A claim that crossed and came back would be reported at its first crossing, which is the conservative reading and matches what a threshold means.

One constraint on the arithmetic is worth stating because it bounds the whole exercise: the age range clamps at zero years above a multiplier of 1.8, past which the scaling is a population that is both wider and older. Every factor reported here is checked against that line rather than assumed to be inside it, and all of them are.

Where the model stops

A headroom is not a probability. It says how far an input would have to move, not how likely it is to be there. Converting one into the other needs a distribution over the inputs and this collection does not have one.

Only declared inputs are audited. The four widths, the census’s five constructed constants, the three bounds and the twenty-five ellipses are the things with numbers attached; the structural choices are not audited and mostly cannot be. That the population is built on a pigment template rather than on the physiological fundamentals, that a colour difference is measured with one formula rather than another, that adaptation is modelled as a diagonal at all — none of those has a multiplier to sweep.

The four audits are not commensurable. A factor on a population’s width, a relative error on a fitted ellipse, a nanometre on a wall’s centre wavelength and a bound on a band’s width are four different kinds of quantity, and nothing here combines them. A reader wanting how wrong could this collection be overall will not find it: each audit conditions on everything else being right, and the joint question needs a joint distribution nobody has.

And an audit of thresholds finds only what has a threshold. Half of what this collection says is prose, and a claim in prose that nothing tests is a sample of size zero — a lesson the previous round learned by finding exactly such a sentence, in a figure’s description, wrong for two phases. Nothing here reaches those, and the number of them is unknown.

The generalisation

The move generalises past this collection and past colour, and it is worth stating as an instruction rather than as an observation.

Write claims with edges in them, and the audit becomes possible. A body of work that reports quantities can be given error bars; a body of work that reports statements with thresholds can be asked which of its statements is nearest to failing, and that is a far more actionable question. The cost is small and comes at writing time: choosing a criterion instead of an adverb.

The second half is the ordering result, and it is the surprise. Rank claims by how near they are to failing, not by how sensitive they are. A sensitive claim that started a long way from its line is safer than an insensitive one that started close, and a table of sensitivities points at the wrong sentence — which it does here, in a table with only six rows in it.

The third half, if there can be one, is about what an audit is for. It did not find an error. Every claim examined is true as stated, at the inputs as declared, and the collection is in better shape after the audit than a reader would have guessed. What it produced is a map of where the weight is, which is the thing a reader needs in order to know which sentences to check against their own knowledge and which to take on trust.

Who found it, and when

Sensitivity analysis is old and formal; the machinery here — one-at-a-time perturbation, elasticities, a threshold search — is the simplest version of it and is standard in every applied field that models anything. The specific inversion, reporting the input change needed to flip a conclusion rather than the output change caused by an input, appears under several names: a break-even analysis in economics, a tipping-point analysis in policy modelling, an E-value in epidemiology, where it is the amount of unmeasured confounding that would explain away an observed effect.

The E-value is the closest relative and its motivation is identical: an unmeasurable quantity, a conclusion that depends on it, and a decision to report the threshold rather than to guess the value. It was proposed in 2017 and its uptake is a good argument that the move is under-used rather than unknown — the arithmetic was always available.

Where the ladder goes next

Four things were audited this round and all four are quantities the collection computes. The next question is about the ones it does not: a bound, a box, a set of candidate objectives — the choices that decide what a search is even searching over.

A worst case turns out to be a statement about the tightest bound anybody was willing to state, and the bound that everybody would nominate is not the one that binds. That is the same result as this essay’s, in a field where the declared inputs are pigments rather than eyes.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 11 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Chromatic adaptationConfusion pointConvergenceDeclared inputDegrees of freedomElasticityMacAdam's ellipsesSampling