What would have to be wrong
Assumes Which measurement is worth making, Twenty-five is a sample of the diagram and An extremum is not a sample.
A collection that writes its claims as thresholds can be asked a question a collection of numbers cannot: which of them is nearest to being false.
The claim
Every claim here that has a threshold has been audited for the smallest change in a declared input that would break it. Most cannot be broken at all. One goes at a factor of 1.34, and it is not the one the numbers would have nominated.
- The audit is possible because the claims are already thresholds. This site writes its assertions with a line in them — outside two standard deviations, inside four, within two units — so each one has an edge to measure a distance to.
- Six claims about the population were audited against four declared widths. Three are statements with thresholds and three are quantities without them; one of the three statements fails at 1.34 times its most dangerous input.
- One claim cannot be broken by any width inside a factor of eight, and it is the control: a statement about the median member, which no width can move by construction.
- The claim nearest to failing is not the claim that responds most. The largest elasticity in the table belongs to a different statement, which starts further from its own line.
- And three other audits ran this round on entirely different machinery — the uniformity ranking, the adaptation census, the searched worst case — with the same shape of answer each time: the extremes hold and the middles do not.
Why this collection can be audited at all
The audit needs something most bodies of work do not have, and it is worth saying what.
A claim has to have an edge. The median disagreement is 9.34 ΔE00 is a number, and asking what would have to be wrong for it to be false is meaningless — any change makes it false and no change makes it importantly false. The nearest published transform’s protan point is at least two standard deviations outside the population’s cloud is a statement, with a line that can be crossed, and the distance to that line is a quantity.
This site has such lines everywhere, and not for this reason. They are there because every claim that two things are identical is an assertion in code and an assertion needs a criterion. Writing at least two σ rather than comfortably outside was a decision about testability made a dozen phases ago, and it turns out to be exactly what a sensitivity audit needs to attach itself to.
A collection that reports numbers cannot be audited this way. That is worth stating because it is the strongest argument this site has yet produced for its own habit, and it arrived as a side effect.
The audit, claim by claim
Four declared inputs — the lens as an age range, the macular pigment’s spread, the cone density’s spread, and the pigment peaks’ spread. Six published quantities. Each input is multiplied by a factor, each claim recomputed, and the factor bisected until the claim crosses its line or the search runs out at a factor of eight.
- The protanope’s point at two σ — measured 2.24 — fails at 1.34 times the pigment-peak spread.
- The deuteranope’s point at two σ — measured 3.74 — fails at 1.55 times the same spread.
- The tritanope’s point inside four σ — measured 2.60 — fails at 1.79 times the lens range.
- The median member within two units of the standard observer — measured 0.97 — fails at nothing. No width inside a factor of eight moves it, exactly.
- The three quantities without thresholds move by a tenth at factors of 1.11, 1.24 and 1.64 on the lens range, which is the same information in a form with no consequence attached.
All three thresholds are inside a factor of two. That is the sentence to carry away. The margins — 1.12, 1.87 and 1.54 times their lines — read as though one claim were tight and two were comfortable, and the audit says all three are within a plausible revision of one number.
One row of that list deserves a second look because it is the audit’s own control. The median member’s distance from the standard observer is immovable — not merely far from its line, but unreachable by any width, in every digit. That is not luck: the residual on this collection’s fitted cone matrix is a statement about the median member, and scaling a width leaves every median exactly where it was. A claim that cannot be moved by the inputs being audited is the audit’s own evidence that it is measuring spread rather than accidentally re-fitting something.
The most exposed claim on this site
The protanope’s row deserves its own paragraph, because it is the answer to the question in the title.
The claim that a fitted adaptation transform is not a set of cone responses rests on two of the three confusion points being far outside the population’s own clouds, and on the third not being — which is why the essay is careful to name the points it rests on. The protan point is the thinnest of the two supports, and a pigment-peak spread a third larger than the declared one takes it under the line.
Whether a third larger is plausible is a question about microspectrophotometry and genetics that this collection cannot settle by computation, which is exactly why the number is reported as a factor rather than as a probability. A reader who thinks the peaks are known to a nanometre should read the claim as safe. One who thinks two nanometres is defensible should read it as marginal.
Two of the collection’s other constants have the same shape and are worth reading beside these, because a headroom is only interesting against a comparable one.
The right response to a small headroom is not to stop believing the claim. It is to find a form of it that does not depend on the width. The underlying observation — that the published transforms are nowhere near anybody’s receptors, on two points of three — does not evaporate at 1.99 σ; what fails is one way of saying it. That repair is available and is not attempted here, because saying clearly which sentence is exposed is the more useful thing to publish first.
Which inputs can break which claim
Its most dangerous input suggests one width matters and the others do not, and for the claim that breaks first the table says otherwise.
The factor by which each width would have to be enlarged to break each statement:
| statement | margin | lens | macular | density | peaks |
|---|---|---|---|---|---|
standard |
2.065 | ∞ | ∞ | ∞ | ∞ |
protanNear |
1.122 | 1.38 | 1.34 | 4.63 | 1.34 |
deutanNear |
1.872 | ∞ | 1.66 | 4.12 | 1.55 |
tritanFar |
1.539 | 1.79 | ∞ | ∞ | ∞ |
The claim that breaks first is close to its line in three directions at once. The macular pigment and the pigment peaks both break protanNear at 1.34 and the lens breaks it at 1.38 — three of the four widths inside a factor of 1.4, with only the cone density far off at 4.63. That is a different situation from one fragile dependency: the statement sits near its edge whichever way it is pushed.
The other two statements have the opposite shape. tritanFar can be broken only by the lens, at 1.79, and is immune to the other three inside a factor of eight. deutanNear is immune to the lens altogether and goes on the pigment peaks at 1.55.
So each statement carries its own profile of immunities, and the profiles do not resemble one another. A claim about the protan point is sensitive to nearly everything and sits near its line; a claim about the tritan point is sensitive to one thing and is comfortably clear of it; a claim about the median observer is immune by construction to all four.
That is worth more than the headline number on its own. 1.34 is not the distance to a weak point. It is the distance to the nearest of three, on the one claim here with no strong direction to lean on.
The other three audits
The same question was asked this round of three other pieces of machinery, each with its own declared inputs, and the answers rhyme.
The uniformity ranking. Eight colour spaces ordered by how nearly they make MacAdam’s ellipses circles. Three of the seven adjacent pairs survive the sampling error of the twenty-five ellipses; the other four do not, and one changes places in a third of resamples. The ends of the table are ordered and its middle is not a ranking.
The adaptation census. Five transforms ranked by their mean residual over fourteen changes of light, five of which are constructed rather than measured. Perturbing those five by amounts plausible in their own units moves the mean residual by up to forty per cent and never changes which transform wins; it does swap the second and third places, which the table separates by six parts in a thousand.
The searched worst case. A worst change of light is a number about the bound it was searched under, and three nested bounds give 28.4, 21.9 and 21.9 ΔE00. The third does not bite, because the worst wall turns out to be dark rather than saturated.
The shape all four share
Read together the four audits say one thing, and it is more useful than any of them alone.
The extremes of every result here are robust and the middles are not. The best and worst colour spaces are separated by five standard errors and the four in between are not ordered. The best adaptation transform wins under every perturbation and the second and third are interchangeable. The two ends of the confusion-point argument hold and the marginal one is marginal.
That is not a coincidence and it is not a fact about this collection. It is what happens when a ranking is produced by a continuous measurement over a set of candidates that were not designed to be far apart. The extremes are extreme because something real puts them there; the middle is a cloud because nothing does.
The practical form of the lesson is a reading habit rather than a method. When a result is a position in a ranking, ask how far the neighbours are. When it is a comparison between the extremes of a set, it is probably safe. When it is this one beat that one, look for the error bar before repeating it.
What an audit is not
Three misreadings are available here and all three are worth blocking, because an audit is the kind of document that gets summarised badly.
It is not a list of errors. Nothing in it is wrong. Every claim examined holds at the inputs as declared, and the collection came out of the exercise with one sentence marked as exposed and everything else confirmed.
It is not a ranking of importance. A claim with a large headroom is not more important than one with a small one, and the three σ-distances that dominate this essay are not the three most interesting things on the site. They are the three that have thresholds and declared inputs on the same page, which is a fact about how they happen to be written.
And it is not complete. The audit covers what can be swept. Which of these is a convention is the essay about the choices that have no multiplier — a wavelength grid, a white point, a difference formula — and those are structural in exactly the way the four widths are not.
What was computed, and how
Every headroom is a bisection on the logarithm of a multiplier, twenty-two halvings over a factor of eight either way, with the claim recomputed at each trial. A claim that does not cross inside that range is reported as unreachable rather than as a large number, because an extrapolated factor of fourteen is not a statement about anything.
The search does not assume monotonicity beyond checking that the far end has crossed. A claim that crossed and came back would be reported at its first crossing, which is the conservative reading and matches what a threshold means.
One constraint on the arithmetic is worth stating because it bounds the whole exercise: the age range clamps at zero years above a multiplier of 1.8, past which the scaling is a population that is both wider and older. Every factor reported here is checked against that line rather than assumed to be inside it, and all of them are.
Where the model stops
A headroom is not a probability. It says how far an input would have to move, not how likely it is to be there. Converting one into the other needs a distribution over the inputs and this collection does not have one.
Only declared inputs are audited. The four widths, the census’s five constructed constants, the three bounds and the twenty-five ellipses are the things with numbers attached; the structural choices are not audited and mostly cannot be. That the population is built on a pigment template rather than on the physiological fundamentals, that a colour difference is measured with one formula rather than another, that adaptation is modelled as a diagonal at all — none of those has a multiplier to sweep.
The four audits are not commensurable. A factor on a population’s width, a relative error on a fitted ellipse, a nanometre on a wall’s centre wavelength and a bound on a band’s width are four different kinds of quantity, and nothing here combines them. A reader wanting how wrong could this collection be overall will not find it: each audit conditions on everything else being right, and the joint question needs a joint distribution nobody has.
And an audit of thresholds finds only what has a threshold. Half of what this collection says is prose, and a claim in prose that nothing tests is a sample of size zero — a lesson the previous round learned by finding exactly such a sentence, in a figure’s description, wrong for two phases. Nothing here reaches those, and the number of them is unknown.
The generalisation
The move generalises past this collection and past colour, and it is worth stating as an instruction rather than as an observation.
Write claims with edges in them, and the audit becomes possible. A body of work that reports quantities can be given error bars; a body of work that reports statements with thresholds can be asked which of its statements is nearest to failing, and that is a far more actionable question. The cost is small and comes at writing time: choosing a criterion instead of an adverb.
The second half is the ordering result, and it is the surprise. Rank claims by how near they are to failing, not by how sensitive they are. A sensitive claim that started a long way from its line is safer than an insensitive one that started close, and a table of sensitivities points at the wrong sentence — which it does here, in a table with only six rows in it.
The third half, if there can be one, is about what an audit is for. It did not find an error. Every claim examined is true as stated, at the inputs as declared, and the collection is in better shape after the audit than a reader would have guessed. What it produced is a map of where the weight is, which is the thing a reader needs in order to know which sentences to check against their own knowledge and which to take on trust.
Who found it, and when
Sensitivity analysis is old and formal; the machinery here — one-at-a-time perturbation, elasticities, a threshold search — is the simplest version of it and is standard in every applied field that models anything. The specific inversion, reporting the input change needed to flip a conclusion rather than the output change caused by an input, appears under several names: a break-even analysis in economics, a tipping-point analysis in policy modelling, an E-value in epidemiology, where it is the amount of unmeasured confounding that would explain away an observed effect.
The E-value is the closest relative and its motivation is identical: an unmeasurable quantity, a conclusion that depends on it, and a decision to report the threshold rather than to guess the value. It was proposed in 2017 and its uptake is a good argument that the move is under-used rather than unknown — the arithmetic was always available.
Where the ladder goes next
Four things were audited this round and all four are quantities the collection computes. The next question is about the ones it does not: a bound, a box, a set of candidate objectives — the choices that decide what a search is even searching over.
A worst case turns out to be a statement about the tightest bound anybody was willing to state, and the bound that everybody would nominate is not the one that binds. That is the same result as this essay’s, in a field where the declared inputs are pigments rather than eyes.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- How far a quadratic can be believed chromatic adaptation · convergence · declared input · degrees of freedom
- A constraint is a direction and a distance chromatic adaptation · confusion point · degrees of freedom
- A lattice is a quadrature rule chromatic adaptation · convergence · sampling
- A mean has a set under it chromatic adaptation · degrees of freedom · sampling
- How long is the bowl chromatic adaptation · degrees of freedom · sampling
- Only the flat directions keep their names chromatic adaptation · declared input · degrees of freedom
What links here
The 8 essays that link to this one and share the most of its objects, of 11 that link here.
The objects this essay names
Each one links to every other essay that touches it.
Chromatic adaptationConfusion pointConvergenceDeclared inputDegrees of freedomElasticityMacAdam's ellipsesSampling