Where the model breaks

The reference had to be built

A five-nanometre error cannot be measured with five-nanometre data. Interpolating the tables and integrating finely measures the interpolator, not the grid — so the audit of this collection's index had to be run against an observer made of formulae, and the price of that is a residual of 1.42 ΔE₀₀ that every number in the section is read beside.

Assumes The index is a choice too, The population rests on a template and A template is mostly its tail.

Every measurement in this section is a distance from a reference, and the reference is the part that could not be bought. A tabulation cannot be audited against a tabulation, so one had to be constructed, and the construction has a price that belongs in front of the results rather than behind them.

How far this collection's analytic observer is from the tabulated one. The construction every figure in this family uses is three pigment absorptances through one fitted 3×3, and this is the residual of that fit against the CIE's 1931 functions: root-mean-square error as a percentage of each curve's peak, and the colour difference it produces over forty-two surfaces. The short-wavelength function is the worst at 16.4 per cent, which is where a pigment template is weakest and where the ocular media are doing most of the work. The median colour difference is 1.42 ΔE₀₀, so this is an observer of the right shape rather than a copy of the table — and every departure in this family should be read beside that number rather than against zero.
Fig. 1 The residual of this collection’s analytic observer against the tabulated 1931 functions, curve by curve and then in colour. Every number in this round is read beside the last two rows rather than against zero.

The claim

Auditing a wavelength grid requires a source of truth finer than the grid, tabulated data cannot supply one, and building one means replacing the measured observer with a modelled one.

  • Interpolating a coarse table and integrating finely measures the interpolator. The true curve never enters the calculation, and the answer that comes back says the coarse table was excellent.
  • So the observer here is analytic — three pigment absorptances through an ocular-media filter, both closed forms — and can be asked about any wavelength.
  • Its residual against the published functions is 2.3 per cent on ȳ and 16.4 on z̄, which is a median of 1.42 ΔE₀₀ over forty-two surfaces.
  • And that residual is not an error in the results. It is the width of the claim they support: statements about the shape of an eye’s response, rather than about the CIE’s table.

The circularity, stated plainly

The obvious way to find out what five nanometres costs is to take the five-nanometre tables, interpolate them to a tenth of a nanometre, integrate both ways, and subtract.

Every step of that is defensible and the result is worthless. The fine integral is taken over a curve that was manufactured from the coarse one, so it agrees with the coarse one at all eighty-one original points by construction and differs only where the interpolator invented something. The difference measures the invention.

Worse, it measures it in the flattering direction. A good interpolator produces a smooth curve through the points, and a smooth curve integrated finely is close to the same curve summed coarsely — so the answer is small, and the smallness is a property of the fill-in rule’s smoothness rather than of the tabulation’s adequacy. The procedure is guaranteed to report that the grid was fine.

This is not a subtle trap. It is the standard way the question gets asked, and the collection’s own earlier attempt at it sidestepped rather than solved it: that essay compared two coarsenings of the same table against each other, which is a legitimate relative measurement and cannot produce an absolute one.

A replacement had to satisfy three demands. A reference for this audit needs three properties and they are demanding taken together.

It must be evaluable at any wavelength, which rules out anything tabulated and every interpolation of anything tabulated.

It must be the right shape, because a statement about what a tabulation costs the human observer is worthless if the curves are not an observer’s. A set of three arbitrary smooth humps would give perfectly self-consistent answers about a fictional eye.

And its own convergence must be checkable, so that the reference can be shown to be a reference rather than another entry in the table.

The first two pull against each other. Anything analytic is a model, and any model of the colour-matching functions is a fit whose residual has to be published. There is no arrangement in which the reference is both closed-form and the standard observer, because the standard observer is a table of measurements and nothing else. Seventeen people were measured in 1928 and the table is what those measurements became; there is no underlying formula that was discretised, so there is nothing to go back to.

What was available already

The collection did not have to invent the model, which is the one piece of luck in the whole exercise.

This collection has held an analytic pigment template since the foundation phase — the Govardovskii form, a sum of exponentials in λmax/λ with a secondary band — and an analytic ocular-media filter, a lens absorbance and a macular absorbance both written as functions of wavelength. A population of two hundred eyes is drawn from those two, and a whole round has already been spent on what the template’s tail decides.

Both were evaluating onto the site’s eighty-one-point grid because nothing had ever asked them for anything else. Adding an optional list of wavelengths to three functions made the entire apparatus answerable at any resolution, which is about twenty lines of change and no new physics.

One detail in that change is worth recording because it is the kind of thing that quietly ruins a measurement. The pigment template is normalised by its own peak, and reading that peak off a coarse grid would rescale the pigment as well as sampling it — so a coarse request would return a curve that was both differently sampled and differently normalised, and the two effects would arrive added together. The peak is taken from the site’s own grid whatever the caller asks to be sampled on, which keeps the normalisation constant across every grid in the audit.

The three cone absorptances at two settings of the cone optical density. Solid and dashed are the same construction at the two ends of two standard deviations of the reported spread. The curves are built from one pigment template through its ocular media, which is the same model its population of two hundred eyes is drawn from. The largest difference between the two sets is 18.8 per cent of the peak, and where it sits along the wavelength axis is what decides which stimuli the two observers disagree about — a departure concentrated in the blue is invisible on a sample with no blue in it.
Fig. 2 The three cone absorptances at two settings of the cone optical density. Everything drawn here is a closed form in the wavelength, which is what makes a finer grid possible at all.
The three cone absorptances at two settings of the macular pigment. Solid and dashed are the same construction at the two ends of two standard deviations of the reported spread. The curves are built from one pigment template through its ocular media, which is the same model its population of two hundred eyes is drawn from. The largest difference between the two sets is 31.5 per cent of the peak, and where it sits along the wavelength axis is what decides which stimuli the two observers disagree about — a departure concentrated in the blue is invisible on a sample with no blue in it.
Fig. 3 The same three curves at two macular densities. The macular absorbance is a Gaussian centred at 460 nanometres, so its effect is confined to a band, and where a departure sits along the wavelength axis decides which stimuli it can reach.

The two curve figures also make the second demand concrete. A reference that was merely three smooth humps would produce internally consistent answers about a fictional eye, and nothing in the arithmetic would object. What ties these curves to a human observer is not their smoothness but their provenance: each one is a pigment absorbance seen through an ocular filter, both of which are separately measurable objects with their own literatures, and the peaks are the peaks that literature reports.

That provenance is what makes the residual interpretable. A fit that is 16.4 per cent wrong on is a statement about how well a two-parameter pigment model plus a two-parameter filter reproduce a curve measured on seventeen people in 1928 — which is a meaningful thing to be wrong about, in a way that a spline’s residual would not be.

The reference, and its own convergence

The reference is a tenth of a nanometre from 300 to 830, which is 5,301 points, and it is checked rather than declared.

Halving it again — to 0.05 nanometres, 10,601 points — moves the sharpest case in the whole file by 3.4 × 10⁻¹³ ΔE₀₀, on a fluorescent tube through a notch filter. That is the floating-point floor, and it is the difference between a reference and a finer entry in the table being audited.

The sharpest case was chosen deliberately for the check. A convergence test on the smoothest case would pass trivially and say nothing, which is the same failure the origin sweep was written against: a check has to be run where it might fail. The tube’s mercury lines are a nanometre wide, so a tenth of a nanometre is a tenfold oversampling of the narrowest thing in the file.

That is not infinite resolution and is not claimed to be. It is a stated factor above the narrowest feature, with the sensitivity to that factor measured.

What the model costs, in the results

The residual is published rather than buried, and it is larger than a reader might expect.

what residual
x̄ against the fit 7.8% rms of peak
ȳ against the fit 2.3%
z̄ against the fit 16.4%
median colour difference over the family 1.42 ΔE₀₀
worst colour difference over the family 6.33 ΔE₀₀

The short-wavelength function is much the worst, and that is where a pigment template is weakest and where the ocular media are doing nearly all the work. The S cone’s absorbance is narrow, its peak is close to the lens’s absorption edge, and small errors in either move a long way.

So the honest statement about every number in this section is that it is what a tabulation costs an observer of roughly the right shape, not what it costs the CIE’s. The structural conclusions — which lights are safe, how the errors scale, which end of the range is expensive, what cancels — do not depend on the third decimal place. The third decimal place does.

Why that is a fair trade and where it is not

For most of this section the trade is clearly good, because the questions are comparative. Whether the range costs more than the step, whether interpolation helps or harms, whether the origin matters — all of those are ratios between two computations using the same curves, and the curves cancel out of a ratio to first order.

There is one place it is not a fair trade and it should be named. The claim that the ultraviolet end of the range costs three thousand times the infrared end is a statement about the observer’s tails, and the tails are exactly where a template is least trustworthy. The direction of that result is safe — every observer has a shoulder below 400 and an abrupt edge above 700, because one is a filter and the other is an absorption edge — and the factor of three thousand is a property of this construction rather than of the tables.

A version of that measurement against the published 360–830 tables would be worth having and is not available, for the reason the whole essay is about: those tables stop at five nanometres and cannot be asked what lies between their points.

What choosing a space to divide the white out in is worth. Three pairs of routes to the same colour, over forty-two surfaces: dividing the white out in tristimulus values, in a published cone space, and in the observer's own cones. The first two agree to 0.59 ΔE₀₀ at the median. Either of them differs from the observer's own cones by more than fifteen. That is why the two exact conditions in this round are exact only in the eye's own coordinates: the identity belongs to the receptors, and every published arithmetic works in a basis somebody else chose.
Fig. 4 Three routes to the same colour, over forty-two surfaces. The two published ones agree to 0.59 ΔE₀₀, which is a useful comparison for the size of the model residual above.
The arguments a standard observer does not have. Seven choices inside a set of colour-matching functions, each with the shape it takes and what it is worth in ΔE₀₀ on a red pigment under a 6500 K radiator. Six are measurements: a field size, an age, a macular density, a cone optical density, three peak wavelengths and a rod contribution. The seventh is not — a change of basis is a change of curves and not a change of observer, and its entry is exactly zero because the space an experiment measures is what an observer is. Printing that zero beside the others is the clearest statement of what the other six are measurements of.
Fig. 5 The seven arguments this construction takes. Every one of them is a knob the tabulated observer does not have, and the model was built to have them rather than as a convenience.

There is a compensation for the residual that is easy to overlook: the model can be asked questions the table cannot answer at all. A tabulated observer has no age, no field size and no macular density; it is one column of numbers. The construction has seven arguments, and the whole second half of this round consists of varying them.

So the same decision that costs 1.42 ΔE₀₀ of fidelity buys an entire audit that would otherwise be impossible. That is not an accident of this project. A model is what gets built when the questions have outgrown the data, and the cost is always the same shape — a residual against what was measured, in exchange for access to what was not.

A residual is not automatically an error bar

There is a temptation to treat 1.42 ΔE₀₀ as an uncertainty and to attach it to every result, and that would be wrong in both directions.

It is too large for the comparative results, because it cancels: the same wrong curves appear on both sides of every subtraction, and what survives is second order in the residual rather than first. A departure measured at 2.38 ΔE₀₀ between two observers is not uncertain by 1.42; it is uncertain by whatever the residual’s derivative with respect to the departure is, which is much smaller and is not measured here.

And it is too small for the absolute ones. A statement about the shape of ’s tail carries the full 16.4 per cent, not the aggregate 1.42.

A single residual number summarising a model’s fidelity cannot be attached to individual results, and the useful thing is to say which results are comparative and which are absolute. This section’s are almost all comparative, which is why the audit is worth doing on a model at all.

What was computed, and how

The 3×3 that carries cone responses to tristimulus values is fitted once by ordinary least squares over the site’s own grid against the 1931 functions, using the median member of the population as the reference eye. Every observer in the round then uses that same matrix, so a difference between two of them is a difference in what the cones caught and never a difference in bookkeeping.

The residual is computed two ways because one of them would have been misleading alone. The root-mean-square per curve says how well the fit reproduces the functions; the colour difference over forty-two surfaces says what that is worth. They disagree about which curve matters: is by far the worst fit and contributes least to most colours, because ’s support is where most reflectances are dark and most lights are weak.

The gate this family carries requires the short-wavelength function to be the worst-fitted of the three — a claim about where a pigment template fails, which the fit could have contradicted and does not.

Where the model stops

The template is Govardovskii’s, fitted to microspectrophotometry of vertebrate pigments generally rather than to human cones specifically, and its secondary band is a caricature. An earlier round found the tail to be most of what a template is, so the choice of template is load-bearing and only one alternative has been tried.

The ocular media are exponentials of stated peak and width rather than fitted tables, which is enough to carry the argument and is honest about being a model. The lens absorbance in particular is a single exponential where the literature reports a two-component form with an age-dependent split.

And the fit is unweighted least squares over the whole grid, which spends its accuracy where the functions are large. A fit weighted towards the tails would have a smaller residual and a worse ȳ one, and no version of this collection’s results has been recomputed under one.

The conditions under which an observer's departure is exactly zero. A departure of the observer is the pairing of something belonging to the observer with something belonging to the stimulus, so emptying either factor empties the product. The axis is logarithmic in what is left when the condition is imposed. Six rows empty the stimulus's factor — a perfectly neutral sample is the same colour for every observer, at any age and any field size — and two empty the observer's, since a gain on each cone and a change of basis are both absorbed exactly. All eight are identities rather than small numbers. The last two are the same two conditions imposed in a published cone space rather than in the observer's own, and they are worth eight and thirteen units: the identity is about the eye, and the arithmetic everybody uses is in somebody else's coordinates.
Fig. 6 Ten conditions under which a departure of the observer vanishes. Eight of them are identities at the floating-point floor, and an identity is the one kind of result a modelled reference cannot get wrong.

That figure is the strongest defence of the whole arrangement and it is worth making explicitly. An identity is a structural claim: it says a quantity is exactly zero under a stated condition, and it is true of any observer of this shape rather than of the particular curves. The model’s residual cannot contaminate it, because the residual would have to be exactly cancelled for the identity to hold spuriously, and it is not.

So the results of this round divide into three kinds with three different exposures to the model. The identities are immune. The comparisons are exposed to second order, which is small. The absolute figures — what the ultraviolet end of the range costs, how far apart two observers sit — carry the residual in full, and they are the ones stated with their construction named in the caption.

The generalisation

The habit is about what to do when a measurement has no reference.

The instinct is to construct one from the same data by a more careful procedure, and that is almost always circular: the more careful procedure inherits the data’s limitation and hides it behind extra arithmetic. The alternative is to change the kind of object being compared against — from a measurement to a model — and to pay for it by publishing the model’s residual in front of the results.

The trade is worth making when the questions are comparative and not when they are absolute, and the way to tell is to ask whether the reference appears on both sides of every subtraction. When it does, its error cancels to first order and a rough model is enough. When it does not, the model’s residual is the answer’s error bar and a model is the wrong tool.

The failure mode is to build the model, produce the numbers, and then quote them as though they had been measured. A fit can be exact and empty, and a model can be adequate for one class of question and useless for the next one asked of it, with nothing in the output to mark the transition.

Who found it, and when

Govardovskii and colleagues published their template in 2000, from microspectrophotometry across many species, and it superseded Dartnall’s nomogram of 1953 for most purposes. Neither was built for this use.

The general problem — that an instrument cannot calibrate itself against its own output — is old enough to be a proverb in metrology, where it is the reason a hierarchy of standards exists at all. The spectral version has an unusual feature: the hierarchy stops. There is no finer tabulation of the standard observer, because the standard observer is defined by its table, and a finer one would be a different observer rather than a better measurement of the same one.

Where the ladder goes next

The grid has now been taken apart into three decisions, each measured, and the section closes by saying what a resolution is and is not.

After that this round leaves the index and opens the third factor of the integral. The observer is not a measurement either — it is a construction with arguments of its own — and the arguments a standard observer does not admit to having are what the rest of the round is about.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AssertionAuditColour-matching functionsConvergenceFalsificationModelling assumptionPigment templateResidualStandard observerWavelength grid