Where the model breaks

A fit can be exact and empty

Every fitted object here reports one number, the residual on the data it was fitted to, and every one of them has two more that nobody publishes — how many of its parameters the data actually determine, and how large the family of equally good answers is. The third column is where the failures live.

Assumes The matches do not name the cones, The chart decides the profile and An image does not determine the light.

Every model in this collection that has ever been fitted arrives with a residual. The camera matrix comes with a mean ΔE00, the adaptation basis with what it leaves after the gain, the profile table with an error between its nodes, the appearance model with an agreement to four significant figures against a published vector.

A residual answers one question, and it is not usually the question anybody has.

What each fitted thing in these essays carries, what its data fix, and what is left. Three columns per row: how many numbers the model has, how many the stated data determine, and the difference — the dimension of the family that fits equally well. The third column is the one nobody publishes. A zero there does not mean the model is right; it means it is determined, which is a much weaker property and is compatible with being determined badly, as the camera row is.
Fig. 1 Six fitted objects from across the site. The first column is how many numbers the model carries, the second is how many the stated data determine, and the third is the difference — the dimension of the family that fits equally well.

The claim

A fit has three quantities and reports one. The residual says how closely the model follows the data it was given. The rank says how much of the parameter space those data actually visited. The orbit says how large the set of equally good answers is — and a residual is silent about both of the others by construction.

  • A model can be exact and empty. Colour matching determines the observer’s three curves to any precision anybody likes and leaves nine numbers entirely free.
  • A model can be determined and badly determined. A camera matrix on a near-neutral chart has all nine parameters pinned, at a conditioning of 178 to 1, and reports 0.19 ΔE00 while delivering 1.65.
  • A model can be underdetermined by a fixed amount however much data arrive. A scene of sixty-four surfaces leaves the same deficit as a scene of two.
  • And a residual cannot detect any of that. It is computed in the directions the data determined, which is exactly where the model is guaranteed to agree.
  • All three quantities are cheap. The rank is one eigendecomposition; the orbit is a perturbation and a re-evaluation.

The three columns

Free is how many numbers the model carries. Nine for a 3×3, three for a linear illuminant model, three thousand six hundred and forty-five for a nine-step profile table with five black levels, eighty-one for a spectrum on this site’s grid.

Fixed is how many of them the data determine. It is a rank, not a count of observations, and the difference between the two is where most of the trouble is: twelve patches supply thirty-six numbers and determine nine parameters at a conditioning that ranges over an order of magnitude depending on which twelve.

Left is the difference, and it is the dimension of the set of parameter values that reproduce the data equally well. Zero means the model is determined. It does not mean the model is right; it means there is exactly one answer, and whether that answer is good is a different question entirely — the camera row has a zero in it and is the worst-behaved row in the table.

Nine inputs, nine numbers, and nothing left over. The singular values of the Jacobian of the nine matrix entries with respect to the nine things that fix them: two coordinates for each of the three dichromat confusion points, and one scale for each row. All nine are well clear of zero, so the map is a bijection — six plus three is nine exactly, and neither a redundancy nor a shortfall is hiding in it. A collapsed bar here would mean one of the inputs is doing nothing, or two are doing the same thing.
Fig. 2 A rank being measured rather than counted. Nine inputs, nine outputs, and a Jacobian whose smallest singular value is well clear of zero — which is what “determined” looks like when it is checked instead of assumed.

Four ways the third column goes wrong

The rows of the table fail in four distinguishable ways and it is worth separating them, because the repairs are different.

A symmetry the data cannot break. The observer’s nine numbers are the pure case: every colour match is invariant under the whole group, so no quantity of matching data narrows it by a single dimension. The repair is a different experimentthree confusion points close it exactly — and no amount of the original measurement helps at all.

A structural deficit that more data cannot close. An image does not determine its illuminant because each further surface supplies three equations and three unknowns. The repair is a fourth sensor, or a smaller model, or a measurement of a different kind; a larger scene is not on the list.

A design that visits few directions. A chart with no chromatic range determines all nine matrix entries and determines several of them from almost nothing. The repair is a better design, and the diagnostic — the condition number — costs nothing and is computed from the design alone, before any measurement is made.

And a model with more parameters than anybody could constrain. A profile table has thousands of entries and is exact at every one of them, so its residual at the nodes is zero and means nothing; everything it is used for is between them.

How short of determining the light a photograph is, as the scene grows. Each cell is the number of unknowns left over after every equation the image supplies: three sensors, a three-dimensional illuminant, and reflectances confined to a linear model of the dimension on the left. At one and two dimensions more surfaces close the gap. At three the gap never closes, because each further surface adds three equations and three unknowns; at four it widens. The count is arithmetic and has no algorithm in it.
Fig. 3 The second kind, as arithmetic. At two dimensions more surfaces close the gap and at three they never do, and which case a problem is in can be decided before any data are collected.
What a camera matrix reports on its own chart, and what it delivers off it. The same camera fitted on charts of increasing chromatic range. The left bar of each pair is the mean error on the chart the matrix was fitted to, which is the number a profile comes with; the right bar is the error on a saturated set it never saw. At the thinnest chart the fit reports 0.19 ΔE00 and delivers 1.65, a factor of 8.6. The gap closes as the chart widens, and it closes because the chart improves rather than because the camera does.
Fig. 4 The third kind. Every bar pair is the same camera; the left bar is what the fit reports and the right is what it delivers, and the gap is a property of the patches.

The two words that get conflated

There is a vocabulary problem underneath all four failures and naming it helps.

Structural identifiability asks whether the parameters would be determined by perfect, noiseless, infinite data. It is a property of the model and the experiment’s design, and it can be settled with algebra before anything is measured. The observer’s nine numbers and the scene’s deficit are both structural: no data of the stated kind would help.

Practical identifiability asks whether the parameters are determined well enough by the data actually available, at the noise actually present. It is a property of the numbers and is settled by a condition number. The near-neutral chart is practically unidentifiable and structurally fine.

The two require different repairs and are routinely described by the same word. A structural failure needs a different experiment; a practical one needs a better version of the same experiment. Trying to fix a structural failure with more data is the commonest wasted effort in this whole family, and it is what a great deal of work on colour constancy amounts to.

Why the residual cannot see it

The reason is one sentence and it is worth stating carefully because it explains all four failures at once.

Least squares chooses the parameters that minimise the residual, so the residual is by construction flat in the directions the data leave free. A parameter combination the data do not constrain is precisely one that does not move the residual; if it did, it would be constrained.

So the residual is not merely uninformative about the orbit — it is guaranteed to be uninformative, in the strongest possible sense. Reporting it and stopping is not an oversight of degree.

The same argument explains why a held-out set is a partial repair rather than a complete one. A held-out residual does detect overfitting, because the held-out data were not used to choose the parameters. It detects an unconstrained direction only if the held-out data happen to vary along it, and if the held-out set was drawn from the same population as the fit set — which is the usual practice and the recommended one — it will tend not to.

What was computed, and how

Each row of the ledger is computed where a computation is available and stated with its source where it is not.

The observer row is the dimension of the general linear group in three dimensions, and its zero in the second column is a theorem rather than a measurement — a match is an equality and a matrix applied to both sides of an equality is applied to nothing.

The camera row’s numbers are the condition number of the patch responses, the fit residual and the held-out residual, all from the same fit.

The scene row is the Maloney–Wandell count at sixty-four surfaces: one hundred and ninety-four unknowns against one hundred and ninety-two equations.

The lamp row’s percentage is the share of a fluorescent tube lying in the null space of three broad filters.

And the profile row’s free is infinite, which is not a rhetorical flourish: the object being fitted is a function on a continuum and the table is a finite sample of it, so the family of functions agreeing at every node is infinite-dimensional. Writing a large finite number there would be a category error about what a table is.

How many numbers a surface takes, on the friendliest set available. The cumulative share of the variance accounted for by the first few principal components of this collection's 240 constructed reflectances. Three components reach 99.94 per cent and two reach 97.99. The counting result allows two, and these are deliberately smooth curves built from a handful of Gaussians — the friendliest possible case, and it already needs 3.
Fig. 5 The measurement behind the scene row. How many numbers it takes to write a reflectance down is not a modelling choice, and the counting result needs it to be two.

Where the model stops

The ledger is curated rather than exhaustive. Six rows chosen because they are the fitted objects this collection argues about most; a complete audit of every fitted quantity on the site would be a larger and less legible thing.

Rank is a threshold decision. Which singular values count as non-zero depends on a tolerance, and the honest tolerance is the measurement’s noise floor rather than machine precision. A chart’s usable rank at a realistic noise floor is smaller than its patch count, and the first draft of that computation used machine precision and reported a rank that no instrument could deliver.

Nor is the ledger’s second column always a rank. For the scene row it is a count of equations, and equations and rank coincide only when the equations are independent — which they are here, generically, and which they would not be in a scene of identical surfaces. Writing the count is the conventional form of the Maloney–Wandell result and it hides that caveat, which is why the construction of the family is a better argument than the count.

And a large orbit is not automatically a problem. If every member of the family makes the same predictions about everything anybody will ask, the model is fine and the parameters were simply over-specified. The orbit matters when its members disagree about something — which is why the useful form of the third column is not a dimension but a consequence, measured in the units the model’s answers are used in.

The scene row hides a subtraction

The scene row’s arithmetic is quoted twice in this essay and the two versions do not agree, which is worth resolving because the disagreement is exactly one unit and the unit has a name.

The prose says each further surface supplies three equations and three unknowns, and a three-parameter illuminant sits on top of that. Counted that way, sixty-four surfaces give 195 unknowns against 192 equations, and the deficit is 3 — the illuminant’s own dimension, which is the intuitive answer.

The ledger says 194, and a deficit of 2. The missing unknown is the overall scale: a scene lit twice as brightly by a lamp half as bright is the same set of measurements, so one of the 195 numbers is fixed by convention before any counting starts. It is 3N + 2 at every N, so the deficit is two rather than three whatever the scene contains, and the count in the ledger has already done the subtraction.

That single unit is not bookkeeping. It is the reason brightness is unrecoverable from an image, stated as a dimension rather than as a fact about photography, and it is the one degeneracy in this collection that everybody already accepts and nobody counts. The other one is the honest deficit and is what the constancy literature is about.

What the camera row’s two numbers are worth against its third

The camera row carries a conditioning of 178, a reported residual of 0.19 ΔE₀₀ and a delivered one of 1.65 — and the essay uses the first as the explanation of the gap between the second and the third without checking whether it is large enough to be one.

It is. The gap is a factor of 8.7, which sits between 1 and 178, so nothing is being violated. The sharper reading is against the square root, 13.3: the delivered error is 65 per cent of the worst amplification a design conditioned that way permits when the conditioning is quoted on the Gram matrix, which is the usual convention for a least-squares design and the one that makes 178 a plausible figure for twelve near-neutral patches.

Sixty-five per cent of a worst case is a bound doing work rather than a bound quoted for decoration. A conditioning of 178 with a delivered error at five per cent of it would say the diagnostic was pessimistic and the chart got lucky; at 65 per cent of the square-root bound it says the chart is close to as bad as its geometry allows. That is the difference between a number that warns and a number that merely does not contradict, and it is available for the price of one square root.

Three of the six rows report the column the essay asks for

The closing recommendation is that the third column should be not a dimension but a consequence, measured in the units the model’s answers are used in. Read against the ledger it audits, the recommendation is not yet met on half its own rows.

Three rows carry a consequence. The camera row’s is 1.65 ΔE₀₀. The sensor row’s is about one per cent on a narrow light, from an alternative sensor exactly indistinguishable on every patch. The lamp row’s is a share of a fluorescent tube lying in a seventy-eight-dimensional null space — 81 bands less three filters — which is a consequence in the units of the thing being measured.

Three rows carry only a dimension. The observer row says nine. The scene row says two. The profile row says infinite, and infinite is the least actionable entry a column can hold: a table exact at 3,645 nodes says nothing at all about the continuum between them, and the number 3,645 is a property of the grid rather than of the object.

That is not a failure of the ledger so much as a statement of what the three rows are. A symmetry the data cannot break has no single consequence, because its consequence depends on what the fitted object is subsequently used for — the observer’s nine numbers cost nothing to a match and a great deal to a confusion line. A dimension is what is left when the consequence depends on the question, which means the recommendation’s not a dimension but a consequence holds where a use is fixed and fails where the fit is a general-purpose object.

The useful form of the rule is therefore narrower than the one stated: report the consequence when the fit has a known consumer, and the dimension when it has many. Half of this ledger is in the second case, which is why half of it reads as abstract.

Who found it, and when

None of this is new anywhere. Identifiability is a standard concept in statistics and in system identification, the singular-value diagnosis of a design matrix is in every numerical-analysis textbook from the 1960s, and the distinction between structural and practical identifiability has a literature of its own in systems biology and pharmacokinetics.

What is unusual is how little of it reaches colour science’s reporting conventions. A camera profile ships a residual. A characterisation ships a residual. A published set of cone fundamentals ships a derivation and no statement of what the derivation’s inputs left free. The vocabulary exists three departments away and the practice has not moved.

The likely reason is that most of these fits are determined — the third column is zero for the camera matrix, for the appearance model, for the adaptation basis — and a zero there is easy to read as a clean bill of health. It is not. Determined and well-determined are different properties, and only the second one makes a residual mean something.

The generalisation

The rule is that a fit should be reported with three numbers, and the second and third are cheap.

The residual on the data. Already reported everywhere.

The conditioning of the design, which is one small eigendecomposition and can be computed before any measurement is taken, from the experimental plan alone. It is the number that says whether to believe the residual.

And the consequence of the orbit — how far the answer can move without the residual changing, expressed in the units the answer is used in. That is one perturbation and one re-evaluation, and it is the number that turns an abstract degeneracy into something a practitioner can act on.

The habit that makes all three worth having is a question rather than a computation: what would have to be different about the world for this fit to come out differently? If the honest answer is nothing that was measured, the fit is reporting a convention.

There is a corollary worth keeping for the cases where the answer is uncomfortable. A model whose orbit is large is not thereby wrong and is not thereby useless; it is over-specified, and the honest response is usually to say so rather than to reduce it. The cone fundamentals are a good example: the nine free numbers are real, the standard choice is defensible, and what the literature owes its readers is one sentence saying which construction produced it — not a smaller model.

One row where the answer is reassuring

It would leave the wrong impression to stop without a case in which the machinery reports that everything is fine, since the point of a diagnostic is that it can come out either way.

The sensitivities of a camera, probed by a chart of reflectances that are bumps at spread wavelengths, are determined to within 3.5 per cent of each curve. An alternative sensor built from the residue is exactly indistinguishable on every patch and differs by about one per cent on a narrow light — a real ambiguity, and a small one.

The reason is the design rule the negative results imply, run forwards. Patches that differ in where in the spectrum they are dark are a crude filter bank, and a filter bank is a good probe of a curve. The near-neutral chart fails for the mirror-image reason: its patches differ in level rather than in position.

So the same site, the same instrument and the same arithmetic give a bad answer to one question and a good answer to another, and the difference is entirely in what the patches were asked. That is what makes the third column worth computing rather than assuming.

The part of a sensor a chart reaches, and the part it does not. Each channel's sensitivity, and the residue left after projecting it onto the span of the 12 illuminated patch spectra — the part no measurement of that chart constrains. A second camera differing by that residue produces identical raw values on every patch, to 2.9e-6 relative, and differs by 1.1% on a narrow light at 546 nm. The residue is 3.5% of each curve, which is small — patches that differ in where they are dark are a good probe, and that is the design rule.
Fig. 6 The reassuring row. Each channel’s sensitivity and the residue no measurement of the chart constrains, at six times scale.

What the columns are worth to a reader

The three-column form is a reporting convention and conventions only travel if they are cheap for the person adopting them. All three columns are cheap, and two of them are cheaper than the one already in use.

The residual costs a fit and a data set, both of which exist.

The rank costs one small eigendecomposition of the design’s Gram matrix, which for a nine-parameter model is a 3 × 3 or a 9 × 9 and takes microseconds. It requires no extra data at all, and — this is the part worth insisting on — it can be computed before any measurement is taken, from the experimental plan alone. A chart that will condition badly can be identified while it is still a list of patches.

The orbit costs one perturbation and one re-evaluation. Move the fitted parameters along the design’s weakest direction until the residual rises by the measurement’s own repeatability, then evaluate the moved model on whatever the model is for. The answer is in the units the model’s answers are used in, which is what makes it actionable where a dimension is not.

None of that is a research programme, and the reason it is absent everywhere is not cost.

Where the ladder goes next

Applying that question to this collection’s own claims sorts them into two piles, and four of nine ordinary sentences on this site turn out to be about the paper rather than the eye.

Which of this collection's own claims survive a change of basis, and which are about the paper. Nine sentences this site says, sorted by whether they mean the same thing after the observer's three curves are replaced by a nonsingular combination of themselves. 5 of the nine do. The four that do not are not thereby wrong — they are statements about a chosen set of coordinates, and they are true of those coordinates. What they cannot be is statements about the eye, which is how every one of them is usually read.
Fig. 7 Which is the audit, and the point at which the machinery in this essay is turned on the site that built it.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 23 that link here.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationDegrees of freedomHeld-out validationIdentifiabilityLeast-squaresLinear modelMeasurement uncertaintyModel selectionNull spaceOverfittingPrincipal componentsRank