What a scene does

A model is a claim about what can be known

The exact answer to chromatic adaptation is nine numbers, and the nine numbers are the change of light itself. A model whose parameters are quantities the observer cannot obtain is not a worse model of the same thing — it is a model of something else, and counting parameters without asking where they come from hides the difference.

Assumes Three numbers the scene supplies, A partial correction is worth its fraction and An image does not determine the light.

Two models with the same number of parameters can be entirely different kinds of claim, and the difference is not in the count.

A correction an observer could have been born with, fitted on half the census and tested on the other. The same six models, each scored twice: on the seven census rows the fixed matrices were fitted to, and on the seven they were not. The split alternates by position so both halves contain daylight changes and discharge lamps. The upper bar is in sample and the lower is out, on a logarithmic axis. For the four models with nothing fitted the two bars differ only because the halves are different questions. For the two fitted ones the gap is the finding, and it is largest where it matters least: bolting a fixed correction onto the von Kries gain takes it from 1.2724 to 1.2592 on the rows it was fitted to, and from 1.3511 to 1.3679 — worse — on the rows it was not. There is no correction to the diagonal that an observer could arrive with.
Fig. 1 Six models scored twice: on the census rows the fitted matrices were fitted to, and on the ones they were not. Two of the six have parameters that were fitted; four do not. The gap between the bars is what fitting bought, and for the eighteen-parameter model it is negative.

The claim

A parameter is a claim that somebody can find its value out, and an adaptation model’s parameters divide sharply into ones that can be and ones that cannot.

  • The exact matrix’s nine numbers are the change of light, so obtaining them requires solving the problem the model exists to solve.
  • A von Kries gain’s three are the white, which is an underdetermined estimate rather than a measurement — but is one a visual system demonstrably makes.
  • A basis’s nine are fixed, so they cost nothing at run time and are a different kind of parameter entirely.
  • Counting the three kinds together makes the diagonal look crude. Counting them separately makes it the only entry on the ladder that gets a large answer from obtainable information.
  • And the distinction is testable, because a parameter that cannot be obtained per scene has to be fitted across scenes — where it fails to generalise, measurably.

Three kinds of number in one model

The published adaptation model has, on the usual count, twelve parameters: a 3×3 basis and three gains. Written out, they are three different things.

Nine fixed numbers. The basis — the cone-like axes the gain is taken in. They are the same in every room, they were settled once, and what settling them well is worth is a question with its own round of work. An observer does not estimate them; a visual system either has them or does not.

Three scene numbers. The white, in that basis. These have to come from the room the observer is in, they change whenever the light does, and getting them is an underdetermined problem with no solution that is not a prior.

And zero free numbers. Nothing in the model is fitted to the census. The gains are computed from the whites by a rule stated in 1902.

The exact matrix has nine numbers and all nine are the third kind gone wrong: they are neither fixed nor estimable, because they are the change of light. An observer that could obtain them would not be doing adaptation; it would be doing measurement.

Why the distinction is not a technicality

There is a tempting way to dismiss all this: every model’s parameters have to come from somewhere, the difference is one of degree, and the residual is the residual.

The reply is that the degrees are not on a continuum. Consider three questions a parameter can be asked.

Can it be settled once? If yes, its cost is a one-off and its risk is that the setting is wrong for some rooms. A basis is like that.

Can it be estimated from the scene, imperfectly? If yes, its cost is a mechanism plus an error, and the error can be measured. A white is like that: this collection has measured what several estimation strategies cost, and the answers are between one and several units.

Or is it the answer? If a parameter’s value can only be known by knowing what the model is for computing, the model is not a model. The exact matrix is like that, and no amount of engineering changes it.

The third case is qualitatively different from the first two, and a table that adds 9 + 3 and reports 12 has erased it.

What happens to three familiar comparisons

The distinction changes the reading of three things this collection has already published.

The census’s residuals stop being a shortfall. Fourteen numbers between 0.26 and 3.37, described for three rounds as what a diagonal gain leaves behind. Read with the count separated, they are what is left when an observer applies everything it can know, and the shortfall framing invites a repair that does not exist.

The basis search stops being a search for a better model. Fitting a basis to the census improves the residual, and it improves it by adjusting nine fixed numbers — so the improvement is available to a visual system in principle, over evolutionary time, and is a genuinely different proposition from an improvement that would need a new estimate per room. That this collection’s held-out basis fit beats the published transforms by a little and not a lot is then a statement about how well-chosen the published ones already are.

And an eighteen-parameter model becomes an obvious mistake rather than a plausible one. Adding a fixed 3×3 correction after the gain adds nine fixed numbers and no scene numbers. It should therefore be free, and it should help if there is anything systematic left. It does not: in sample 1.2592 against 1.2724, out of sample 1.3679 against 1.3511 — worse. There is nothing systematic left, which is a stronger statement than the residual alone could make.

What an observer is left with, by how much it is allowed to know about the room. Six ways of discounting a change of light, averaged over the fourteen changes in the adaptation census and 125 test surfaces each. The bar is what each leaves behind, on a logarithmic axis because the models span two orders of magnitude. The second line under each name is the count that matters: how many numbers about this room the model has to be given. Doing nothing leaves 15.7 ΔE₀₀. A single gain read off the two whites' luminances leaves 15.3. A matrix fitted across half the census and then applied everywhere, knowing nothing about the room at all, leaves 12.5. The published von Kries gain, which is told the white and nothing else, leaves 1.312 — and bolting a fixed correction onto it, at no cost in scene information, leaves 1.368, which is very slightly worse. The exact matrix leaves nothing and is not on the chart: its nine numbers are the change of light, which is the quantity being discounted.
Fig. 2 The ladder with the counts on it. Reading down the second line under each name is reading the argument: nothing under three scene numbers gets close, and the model above three requires a quantity nobody has.

The negative result, assembled

Three findings, from three directions, all pointing the same way. Together they are as close to a complete account of the residual as this collection is likely to reach.

Nothing fixed removes any of it. A 3×3 fitted across half the census and applied to the other half makes things slightly worse. If the residual contained a component common to different changes of light, one matrix would capture it.

Nothing partial removes more than its share. The line from the diagonal to the exact answer is straight in every unit that is a norm, so there is no cheap approximation to be had.

And what remains is per-scene by construction. The exact correction is a different matrix for every row, and its nine numbers are the change of light.

So the residual is neither free nor cheap nor generic. It is information about the particular room, and the observer’s inability to have it is not a limitation of the model but the definition of the problem.

That is a satisfying place to arrive and it is worth flagging as the point at which to be suspicious. A tidy negative result assembled from three measurements on one constructed family of surfaces is exactly the shape of a conclusion that is true of the construction. The family is three-dimensional by fiat, and a fourth dimension at realistic amplitude puts 0.33 ΔE2000 into the theorem the whole argument rests on — which is the size of the mildest census row.

How much of the residual a partial correction removes. Between the diagonal gain and the exact matrix there is a line: apply the correction that would make a row exact, but only a fraction of it. The horizontal axis is that fraction and the vertical is the share of the row's residual it removes, for all fourteen census rows. The straight diagonal is where a correction worth exactly its fraction would fall, and in the published unit every curve lies on it to within 2.2 percentage points. The lower band of curves is the same interpolation measured in CAM16-UCS, which departs by up to 17 points — because its distance is a power of the Euclidean one and a power is not homogeneous along a ray, where every ordinary norm is. The straight line is therefore a property of the ruler rather than of the correction, and the exception is what says so.
Fig. 3 The second of the three: a partial correction is worth its fraction, in the published unit, on every row. The lower band is the same interpolation in the appearance unit, whose distance is a power law and which is the exception that establishes the rule.

The same count applied elsewhere in this collection

The distinction is not about adaptation, and pointing it at three other models here shows what it does and does not settle.

A camera profile. Nine numbers, all fixed, all fitted once at calibration time. The camera does not estimate them per scene and could not — so a profile’s error is entirely the price of having settled them in advance, and changing the objective it was settled under moves the matrix without changing what the camera knows. That is a pure fixed-parameter model, and it is why a profile fitted on one chart and used on another fails in a way that is predictable rather than random.

A colour appearance model. Dozens of constants, all fixed, plus four scene numbers: the adapting white, the adapting luminance, the background and the surround. Three of those four are things an observer could plausibly estimate; the fourth, the surround, is a category with three tabulated values standing in for a continuum, and nobody has said how an observer would know which of the three it is in. The surround is the appearance model’s weakest parameter on this criterion, and it is not usually named as weak.

And an illuminant estimator. Zero fixed numbers and a prior — which is a fixed number in disguise, and a large one. The grey-world assumption is one number’s worth of prior and the gamut-mapping methods are many, and their failures are failures of the prior on scenes that violate it rather than of arithmetic.

Three models, three different distributions of the same twelve-to-forty parameters across the three kinds, and in each case the model’s characteristic failure is a failure of whichever kind dominates it.

What the count does not say

Two things it would be easy to read into this and should not be.

It does not say the visual system computes a diagonal gain. The mechanisms are receptor gain control, post-receptoral normalisation and processes further up, and the diagonal is a description of the net effect at the level colorimetry works at. A physiological account with more parameters is not refuted by an information argument about a phenomenological model.

And it does not say more parameters are always suspect. A model with many fixed parameters is entirely respectable — a colour appearance model has dozens and most of them are fixed constants fitted once to psychophysical data. The question is never how many; it is where each one comes from, and whether the answer is available to whatever is supposed to be running the model.

Every test surface under each model — a triphosphor tube. The 125 surfaces of the test set, sorted for each model from the one it costs least to the one it costs most. Doing nothing is a nearly flat line high on the chart: a change of light of this kind moves almost every surface by about the same amount, which is why it looks like a change of light rather than like a change of scene. A luminance gain barely moves that line. The von Kries gain drops it to the floor and changes its shape: 5 of the 125 surfaces cost exactly nothing, because they are flat greys on which an adapted observer's gain is exactly right, and what is left rises steeply over the most saturated members. The published residual for this row is a mean over that curve.
Fig. 4 The row the models struggle with most, surface by surface. A triphosphor tube’s emission lines are structure narrower than a cone, and no number the observer could read off the room would help — which is what a per-scene residual looks like when it is drawn rather than summarised.

Two more rows say whether the ordering of the six models is a property of the models or of the row they were read on.

Every test surface under each model — a green wall, two bounces. The 125 surfaces of the test set, sorted for each model from the one it costs least to the one it costs most. Doing nothing is a nearly flat line high on the chart: a change of light of this kind moves almost every surface by about the same amount, which is why it looks like a change of light rather than like a change of scene. A luminance gain barely moves that line. The von Kries gain drops it to the floor and changes its shape: 5 of the 125 surfaces cost exactly nothing, because they are flat greys on which an adapted observer's gain is exactly right, and what is left rises steeply over the most saturated members. The published residual for this row is a mean over that curve.
Fig. 5 The harshest row in the census, surface by surface, under each model. The ordering of the six is the same as it was on the triphosphor row, which is what makes the ladder a ranking rather than an anecdote.
Every test surface under each model — the macular pigment. The 125 surfaces of the test set, sorted for each model from the one it costs least to the one it costs most. Doing nothing is a nearly flat line high on the chart: a change of light of this kind moves almost every surface by about the same amount, which is why it looks like a change of light rather than like a change of scene. A luminance gain barely moves that line. The von Kries gain drops it to the floor and changes its shape: 5 of the 125 surfaces cost exactly nothing, because they are flat greys on which an adapted observer's gain is exactly right, and what is left rises steeply over the most saturated members. The published residual for this row is a mean over that curve.
Fig. 6 And the mildest, where every model has less to remove. The gaps have closed and none of the six has changed places, so what the ladder measures survives the row it is read on.

The row above is the criterion made concrete. A triphosphor tube is the change of light the diagonal gain handles worst, at 67.3 per cent removed against the census’s 91.6, and the reason is not that three numbers are too few. It is that the information the residual would need is in the lamp’s line spectrum, which three cone channels have already integrated away before any model gets to run. No count of parameters helps with a quantity the instrument cannot represent.

Where the model stops

The argument is about a phenomenological model of adaptation and a census of constructed changes of light. It is not about vision.

The three scene numbers are handed to every model here exactly, which no observer gets. An error in the estimated white propagates, and the propagated error is comparable in size to the residual this essay is about — so a full account would have the diagonal gain’s real performance somewhere below its ideal 91.6 per cent, and the ladder’s gaps correspondingly compressed.

And “obtainable” is doing philosophical work that has not been earned. It is used here to mean obtainable from the light arriving, by something with the structure of a visual system, which is a claim about mechanisms this collection has no access to. A different observer — a spectroradiometer on a tripod — obtains all nine numbers easily, and for it the exact matrix is a perfectly good model. The count is relative to the knower, which is the whole point and is also the limit of it.

How much of each change of light three numbers about the room remove. The ladder's headline figure, per row rather than averaged. The bar is the share of the unadapted difference that a von Kries gain removes — the gain being three numbers, the white, read off the room. It runs from 67.3 per cent on a triphosphor tube to 97.6 per cent on the macular pigment. The two ends are the two kinds of change: a filter inside the observer or a wall in the room changes the light in a way a gain handles almost perfectly, and a discharge lamp with emission lines in it does not, because a gain cannot see structure narrower than a cone.
Fig. 7 The white’s three numbers, row by row. From two thirds to almost all of each change of light, from a quantity that has to be guessed and can be guessed reasonably well.

What a held-out test measures when the count is separated

The held-out discipline in the figure at the top of this essay is doing something more specific than it looks, and it is worth naming because it is the criterion’s one empirical handle.

A fixed parameter is a claim that one value works everywhere. That claim is testable: fit the parameter on some rooms and score it on others. If it generalises, the claim is good; if it does not, the parameter was never fixed — it was a scene number being estimated badly by averaging.

That is exactly what the eighteen-parameter model fails. Its nine extra numbers are declared fixed, are fitted on seven changes of light, and make the other seven worse. So the declaration was wrong: whatever those nine numbers were capturing is a property of particular changes rather than of changes in general.

And it is what the basis passes. A basis fitted on half the census and scored on the other half beats the published transforms on the half it never saw. Nine numbers, declared fixed, and the declaration survives the test.

The two results sit in the same table and are usually read as one thing — fitting helps a bit — when they are opposite. One fixed-parameter claim is vindicated and one is refuted, by the same instrument, in the same run. This collection’s habit of holding out exists for the second case and turns out to be the way to check the first.

What the four numbers in the held-out test are each measuring

The eighteen-parameter result is carried by four numbers — 1.2592 and 1.2724 in sample, 1.3679 and 1.3511 out of it — and reading them as two comparisons rather than four readings makes the evidence look stronger than it is. Three of the four differences they contain are worth naming separately, because only one of them is the finding.

The in-sample gain is 1.04 per cent. Nine free numbers, fitted on the rows they are then scored on, improve the mean residual from 1.2724 to 1.2592. That is a small return for nine parameters, and it is the first piece of evidence for the essay’s conclusion rather than a preliminary to it: a fixed 3×3 applied after the gain can be fitted to the residual and still recover only a hundredth of it, which says the residual is very nearly orthogonal to everything a fixed matrix can express.

The out-of-sample loss is 1.24 per cent, 1.3679 against 1.3511. This is the finding, and it is a valid comparison: both figures are on the same held-out half, so whatever makes that half harder affects them equally and cancels.

And the generalisation gap tells the two models apart more clearly than either. The eighteen-parameter model is 8.63 per cent worse out of sample than in it; the twelve-parameter model is 6.19 per cent worse. The nine extra numbers add 2.45 points of gap — twice the size of the effect the essay reports and in the same direction, which is the textbook signature of parameters that are memorising rather than describing.

The yardstick the four numbers arrive with

The last of those readings is the one worth dwelling on, because the twelve-parameter model has no fitted parameters at all. Nothing in it was tuned on either half. So its 6.19 per cent is not a generalisation gap in any ordinary sense: it is a measurement of how much harder the held-out half of the census is than the fitted half, and it arrives free with the experiment.

That gives a scale for the other two numbers, and the scale is not flattering. The in-sample gain is 0.17 of the half-to-half difference. The out-of-sample loss is 0.20 of it. Both effects the essay rests on are a fifth of the variation between two arbitrary halves of a fourteen-row census, measured on one split.

This does not undo the conclusion. The comparison that carries it is within-half and the sign is unambiguous on that half. What it bounds is how much a single split can establish: with an effect a fifth the size of the between-half variation and one draw, a different partition of the fourteen rows could plausibly report a smaller loss, and could conceivably report none. The honest form of the claim is that the nine extra numbers do not help and appear to hurt, with the appearance resting on one arrangement of a small census.

The repair is cheap and worth naming since the essay’s own argument is about what a held-out test can establish. Fourteen rows admit 3,432 balanced splits; running all of them, or a few hundred, would turn one reading into a distribution and say whether the loss is reliable or is one draw’s worth of which rows landed where. That costs a few seconds of the machinery that already exists, and it is the difference between the fixed correction fails to generalise and the fixed correction failed to generalise on this split.

The basis result is on firmer ground for a reason that has nothing to do with sample size. It beats the published transforms on the half it never saw — a comparison against an external benchmark rather than against a sibling fit — so it does not depend on the two halves being comparable at all. The two results in that table really are opposite, as the essay says, and they are also unequally supported.

Who found it, and when

The observation that colour constancy is ill-posed is the foundation of computational colour constancy and dates to the 1980s; Land’s retinex and everything after it is an attempt to supply the missing prior. That literature is about the estimate — how to guess the illuminant — and takes the model that uses the estimate for granted.

The observation that a model’s parameters are claims about available information is not from colour science at all. It is closest to the distinction in control theory between the state a controller can observe and the state it cannot, and to the older statistical distinction between parameters and latent variables. Neither is usually applied to a perceptual model, because a perceptual model is normally judged by how well it predicts what people report rather than by whether its inputs are things people have.

Applying it here costs one extra column in a table. What it buys is that three familiar quantities stop being a shortfall and become a bound, which is a change in what the collection is entitled to say next.

Where the ladder goes next

Two of the three structural choices are now reached. The third is under every observer on this site rather than under any one calculation: the population’s two hundred members are built from a pigment template at stated peaks, and the template is a formula somebody fitted to measurements of individual receptors in 2000. What happens when it is replaced is a question about the whole population at once.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

BasisChromatic adaptationConstancyHeld-outIlluminant estimationInverse problemMatrixModel complexityResidualThe von Kries transform