What a scene does

An image does not determine the light

A photograph of a scene under one illuminant gives three numbers per surface and asks for the illuminant plus three numbers per surface. The count closes only if reflectances lie in a two-dimensional model, and no number of surfaces helps — at three dimensions the alternative scenes can be written down, and they reproduce every sensor response exactly.

Assumes A scene has no white point, The algorithms that guess the light and A bounce is a multiplication.

Every estimator this site has guesses the illuminant of a scene and every one of them is wrong on some scenes. Grey-world reaches 43 degrees of angular error on an all-green scene; grey-edge and max-RGB fail elsewhere; no estimator wins on every scene, so a camera choosing one has chosen which photographs to be wrong about.

That reads like an engineering problem awaiting a better algorithm. It is not.

How short of determining the light a photograph is, as the scene grows. Each cell is the number of unknowns left over after every equation the image supplies: three sensors, a three-dimensional illuminant, and reflectances confined to a linear model of the dimension on the left. At one and two dimensions more surfaces close the gap. At three the gap never closes, because each further surface adds three equations and three unknowns; at four it widens. The count is arithmetic and has no algorithm in it.
Fig. 1 The count. Each cell is how many unknowns are left over after every equation the image supplies. At one and two dimensions more surfaces close the gap; at three the gap never closes.

The claim

A scene of any number of surfaces, seen by three sensors under one unknown light, supplies three numbers per surface and asks for three numbers per surface plus the light. The count closes only when reflectances are confined to a linear model of dimension at most two — one less than the number of sensors — and real reflectances are not.

  • The arithmetic is one line. With N surfaces, p sensors, a k-dimensional reflectance model and an m-dimensional illuminant model, the problem is solvable when N(p − k) ≥ m − 1. At k = p the left side is zero for every N.
  • So sixty-four surfaces leave the same deficit as two. Each further surface adds three equations and three unknowns.
  • At three dimensions the alternatives can be constructed, not merely counted. Given any second illuminant, the reflectances that reproduce every sensor response exactly are one matrix multiplication away.
  • Five of seven alternative daylights tested stay physical — every reflectance inside zero and one — spanning 4,500 kelvin, with a worst residual of 10⁻¹⁵.
  • And two dimensions is not enough for reflectance. This collection’s own deliberately smooth constructed surfaces need three components to reach 99.8 per cent of their variance.

The count

Write a reflectance as ρᵢ(λ) = Σₐ σᵢₐ Bₐ(λ) on a fixed basis of k functions, and the illuminant as E(λ) = Σ_b ε_b Dₑ(λ) on a fixed basis of m. A sensor’s response to surface i is

rᵢ𝒸 = ∫ S𝒸(λ) E(λ) ρᵢ(λ) dλ

which is bilinear in ε and σᵢ. Every quantity in it is known except ε and the σᵢ.

The data are N p numbers. The unknowns are m + N k, less one for the scale that illuminant and reflectance trade between them — a light twice as bright with every surface half as reflective is the same image, always. So the deficit is

m + N k − 1 − N p = N(k − p) + m − 1

The whole argument is in the sign of k − p. If k < p, each additional surface subtracts from the deficit and enough surfaces will close it. If k = p, additional surfaces change nothing at all. If k > p, they make it worse.

With three sensors, that means reflectance must be at most two-dimensional. Not approximately; the count is exact.

How many numbers a surface takes, on the friendliest set available. The cumulative share of the variance accounted for by the first few principal components of this collection's 240 constructed reflectances. Three components reach 99.94 per cent and two reach 97.99. The counting result allows two, and these are deliberately smooth curves built from a handful of Gaussians — the friendliest possible case, and it already needs 3.
Fig. 2 And how many dimensions a reflectance actually takes, on the friendliest set available. Two components reach 98 per cent and three reach 99.94, on curves built deliberately from a handful of Gaussians.

The family, written down

Counting says a family exists. Constructing it says what the family contains, and at k = p the construction is short enough to state.

Collect the bilinear form into a matrix: rᵢ = T(ε) σᵢ, where T is p × k and depends only on the illuminant. When p = k this is square. So for any second illuminant ε′ whose T is invertible, setting

σ′ᵢ = T(ε′)⁻¹ T(ε) σᵢ

reproduces every sensor response exactly. There is no optimisation, no residual and no approximation — the equality is algebraic, and the computed residual of about 10⁻¹⁵ is machine arithmetic rather than a fit.

Six other lights that would have made exactly the same photograph. The scene is six surfaces from a three-dimensional reflectance model under a 6500 K daylight. Each mark is an alternative daylight, with the reflectances that reproduce every sensor response exactly — the residual is 1.8e-15, which is machine zero. Filled marks are the ones whose reflectances stay inside zero and one and are therefore real surfaces; 5 of 7 do, spanning 4500 kelvin. Nothing in the image chooses between them.
Fig. 3 Seven alternative daylights, each with the reflectances that reproduce the image exactly. Filled marks are the ones whose reflectances stay inside zero and one and are therefore real surfaces.

What the construction does need is a check that the alternative surfaces are still surfaces. A family of explanations requiring a reflectance of 1.33 is not a family of explanations, and two of the seven alternatives fail exactly that way — both of them cooler than the true light, which is the direction that demands brighter surfaces.

The five that survive span 5,500 K to 10,000 K. Nothing in the image distinguishes them. Something outside the image must, and that something is a prior.

Why the scale ambiguity is only subtracted once

The − 1 in the count is small and is worth a paragraph, because getting it wrong would change the conclusion for the two-dimensional case and it is the kind of term that is easy to double-count.

An overall scale can be moved between the illuminant and the surfaces: halve the light, double every reflectance, and every sensor response is unchanged. That is one degree of freedom in the whole problem, not one per surface, because the same factor has to be applied to every surface at once — a per-surface factor would change the ratios between surfaces, which the sensors do see.

So the unknowns are m + Nk with one of them redundant, and the deficit carries a single − 1. At k = p that means the deficit is m − 1, which for a three-dimensional illuminant family is two: two numbers, however large the scene. The brightness half of the ambiguity is separately unresolvable for reasons that have nothing to do with colour, which is why it is conventional to normalise it away and speak of the light’s chromaticity.

Which is what an estimator is

Once the family is explicit, every constancy algorithm reads differently. None of them is extracting information from the image that a cleverer one might extract better; each is a rule for choosing a member of the family, and the rule is an assumption about scenes.

Grey-world chooses the member whose surfaces average to neutral. It fails on an all-green scene because the assumption is false there, and no amount of care in implementing it helps.

Max-RGB chooses the member whose brightest surface is white, which is a claim that some surface in the scene is white.

Grey-edge makes the same assumption about differences rather than about values, which is a different and often better claim about natural scenes and is still a claim.

And the highlight estimator is the odd one out, because a specular reflection is the one part of a scene not multiplied by a reflectance. It does not choose a member of the family; it adds a measurement that breaks the degeneracy, which is a different kind of solution and is why it works when it works and fails completely when there is no highlight.

The distinction between those two kinds of solution is the useful thing the counting argument gives a practitioner. A method that chooses a member of the family can be evaluated only against a distribution of scenes, because its answer on any single scene is a prior’s answer; a method that adds a measurement can be evaluated on one scene, because it is measuring something. The first kind has a best-case error that is a property of the world; the second has one that is a property of the instrument. A whiter surface somewhere in the scene is the commonest assumption of the first kind and the commonest reason a camera gets it wrong.

What was computed, and how

The reflectance basis is a raised-cosine series, which is the standard smooth basis and is close to what this collection’s constructed surfaces already are. The illuminant basis is the CIE daylight reconstruction, which is genuinely three-dimensional because the CIE publishes three basis functions — this is not a convenient choice, it is the one the standard makes.

The surfaces are drawn with coefficients kept well inside the physical range, so that the alternatives have room to move before they stop being reflectances. That is a deliberate bias in favour of the family being large and it is stated rather than hidden: surfaces close to the boundary would make more alternatives non-physical and shrink the reported span.

The residual is computed by pushing each alternative’s reflectances back through its own illuminant and comparing sensor responses to the originals, relative to the originals. It is not a fit statistic; it is a check that the algebra was implemented as written.

The effective dimension is the uncentred principal components of the constructed set. Uncentred is the right choice because a linear model of reflectance has no mean term to remove — the first component is the average surface, and centring would answer a question about variation when the question is how many numbers it takes to write a surface down.

What each fitted thing in these essays carries, what its data fix, and what is left. Three columns per row: how many numbers the model has, how many the stated data determine, and the difference — the dimension of the family that fits equally well. The third column is the one nobody publishes. A zero there does not mean the model is right; it means it is determined, which is a much weaker property and is compatible with being determined badly, as the camera row is.
Fig. 4 The same three columns applied across the site. The scene row is the one this essay is about, and its last column is two — small, and never zero however large the scene gets.

The same three columns can be filled in for an instrument and for a chart, and both are cases where the count comes out at zero.

The part of a sensor a chart reaches, and the part it does not. Each channel's sensitivity, and the residue left after projecting it onto the span of the 12 illuminated patch spectra — the part no measurement of that chart constrains. A second camera differing by that residue produces identical raw values on every patch, to 2.9e-6 relative, and differs by 1.1% on a narrow light at 546 nm. The residue is 3.5% of each curve, which is small — patches that differ in where they are dark are a good probe, and that is the design rule.
Fig. 5 A camera’s residue: the part of each channel’s sensitivity that no measurement of any chart constrains. Here the unknowns outnumber the equations too, and the gap is filled by a convention rather than by an estimate.
What a camera matrix reports on its own chart, and what it delivers off it. The same camera fitted on charts of increasing chromatic range. The left bar of each pair is the mean error on the chart the matrix was fitted to, which is the number a profile comes with; the right bar is the error on a saturated set it never saw. At the thinnest chart the fit reports 0.19 ΔE00 and delivers 1.65, a factor of 8.6. The gap closes as the chart widens, and it closes because the chart improves rather than because the camera does.
Fig. 6 And a fit whose count comes out determined: the same camera on charts of increasing range. Nine numbers, nine equations, nothing left over — and a number that is nevertheless wrong about a photograph.

How large the family is, in the units the estimators are scored in

The five surviving alternatives are counted and their span is given in kelvin, which is the one unit that says nothing about how different they look. In the units the rest of this collection uses:

  • Across the family, 5500 K to 10000 K, is ΔE₀₀ 23.6 on the white itself, and 19.4 degrees of angular error — the measure every estimator on this site is scored by.
  • Against the true light, the two ends sit at 8.8 and 14.3 ΔE₀₀, or 6.8 and 12.6 degrees.

Those numbers put the ambiguity beside the estimators rather than above them. Grey-world’s celebrated failure on an all-green scene is 43 degrees; the family the image cannot distinguish is nineteen, and an estimator landing anywhere inside it has made no error the data could have caught. So the distance between a good estimator and a bad one on a single scene is of the same order as the distance the scene itself does not determine, which is why the essay is right that a method choosing a member can be evaluated only against a distribution of scenes.

It also gives the practical version of the count. Two numbers of deficit is a nineteen-degree interval, not an abstract pair of unknowns, and a camera that lands at either end produces a photograph a person would call plainly miscoloured. The count says the deficit does not shrink with scene size; this says the deficit is large enough to see.

The family is one-sided here and cannot be in general

Two of the seven alternatives are rejected for demanding a reflectance above one, and both fail on the same side. That is a property of these surfaces rather than of the problem.

An alternative illuminant with less power than the true one in some band forces the surfaces to reflect more there, and far enough in either direction the requirement passes one — a much bluer light demands impossible long-wave reflectance, and a much redder one demands impossible short-wave reflectance. So the physical family is bounded on both sides, and a test finding all its failures on one side has not reached the other boundary yet.

The essay says why: the surfaces were drawn with coefficients kept well inside the physical range, and names that as a deliberate bias in favour of the family being large. The one-sidedness is the same bias showing through — the tested range simply stopped before the other wall. A set of surfaces drawn nearer the boundary would fail on both sides and would report a narrower family, and the honest reading of the 4,500-kelvin span is therefore an upper bound on the ambiguity for these surfaces rather than a measurement of it.

That does not weaken the argument, because the argument needs the family to be non-empty rather than wide. A family of two members that both look like daylight is still a family the image cannot choose between.

Two figures for one variance

The effective dimension is quoted twice and the two do not match: 99.8 per cent at three components in the claim list, 99.94 in the figure caption, with the caption also giving 98 at two.

Both cannot be the same measurement, and the gap matters more than its size, because the whole argument’s second half turns on whether three components are needed or merely nice. At 98 per cent for two components the residual is a fiftieth; at 99.94 for three it is a sixteen-hundredth. The ratio between those two residuals is 33, so the third component is not a marginal refinement — it removes 97 per cent of what two components leave.

Which is the number the argument wants, and it is worth stating that way rather than as a percentage of variance explained. A percentage close to a hundred always reads as nearly enough; the residual ratio reads as what it is, which is that dropping to two dimensions multiplies the reflectance error by more than thirty on the friendliest set this collection has.

Where the model stops

Everything here is linear-model constancy. Real scenes have shading, mutual illumination, specularities, more than one light source and objects that are not flat matte patches — and several of those add information rather than removing it. A corner receives light carrying a reflectance twice, which is a nonlinearity in reflectance and is therefore a constraint the count above does not include.

That is the honest limit of the argument: the counting result applies to the idealisation, and the idealisation is what every classical estimator assumes. A method exploiting interreflection, or shadow boundaries, or the statistics of natural images is not refuted by it.

The dimension of real reflectance is a measurement, not a constant. Different sets give different answers — measured Munsell chips, natural surfaces, printing inks — and none of them gives two. What is computed here is the friendliest case available, on constructed curves, and even that gives three.

The illuminant model is a strong assumption too, and it is the one doing most of the work in the favourable cases. Restricting lights to the daylight family is what makes m three rather than eighty-one, and a scene lit by a fluorescent tube or a white LED is not in that family at all. A method assuming daylight and given a discharge lamp is not choosing badly among the alternatives; it is choosing among a set that does not contain the answer.

And a camera is not three ideal sensors. Its filters do not satisfy the Luther condition, which means its responses are not linear combinations of tristimulus values at all, and the count is unaffected: three sensors are three sensors whatever their shapes.

Who found it, and when

Laurence Maloney and Brian Wandell set the count out in 1986, and the result is usually stated as its conclusion — that with three sensors, surfaces must be two-dimensional. The paper is careful about the accompanying empirical fact that they are not, and about what follows: constancy in the strict sense is unavailable and what people and cameras do must be something else.

The something else has been the subject of the forty years since. Bayesian formulations put a prior on illuminants and surfaces explicitly and are the direct descendants of the counting argument; the gamut-mapping methods use the constraint that reflectances lie in a known convex set; learned methods estimate the prior from data. All of them are answers to “which member of the family”, and the counting argument is what makes that the right question.

The generalisation

The pattern is a problem where more data of the same kind cannot help, and where this is decidable in advance by counting.

The diagnostic is the per-datum balance: how many equations does one more observation supply, and how many unknowns does it bring with it? If the two are equal, the problem has a fixed deficit and a larger data set is wasted effort. That is a question about the structure of the model and can be answered before any data are collected, which makes it one of the cheapest checks available and one of the least often run.

Where the balance is unfavourable there are exactly three moves: add sensors, reduce the model’s dimension, or add a measurement of a different kind. A better algorithm is not on the list.

What people do, which is not this

It would be wrong to leave the impression that human colour constancy is refuted by the count, and it is worth saying exactly why not.

People are demonstrably good at constancy and demonstrably imperfect at it, and the counting argument predicts precisely that combination. What it forbids is exact recovery of surface reflectance from a three-channel image of a flat scene under a single unknown light. It forbids nothing about a visual system that uses shading, mutual illumination, familiar objects, the statistics of natural spectra, and a memory of what the light was doing a second ago — all of which are extra measurements or extra priors, and several of which this collection models elsewhere.

The count is a statement about an idealisation, and the idealisation is the one every textbook diagram of colour constancy draws. Its value is that it converts a vague sense that the problem is hard into an exact statement about what would have to be added, and the list of what would have to be added is short and checkable.

The move that does work

Of the three ways out the count allows, one has quietly become ordinary and is worth naming as the successful case.

Adding sensors changes the sign of k − p. A four-channel or six-channel camera turns an unidentifiable problem into a solvable one, and the arithmetic is not marginal: with four sensors a three-dimensional reflectance model leaves one equation per surface to spare, so a scene of a few surfaces determines the light outright.

That is why multispectral imaging exists in the places where getting the light right actually matters — art reproduction, textiles, food inspection, remote sensing — and it is a much better argument for it than the usual one about capturing spectra. The usual argument is about fidelity and this one is about possibility, and the second is the harder claim: a three-channel system is not merely less accurate at illuminant estimation, it is doing something the data do not support.

The other two moves are worse bargains. Reducing the reflectance model to two dimensions is choosing to be wrong about most surfaces; adding a measurement of a different kind means finding a highlight, a known surface or a second light, none of which a photographer controls.

Where the ladder goes next

The same counting question asked of a camera’s own calibration gives a sharper answer, because there the data are chosen rather than found: a chart determines the profile it can determine, and a chart with no chromatic range determines it badly while reporting excellent numbers.

And the instrument itself is the other direction. Three readings of a spectrum determine a three-dimensional projection of it and nothing else, which is why a colorimeter cannot see a mercury line and why the colour it reports can be right while the spectrum it implies is wrong.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 16 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Colour constancyThe D-series daylight illuminantsDegrees of freedomThe grey-world assumptionIdentifiabilityIlluminant estimationLinear modelPrincipal componentsReflectanceScene interpretationWhite balance