An image does not determine the light
Assumes A scene has no white point, The algorithms that guess the light and A bounce is a multiplication.
Every estimator this site has guesses the illuminant of a scene and every one of them is wrong on some scenes. Grey-world reaches 43 degrees of angular error on an all-green scene; grey-edge and max-RGB fail elsewhere; no estimator wins on every scene, so a camera choosing one has chosen which photographs to be wrong about.
That reads like an engineering problem awaiting a better algorithm. It is not.
The claim
A scene of any number of surfaces, seen by three sensors under one unknown light, supplies three numbers per surface and asks for three numbers per surface plus the light. The count closes only when reflectances are confined to a linear model of dimension at most two — one less than the number of sensors — and real reflectances are not.
- The arithmetic is one line. With
Nsurfaces,psensors, ak-dimensional reflectance model and anm-dimensional illuminant model, the problem is solvable whenN(p − k) ≥ m − 1. Atk = pthe left side is zero for everyN. - So sixty-four surfaces leave the same deficit as two. Each further surface adds three equations and three unknowns.
- At three dimensions the alternatives can be constructed, not merely counted. Given any second illuminant, the reflectances that reproduce every sensor response exactly are one matrix multiplication away.
- Five of seven alternative daylights tested stay physical — every reflectance inside zero and one — spanning 4,500 kelvin, with a worst residual of 10⁻¹⁵.
- And two dimensions is not enough for reflectance. This collection’s own deliberately smooth constructed surfaces need three components to reach 99.8 per cent of their variance.
The count
Write a reflectance as ρᵢ(λ) = Σₐ σᵢₐ Bₐ(λ) on a fixed basis of k functions, and the illuminant as E(λ) = Σ_b ε_b Dₑ(λ) on a fixed basis of m. A sensor’s response to surface i is
rᵢ𝒸 = ∫ S𝒸(λ) E(λ) ρᵢ(λ) dλ
which is bilinear in ε and σᵢ. Every quantity in it is known except ε and the σᵢ.
The data are N p numbers. The unknowns are m + N k, less one for the scale that illuminant and reflectance trade between them — a light twice as bright with every surface half as reflective is the same image, always. So the deficit is
m + N k − 1 − N p = N(k − p) + m − 1
The whole argument is in the sign of k − p. If k < p, each additional surface subtracts from the deficit and enough surfaces will close it. If k = p, additional surfaces change nothing at all. If k > p, they make it worse.
With three sensors, that means reflectance must be at most two-dimensional. Not approximately; the count is exact.
The family, written down
Counting says a family exists. Constructing it says what the family contains, and at k = p the construction is short enough to state.
Collect the bilinear form into a matrix: rᵢ = T(ε) σᵢ, where T is p × k and depends only on the illuminant. When p = k this is square. So for any second illuminant ε′ whose T is invertible, setting
σ′ᵢ = T(ε′)⁻¹ T(ε) σᵢ
reproduces every sensor response exactly. There is no optimisation, no residual and no approximation — the equality is algebraic, and the computed residual of about 10⁻¹⁵ is machine arithmetic rather than a fit.
What the construction does need is a check that the alternative surfaces are still surfaces. A family of explanations requiring a reflectance of 1.33 is not a family of explanations, and two of the seven alternatives fail exactly that way — both of them cooler than the true light, which is the direction that demands brighter surfaces.
The five that survive span 5,500 K to 10,000 K. Nothing in the image distinguishes them. Something outside the image must, and that something is a prior.
Why the scale ambiguity is only subtracted once
The − 1 in the count is small and is worth a paragraph, because getting it wrong would change the conclusion for the two-dimensional case and it is the kind of term that is easy to double-count.
An overall scale can be moved between the illuminant and the surfaces: halve the light, double every reflectance, and every sensor response is unchanged. That is one degree of freedom in the whole problem, not one per surface, because the same factor has to be applied to every surface at once — a per-surface factor would change the ratios between surfaces, which the sensors do see.
So the unknowns are m + Nk with one of them redundant, and the deficit carries a single − 1. At k = p that means the deficit is m − 1, which for a three-dimensional illuminant family is two: two numbers, however large the scene. The brightness half of the ambiguity is separately unresolvable for reasons that have nothing to do with colour, which is why it is conventional to normalise it away and speak of the light’s chromaticity.
Which is what an estimator is
Once the family is explicit, every constancy algorithm reads differently. None of them is extracting information from the image that a cleverer one might extract better; each is a rule for choosing a member of the family, and the rule is an assumption about scenes.
Grey-world chooses the member whose surfaces average to neutral. It fails on an all-green scene because the assumption is false there, and no amount of care in implementing it helps.
Max-RGB chooses the member whose brightest surface is white, which is a claim that some surface in the scene is white.
Grey-edge makes the same assumption about differences rather than about values, which is a different and often better claim about natural scenes and is still a claim.
And the highlight estimator is the odd one out, because a specular reflection is the one part of a scene not multiplied by a reflectance. It does not choose a member of the family; it adds a measurement that breaks the degeneracy, which is a different kind of solution and is why it works when it works and fails completely when there is no highlight.
The distinction between those two kinds of solution is the useful thing the counting argument gives a practitioner. A method that chooses a member of the family can be evaluated only against a distribution of scenes, because its answer on any single scene is a prior’s answer; a method that adds a measurement can be evaluated on one scene, because it is measuring something. The first kind has a best-case error that is a property of the world; the second has one that is a property of the instrument. A whiter surface somewhere in the scene is the commonest assumption of the first kind and the commonest reason a camera gets it wrong.
What was computed, and how
The reflectance basis is a raised-cosine series, which is the standard smooth basis and is close to what this collection’s constructed surfaces already are. The illuminant basis is the CIE daylight reconstruction, which is genuinely three-dimensional because the CIE publishes three basis functions — this is not a convenient choice, it is the one the standard makes.
The surfaces are drawn with coefficients kept well inside the physical range, so that the alternatives have room to move before they stop being reflectances. That is a deliberate bias in favour of the family being large and it is stated rather than hidden: surfaces close to the boundary would make more alternatives non-physical and shrink the reported span.
The residual is computed by pushing each alternative’s reflectances back through its own illuminant and comparing sensor responses to the originals, relative to the originals. It is not a fit statistic; it is a check that the algebra was implemented as written.
The effective dimension is the uncentred principal components of the constructed set. Uncentred is the right choice because a linear model of reflectance has no mean term to remove — the first component is the average surface, and centring would answer a question about variation when the question is how many numbers it takes to write a surface down.
The same three columns can be filled in for an instrument and for a chart, and both are cases where the count comes out at zero.
How large the family is, in the units the estimators are scored in
The five surviving alternatives are counted and their span is given in kelvin, which is the one unit that says nothing about how different they look. In the units the rest of this collection uses:
- Across the family, 5500 K to 10000 K, is ΔE₀₀ 23.6 on the white itself, and 19.4 degrees of angular error — the measure every estimator on this site is scored by.
- Against the true light, the two ends sit at 8.8 and 14.3 ΔE₀₀, or 6.8 and 12.6 degrees.
Those numbers put the ambiguity beside the estimators rather than above them. Grey-world’s celebrated failure on an all-green scene is 43 degrees; the family the image cannot distinguish is nineteen, and an estimator landing anywhere inside it has made no error the data could have caught. So the distance between a good estimator and a bad one on a single scene is of the same order as the distance the scene itself does not determine, which is why the essay is right that a method choosing a member can be evaluated only against a distribution of scenes.
It also gives the practical version of the count. Two numbers of deficit is a nineteen-degree interval, not an abstract pair of unknowns, and a camera that lands at either end produces a photograph a person would call plainly miscoloured. The count says the deficit does not shrink with scene size; this says the deficit is large enough to see.
The family is one-sided here and cannot be in general
Two of the seven alternatives are rejected for demanding a reflectance above one, and both fail on the same side. That is a property of these surfaces rather than of the problem.
An alternative illuminant with less power than the true one in some band forces the surfaces to reflect more there, and far enough in either direction the requirement passes one — a much bluer light demands impossible long-wave reflectance, and a much redder one demands impossible short-wave reflectance. So the physical family is bounded on both sides, and a test finding all its failures on one side has not reached the other boundary yet.
The essay says why: the surfaces were drawn with coefficients kept well inside the physical range, and names that as a deliberate bias in favour of the family being large. The one-sidedness is the same bias showing through — the tested range simply stopped before the other wall. A set of surfaces drawn nearer the boundary would fail on both sides and would report a narrower family, and the honest reading of the 4,500-kelvin span is therefore an upper bound on the ambiguity for these surfaces rather than a measurement of it.
That does not weaken the argument, because the argument needs the family to be non-empty rather than wide. A family of two members that both look like daylight is still a family the image cannot choose between.
Two figures for one variance
The effective dimension is quoted twice and the two do not match: 99.8 per cent at three components in the claim list, 99.94 in the figure caption, with the caption also giving 98 at two.
Both cannot be the same measurement, and the gap matters more than its size, because the whole argument’s second half turns on whether three components are needed or merely nice. At 98 per cent for two components the residual is a fiftieth; at 99.94 for three it is a sixteen-hundredth. The ratio between those two residuals is 33, so the third component is not a marginal refinement — it removes 97 per cent of what two components leave.
Which is the number the argument wants, and it is worth stating that way rather than as a percentage of variance explained. A percentage close to a hundred always reads as nearly enough; the residual ratio reads as what it is, which is that dropping to two dimensions multiplies the reflectance error by more than thirty on the friendliest set this collection has.
Where the model stops
Everything here is linear-model constancy. Real scenes have shading, mutual illumination, specularities, more than one light source and objects that are not flat matte patches — and several of those add information rather than removing it. A corner receives light carrying a reflectance twice, which is a nonlinearity in reflectance and is therefore a constraint the count above does not include.
That is the honest limit of the argument: the counting result applies to the idealisation, and the idealisation is what every classical estimator assumes. A method exploiting interreflection, or shadow boundaries, or the statistics of natural images is not refuted by it.
The dimension of real reflectance is a measurement, not a constant. Different sets give different answers — measured Munsell chips, natural surfaces, printing inks — and none of them gives two. What is computed here is the friendliest case available, on constructed curves, and even that gives three.
The illuminant model is a strong assumption too, and it is the one doing most of the work in the favourable cases. Restricting lights to the daylight family is what makes m three rather than eighty-one, and a scene lit by a fluorescent tube or a white LED is not in that family at all. A method assuming daylight and given a discharge lamp is not choosing badly among the alternatives; it is choosing among a set that does not contain the answer.
And a camera is not three ideal sensors. Its filters do not satisfy the Luther condition, which means its responses are not linear combinations of tristimulus values at all, and the count is unaffected: three sensors are three sensors whatever their shapes.
Who found it, and when
Laurence Maloney and Brian Wandell set the count out in 1986, and the result is usually stated as its conclusion — that with three sensors, surfaces must be two-dimensional. The paper is careful about the accompanying empirical fact that they are not, and about what follows: constancy in the strict sense is unavailable and what people and cameras do must be something else.
The something else has been the subject of the forty years since. Bayesian formulations put a prior on illuminants and surfaces explicitly and are the direct descendants of the counting argument; the gamut-mapping methods use the constraint that reflectances lie in a known convex set; learned methods estimate the prior from data. All of them are answers to “which member of the family”, and the counting argument is what makes that the right question.
The generalisation
The pattern is a problem where more data of the same kind cannot help, and where this is decidable in advance by counting.
The diagnostic is the per-datum balance: how many equations does one more observation supply, and how many unknowns does it bring with it? If the two are equal, the problem has a fixed deficit and a larger data set is wasted effort. That is a question about the structure of the model and can be answered before any data are collected, which makes it one of the cheapest checks available and one of the least often run.
Where the balance is unfavourable there are exactly three moves: add sensors, reduce the model’s dimension, or add a measurement of a different kind. A better algorithm is not on the list.
What people do, which is not this
It would be wrong to leave the impression that human colour constancy is refuted by the count, and it is worth saying exactly why not.
People are demonstrably good at constancy and demonstrably imperfect at it, and the counting argument predicts precisely that combination. What it forbids is exact recovery of surface reflectance from a three-channel image of a flat scene under a single unknown light. It forbids nothing about a visual system that uses shading, mutual illumination, familiar objects, the statistics of natural spectra, and a memory of what the light was doing a second ago — all of which are extra measurements or extra priors, and several of which this collection models elsewhere.
The count is a statement about an idealisation, and the idealisation is the one every textbook diagram of colour constancy draws. Its value is that it converts a vague sense that the problem is hard into an exact statement about what would have to be added, and the list of what would have to be added is short and checkable.
The move that does work
Of the three ways out the count allows, one has quietly become ordinary and is worth naming as the successful case.
Adding sensors changes the sign of k − p. A four-channel or six-channel camera turns an unidentifiable problem into a solvable one, and the arithmetic is not marginal: with four sensors a three-dimensional reflectance model leaves one equation per surface to spare, so a scene of a few surfaces determines the light outright.
That is why multispectral imaging exists in the places where getting the light right actually matters — art reproduction, textiles, food inspection, remote sensing — and it is a much better argument for it than the usual one about capturing spectra. The usual argument is about fidelity and this one is about possibility, and the second is the harder claim: a three-channel system is not merely less accurate at illuminant estimation, it is doing something the data do not support.
The other two moves are worse bargains. Reducing the reflectance model to two dimensions is choosing to be wrong about most surfaces; adding a measurement of a different kind means finding a highlight, a known surface or a second light, none of which a photographer controls.
Where the ladder goes next
The same counting question asked of a camera’s own calibration gives a sharper answer, because there the data are chosen rather than found: a chart determines the profile it can determine, and a chart with no chromatic range determines it badly while reporting excellent numbers.
And the instrument itself is the other direction. Three readings of a spectrum determine a three-dimensional projection of it and nothing else, which is why a colorimeter cannot see a mercury line and why the colour it reports can be right while the spectrum it implies is wrong.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- Constancy is the default colour constancy · illuminant estimation · reflectance · white balance
- The average surface does not look average colour constancy · the grey-world assumption · illuminant estimation · reflectance
- A confusion point is a missing pigment degrees of freedom · identifiability
- A constraint costs what it points at degrees of freedom · identifiability
- A constraint is a direction and a distance degrees of freedom · identifiability
- A corner moves both terms colour constancy · reflectance
What links here
The 8 essays that link to this one and share the most of its objects, of 16 that link here.
The objects this essay names
Each one links to every other essay that touches it.
Colour constancyThe D-series daylight illuminantsDegrees of freedomThe grey-world assumptionIdentifiabilityIlluminant estimationLinear modelPrincipal componentsReflectanceScene interpretationWhite balance