What the brain does

Constancy is the default

A sheet of paper looks white in daylight and white under a tungsten lamp, although the light reaching the eye differs enormously. The visual system is solving one equation with two unknowns, and it solves it by assumption.

A sheet of paper looks white in daylight and white under a tungsten lamp. The light reaching the eye from it differs by a large factor in the ratio of long to short wavelengths — a photograph taken under the lamp without correction comes out plainly orange — and the paper looks white in both.

One reflectance, two illuminants, two coloursA reflectance peaking near 580 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.400450500550600650700wavelength / nmunder D650.428, 0.474under A0.500, 0.467reflectance is a fraction, 0 to 1CIE 1931 2° observer
Fig. 1 A single surface under two illuminants, as the stimulus rather than as the appearance. The two swatches are what a camera records. What an observer sitting in each room perceives is much closer to the same colour, and the difference between those two statements is what this essay is about.

This is colour constancy, and it is not a curiosity. It is the normal operation of the system, and the effects usually presented as illusions are cases where it produces an answer the viewer can recognise as wrong.

The problem being solved

The light arriving from a surface is the product of two things: the surface’s reflectance and the illumination falling on it. The eye measures the product and needs the first factor.

stimulus(λ)=reflectance(λ)×illuminant(λ)\text{stimulus}(\lambda) = \text{reflectance}(\lambda) \times \text{illuminant}(\lambda)

One measurement, two unknowns. The problem is underdetermined and no amount of cleverness makes it otherwise — any stimulus is consistent with a bright surface under dim light or a dark surface under bright light, and with every intermediate.

There is a good reason to want the reflectance rather than the product. Reflectance is a property of the object; it identifies materials, it stays constant as conditions change, it is what makes an object recognisable across a day. Illumination varies by orders of magnitude and carries little useful information about anything.

So the visual system estimates the illuminant and divides it out. Since the problem is underdetermined, the estimate rests on assumptions, and the assumptions are what everything else here follows from.

The assumptions

Several are used, and none of them is guaranteed.

The brightest thing is white. A useful heuristic: scenes usually contain something near-white, and the brightest surface is a reasonable estimate of the illuminant’s colour. This is the anchoring principle that also fixes the lightness scale.

The average is grey. The grey-world assumption: averaged over a whole scene, reflectances tend toward neutral, so any tint in the average is attributable to the light. It works well for varied scenes and fails badly for a photograph of a green field.

Specular highlights carry the illuminant. A glossy surface reflects some light without spectral modification, so a highlight is a direct sample of the illumination. This is a genuinely good cue and the visual system appears to use it.

Illumination varies smoothly, surfaces do not. Sharp changes are attributed to surfaces, gradual ones to lighting. This is what the Cornsweet effect exploits.

Each is a sound generalisation about the world and each can be violated, deliberately or by accident. Everything called a colour illusion is a stimulus that violates one.

The arithmetic

The computational core is simpler than the phenomenon suggests, and is called von Kries adaptation: scale each cone response by a factor set by the illuminant.

L=LLw,M=MMw,S=SSwL' = \frac{L}{L_w},\qquad M' = \frac{M}{M_w},\qquad S' = \frac{S}{S_w}

where LwL_w, MwM_w, SwS_w are the responses to the estimated white. Each receptor class independently normalises to what it is receiving on average, which is exactly what a gain control does, and there is physiological evidence for adaptation at the receptor level.

Modern chromatic adaptation transforms — Bradford, CAT02, CAT16 — are refinements of this. Each converts to a cone-like space chosen to make the scaling work better, applies the diagonal scaling, and converts back. This site uses Bradford wherever a white point has to be changed.

The important structural point is that adaptation is a diagonal transformation in a well-chosen basis. It is not a complicated operation, and the difficulty in colour constancy is not the correction — it is estimating the illuminant to correct by, which is the underdetermined part.

Where it fails

Three familiar failures, all of them assumption violations.

Photographs. A camera does not adapt; it records the product. Automatic white balance is an attempt to estimate the illuminant using the same heuristics the visual system uses, and it fails in the same situations — a scene dominated by one colour defeats the grey-world assumption, which is why photographs of sunsets and forests are so often corrected in the wrong direction.

Mixed lighting. A room with daylight from a window and tungsten from a lamp has two illuminants. The system adapts to something in between, and surfaces in each region look tinted in opposite directions. Neither can be right, and the effect is visible to anyone who looks for it at dusk.

Narrow-band sources. Sodium street lighting emits essentially one wavelength. There is nothing to estimate and nothing to divide out — all surfaces return a scaled version of the same spectrum, so colour vision fails almost entirely. The scene appears in shades of yellow, and no amount of adaptation recovers anything, because the information is not present in the stimulus.

That last case is the clearest demonstration that constancy is an inference and not a correction. Where the light carries no information about the surfaces, no processing recovers the surfaces.

The dress

The photograph that circulated in 2015, seen by some as white and gold and by others as blue and black, is the best-known instance of the estimation genuinely being ambiguous.

The image was poorly exposed and cropped so that almost every illuminant cue was absent — no clear white reference, no reliable highlight, an uncertain background. Under those conditions the assumptions do not converge, and different viewers settled on different illuminant estimates. Assume a bluish daylight illuminant and divide it out, and the fabric is white and gold; assume a warm artificial one, and it is blue and black.

Both processes were operating exactly as designed. The stimulus was constructed — accidentally — to leave the problem underdetermined in a way that ordinary scenes do not, and the disagreement was between priors rather than between visual systems.

What constancy costs

The mechanism has a price, and the price is that a colour has no fixed appearance.

Two identical grey patches on different surroundsBoth inner squares are #868686. The one on the dark field looks lighter. The values are asserted equal in code before the figure is drawn, so the claim is a fact about the drawing rather than a promise.on a dark fieldon a light fieldboth patches are #868686asserted equal
Fig. 2 The same machinery operating on a local scale. Two identical patches, two surrounds, two lightnesses. The system is estimating the illumination from the local neighbourhood, and the neighbourhoods differ.

This is why a specification of a colour value does not specify an appearance, why a paint chip is misleading in the shop, and why a colour chosen in one context has to be rechecked in another. The visual system is not reporting the stimulus; it is reporting its best estimate of the surface, and the estimate depends on everything else in view.

It is also why colorimetry stops where it does. Predicting a match under identical conditions requires no model of the estimation. Predicting appearance requires knowing what the system has adapted to, which is why appearance models take viewing conditions as an argument.

The failure cases, drawn

Each of the assumptions can be violated, and the violations are what the illusion literature is made of.

A checkerboard under a shadow, with two squares markedThe two marked squares are both #797979. One is a dark square in full light; the other is a light square inside the shadow. Nothing about the drawing enforces the illusion — the two values are set equal in code and asserted before drawing.ABA and B are both #797979asserted equal
Fig. 3 The illumination assumption violated deliberately. The visual system identifies a shadow, discounts the illumination inside it, and reports reflectances — which is correct about the world and wrong about the image, because there is no shadow and no illumination to discount.
The same checkerboard with the two squares joinedThe two marked squares are both #797979. One is a dark square in full light; the other is a light square inside the shadow. The connecting bar is the same value again, and joining them makes the equality obvious.ABA and B are both #797979asserted equal
Fig. 4 The same board with the two squares joined by a bar of the same value. A continuous region of one value cannot be two surfaces at different illuminations, so the parse that produced the effect is no longer available.

The connecting bar is the clearest demonstration that constancy is an inference rather than a filter. Nothing about the two squares changed; what changed is the evidence available about how the scene is organised, and the inference changed with it.

The Cornsweet edge, with its luminance profileThe two plateaux are both #898989 and are flat to exactly — the profile below shows that everything which differs lies within a narrow band at the boundary. Cover the centre line and the two halves become obviously identical.one value on the left, the same value on the rightluminance profileboth plateaux are #898989asserted equal
Fig. 5 The smoothness assumption violated. Gradual changes are attributed to illumination and discounted; sharp ones are attributed to surfaces and believed. This profile is constructed so that its gradual part is discounted and its sharp part is believed, and the result propagates across regions where nothing differs.

Constancy is not perfect, and the imperfection is measurable

It would be easy to read the above as claiming the visual system fully discounts the illuminant. It does not.

Measured constancy indices — how much of the illuminant change is compensated — typically fall between 60 and 90 per cent under good conditions, and lower when cues are sparse. Surfaces do look somewhat warmer under tungsten; the correction is substantial and partial.

This is worth stating because the two extreme positions are both wrong. Colour is not simply a property of the surface, recovered perfectly; nor is it simply the stimulus, with no correction at all. It is an estimate, it is usually good, and it degrades gracefully as the evidence gets worse — which is what one would expect of an inference under uncertainty.

The degradation is why photographs of unusual lighting look wrong in a way the scene did not, and why two people can settle on different estimates when the cues are ambiguous enough.

What was computed here

The chromatic adaptation transforms are implemented in this site’s colour library — Bradford and CAT16 — and are used wherever a white point changes, in particular when converting between D50 and D65 based spaces.

They are checked by round trip: adapting from D65 to D50 and back must return the starting XYZ, and it does to better than 101010^{-10}. That catches an inverted matrix, which otherwise produces a plausible warm or cool shift that looks like a deliberate choice.

What is not computed on this page is any constancy prediction. The figures show stimuli, and the essay describes what the visual system does with them. Modelling the estimation would require an appearance model, which is the next phase’s work rather than something the foundation machinery can support — and claiming otherwise would be exactly the overreach this site is built to avoid.

The adaptation transforms are checked by round trip: adapting D65 to D50 and back must return the starting XYZ, and does to better than 1e-10. That catches an inverted matrix, which otherwise produces a plausible warm or cool shift indistinguishable from a deliberate choice.

What a camera does instead

Comparing the two systems is the clearest way to see what constancy is doing, because a camera has the same problem and no visual system.

A camera records the product of reflectance and illumination and nothing else. Automatic white balance estimates the illuminant using the same heuristics — grey world, brightest-is-white, sometimes highlight detection — and divides it out, which is von Kries adaptation implemented in software.

It fails in the same situations, and for the same reasons. A photograph dominated by one colour defeats the grey-world assumption, so pictures of sunsets are corrected toward neutral and lose the warmth that was the subject. Mixed lighting produces a compromise wrong in both regions.

The difference is that a camera’s failure is visible afterwards and the visual system’s is not, because there is no unadapted view to compare against.

Four standard illuminants, and how little they have in commonSpectral power distributions for A, D65, E, on one scale. Illuminant A rises steeply toward the red; the daylight illuminants carry the atmosphere's absorption structure; E is flat by definition. All four are ordinarily called white.400450500550600650700wavelength / nmA0.448, 0.407D650.313, 0.329E0.333, 0.333normalised to 100 at 560 nmCIE 1931 2° observer
Fig. 6 The illuminants whose differences the system is discounting. Something adapted to the first and something adapted to the second report similar colours for the same surfaces, and a camera adapted to neither reports the difference in full.

The limit case

Low-pressure sodium lighting is the cleanest demonstration that constancy is an inference rather than a correction.

The lamp emits essentially one wavelength. Every surface returns a scaled copy of the same spectrum, differing only in how much. There is no chromatic information in the reflected light at all, so there is nothing to estimate and nothing to divide out, and colour vision fails almost completely — the scene appears in shades of one hue.

No processing recovers the surfaces, because the information is not present. This is worth holding onto as the boundary condition: constancy works by exploiting structure in the light, and where the light has no structure it has nothing to work with.

How good the estimate is

Constancy indices — the fraction of an illuminant change that gets compensated — typically fall between 60 and 90 per cent under good conditions and lower when cues are sparse.

That figure disposes of both extreme positions. Colour is not simply a property of the surface recovered perfectly, and it is not simply the stimulus with no correction. It is an estimate, usually good, degrading gracefully as evidence worsens — which is what an inference under uncertainty should do.

It is also worth noticing what constancy implies about the word colour. If the visual system’s output is an estimate of reflectance rather than a report of the stimulus, then everyday colour language refers to the estimate — the property of the object — and colorimetry refers to the stimulus. Two different quantities, both reasonably called colour, and most disputes about whether two things are the same colour turn out to be disputes about which one is meant.

What the pictures cannot show

The central limitation is unavoidable and worth being blunt about. Demonstrating adaptation requires controlling what the observer is adapted to, and this page controls none of it. A reader adapted to their room is looking at a screen inside that room, so every swatch here is already being interpreted against a state the figure knows nothing about.

The two swatches in the hero figure are a particular case. They are the stimuli a camera would record under two illuminants, presented side by side on one screen under a third. That is not the comparison an observer makes when walking from one room to another, and it exaggerates the difference substantially — side-by-side presentation defeats adaptation, which is exactly the point of the figure and also its limitation.

Nothing here can show what the paper looks like to somebody in the tungsten room, because showing it would require putting the reader in that room.

Who found it, and when

Monge described constancy in 1789. Helmholtz gave the standard formulation in the nineteenth century, calling it unconscious inference — the visual system draws conclusions from assumptions without any of it being available to introspection, which remains a fair description.

Johannes von Kries proposed the independent receptor-scaling rule around 1902, and it is still the core of every adaptation transform in use.

Edwin Land’s experiments in the 1950s and 1960s were the most striking demonstrations: scenes lit by two or three narrow-band sources in which surfaces retained their apparent colours far better than the stimulus alone could explain. His retinex theory has not survived as a mechanism, and the demonstrations remain the clearest evidence that constancy is computed from spatial relationships rather than from local measurements.

Where this goes next

The local version of the same machinery is these two patches are identical and brightness is inferred from edges. The physics being discounted is the illuminant is half the answer. And the boundary this places on colorimetry is matching is not appearance.