What the brain does

Constancy is the default

A sheet of paper looks white in daylight and white under a tungsten lamp, although the light reaching the eye differs enormously. The visual system is solving one equation with two unknowns, and it solves it by assumption.

Assumes These two patches are identical.

A sheet of paper looks white in daylight and white under a tungsten lamp. The light reaching the eye from it differs by a large factor in the ratio of long to short wavelengths — a photograph taken under the lamp without correction comes out plainly orange — and the paper looks white in both.

One reflectance, two illuminants, two coloursA reflectance peaking near 580 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.400450500550600650700wavelength / nmunder D650.423, 0.472under A0.498, 0.466reflectance is a fraction, 0 to 1CIE 1931 2° observer
Fig. 1 A single surface under two illuminants, as the stimulus rather than as the appearance. The two swatches are what a camera records. What an observer sitting in each room perceives is much closer to the same colour, and the difference between those two statements is what this essay is about.

One surface is one case. The claim is about every surface, and the stimulus moves for all of them.

One reflectance, two illuminants, two colours. A reflectance peaking near 480 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.
Fig. 2 A surface peaking in the blue, under the same two lights. The pair of colours it produces is further apart than the pair above, because a tungsten lamp has least to offer exactly where this surface reflects most.

A surface at the other end of the spectrum is the opposite case, and the two together are why a single gain cannot be right for both.

One reflectance, two illuminants, two colours. A reflectance peaking near 640 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.
Fig. 3 And a narrow red one, where the same change of light barely moves the reading at all. The stimulus changes by a different amount for every surface in the room, and a single gain is what the observer has to answer all of them with.

The size of the change is a property of the pair of lights as much as of the surface, which is checked by replacing one of the lights with a reference that is not a light at all.

One reflectance, two illuminants, two colours. A reflectance peaking near 580 nm, and the colours it produces under D65 and E. The object has not changed. The light has, and colour is a property of the pair.
Fig. 4 The first surface again under an equal-energy reference rather than tungsten — a light nobody has ever built, which is what makes it the control. The two readings are close because the two lights are, not because the surface is doing anything different.

Two more surfaces fill in the middle of the range, and between the five of them the change of light is worth a factor of several.

One reflectance, two illuminants, two colours. A reflectance peaking near 520 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.
Fig. 5 A surface peaking in the green, which is where the eye’s own sensitivity is largest. The two readings are far apart and the observer discounts the difference anyway, which is the whole of what the default is.
One reflectance, two illuminants, two colours. A reflectance peaking near 610 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.
Fig. 6 And one in the orange, where tungsten has most to offer. The stimulus moves by a different amount for every surface in the room, and one gain is what the observer answers all of them with.

This is colour constancy, and it is not a curiosity. It is the normal operation of the system, and the effects usually presented as illusions are cases where it produces an answer the viewer can recognise as wrong.

The problem being solved

The light arriving from a surface is the product of two things: the surface’s reflectance and the illumination falling on it. The eye measures the product and needs the first factor.

stimulus(λ)=reflectance(λ)×illuminant(λ)\text{stimulus}(\lambda) = \text{reflectance}(\lambda) \times \text{illuminant}(\lambda)

One measurement, two unknowns. The problem is underdetermined and no amount of cleverness makes it otherwise — any stimulus is consistent with a bright surface under dim light or a dark surface under bright light, and with every intermediate.

There is a good reason to want the reflectance rather than the product. Reflectance is a property of the object; it identifies materials, it stays constant as conditions change, it is what makes an object recognisable across a day. Illumination varies by orders of magnitude and carries little useful information about anything.

So the visual system estimates the illuminant and divides it out. Since the problem is underdetermined, the estimate rests on assumptions, and the assumptions are what everything else here follows from.

The assumptions

Several are used, and none of them is guaranteed.

The brightest thing is white. A useful heuristic: scenes usually contain something near-white, and the brightest surface is a reasonable estimate of the illuminant’s colour. This is the anchoring principle that also fixes the lightness scale.

The average is grey. The grey-world assumption: averaged over a whole scene, reflectances tend toward neutral, so any tint in the average is attributable to the light. It works well for varied scenes and fails badly for a photograph of a green field.

Specular highlights carry the illuminant. A glossy surface reflects some light without spectral modification, so a highlight is a direct sample of the illumination. This is a genuinely good cue and the visual system appears to use it.

Illumination varies smoothly, surfaces do not. Sharp changes are attributed to surfaces, gradual ones to lighting. This is what the Cornsweet effect exploits.

Each is a sound generalisation about the world and each can be violated, deliberately or by accident. Everything called a colour illusion is a stimulus that violates one.

The arithmetic

The computational core is simpler than the phenomenon suggests, and is called von Kries adaptation: scale each cone response by a factor set by the illuminant.

L=LLw,M=MMw,S=SSwL' = \frac{L}{L_w},\qquad M' = \frac{M}{M_w},\qquad S' = \frac{S}{S_w}

where LwL_w, MwM_w, SwS_w are the responses to the estimated white. Each receptor class independently normalises to what it is receiving on average, which is exactly what a gain control does, and there is physiological evidence for adaptation at the receptor level.

Modern chromatic adaptation transforms — Bradford, CAT02, CAT16 — are refinements of this. Each converts to a cone-like space chosen to make the scaling work better, applies the diagonal scaling, and converts back. This site uses Bradford wherever a white point has to be changed.

The important structural point is that adaptation is a diagonal transformation in a well-chosen basis. It is not a complicated operation, and the difficulty in colour constancy is not the correction — it is estimating the illuminant to correct by, which is the underdetermined part.

Where it fails

Three familiar failures, all of them assumption violations.

Photographs. A camera does not adapt; it records the product. Automatic white balance is an attempt to estimate the illuminant using the same heuristics the visual system uses, and it fails in the same situations — a scene dominated by one colour defeats the grey-world assumption, which is why photographs of sunsets and forests are so often corrected in the wrong direction.

Mixed lighting. A room with daylight from a window and tungsten from a lamp has two illuminants. The system adapts to something in between, and surfaces in each region look tinted in opposite directions. Neither can be right, and the effect is visible to anyone who looks for it at dusk.

Narrow-band sources. Sodium street lighting emits essentially one wavelength. There is nothing to estimate and nothing to divide out — all surfaces return a scaled version of the same spectrum, so colour vision fails almost entirely. The scene appears in shades of yellow, and no amount of adaptation recovers anything, because the information is not present in the stimulus.

That last case is the clearest demonstration that constancy is an inference and not a correction. Where the light carries no information about the surfaces, no processing recovers the surfaces.

The dress

The photograph that circulated in 2015, seen by some as white and gold and by others as blue and black, is the best-known instance of the estimation genuinely being ambiguous.

The image was poorly exposed and cropped so that almost every illuminant cue was absent — no clear white reference, no reliable highlight, an uncertain background. Under those conditions the assumptions do not converge, and different viewers settled on different illuminant estimates. Assume a bluish daylight illuminant and divide it out, and the fabric is white and gold; assume a warm artificial one, and it is blue and black.

Both processes were operating exactly as designed. The stimulus was constructed — accidentally — to leave the problem underdetermined in a way that ordinary scenes do not, and the disagreement was between priors rather than between visual systems.

What constancy costs

The mechanism has a price, and the price is that a colour has no fixed appearance.

This is why a specification of a colour value does not specify an appearance, why a paint chip is misleading in the shop, and why a colour chosen in one context has to be rechecked in another. The visual system is not reporting the stimulus; it is reporting its best estimate of the surface, and the estimate depends on everything else in view.

It is also why colorimetry stops where it does. Predicting a match under identical conditions requires no model of the estimation. Predicting appearance requires knowing what the system has adapted to, which is why appearance models take viewing conditions as an argument.

The failure cases, drawn

Each of the assumptions can be violated, and the violations are what the illusion literature is made of.

The connecting bar is the clearest demonstration that constancy is an inference rather than a filter. Nothing about the two squares changed; what changed is the evidence available about how the scene is organised, and the inference changed with it.

Constancy is not perfect, and the imperfection is measurable

It would be easy to read the above as claiming the visual system fully discounts the illuminant. It does not.

Measured constancy indices — how much of the illuminant change is compensated — typically fall between 60 and 90 per cent under good conditions, and lower when cues are sparse. Surfaces do look somewhat warmer under tungsten; the correction is substantial and partial.

This is worth stating because the two extreme positions are both wrong. Colour is not simply a property of the surface, recovered perfectly; nor is it simply the stimulus, with no correction at all. It is an estimate, it is usually good, and it degrades gracefully as the evidence gets worse — which is what one would expect of an inference under uncertainty.

The degradation is why photographs of unusual lighting look wrong in a way the scene did not, and why two people can settle on different estimates when the cues are ambiguous enough.

What was computed here

The chromatic adaptation transforms are implemented in this site’s colour library — Bradford and CAT16 — and are used wherever a white point changes, in particular when converting between D50 and D65 based spaces.

They are checked by round trip: adapting from D65 to D50 and back must return the starting XYZ, and it does to better than 101010^{-10}. That catches an inverted matrix, which otherwise produces a plausible warm or cool shift that looks like a deliberate choice.

What is not computed on this page is any constancy prediction. The figures show stimuli, and the essay describes what the visual system does with them. Modelling the estimation would require an appearance model, which is the next phase’s work rather than something the foundation machinery can support — and claiming otherwise would be exactly the overreach this site is built to avoid.

The adaptation transforms are checked by round trip: adapting D65 to D50 and back must return the starting XYZ, and does to better than 1e-10. That catches an inverted matrix, which otherwise produces a plausible warm or cool shift indistinguishable from a deliberate choice.

What a camera does instead

Comparing the two systems is the clearest way to see what constancy is doing, because a camera has the same problem and no visual system.

A camera records the product of reflectance and illumination and nothing else. Automatic white balance estimates the illuminant using the same heuristics — grey world, brightest-is-white, sometimes highlight detection — and divides it out, which is von Kries adaptation implemented in software.

It fails in the same situations, and for the same reasons. A photograph dominated by one colour defeats the grey-world assumption, so pictures of sunsets are corrected toward neutral and lose the warmth that was the subject. Mixed lighting produces a compromise wrong in both regions.

The difference is that a camera’s failure is visible afterwards and the visual system’s is not, because there is no unadapted view to compare against.

The limit case

Low-pressure sodium lighting is the cleanest demonstration that constancy is an inference rather than a correction.

The lamp emits essentially one wavelength. Every surface returns a scaled copy of the same spectrum, differing only in how much. There is no chromatic information in the reflected light at all, so there is nothing to estimate and nothing to divide out, and colour vision fails almost completely — the scene appears in shades of one hue.

No processing recovers the surfaces, because the information is not present. This is worth holding onto as the boundary condition: constancy works by exploiting structure in the light, and where the light has no structure it has nothing to work with.

How good the estimate is

Constancy indices — the fraction of an illuminant change that gets compensated — typically fall between 60 and 90 per cent under good conditions and lower when cues are sparse.

That figure disposes of both extreme positions. Colour is not simply a property of the surface recovered perfectly, and it is not simply the stimulus with no correction. It is an estimate, usually good, degrading gracefully as evidence worsens — which is what an inference under uncertainty should do.

It is also worth noticing what constancy implies about the word colour. If the visual system’s output is an estimate of reflectance rather than a report of the stimulus, then everyday colour language refers to the estimate — the property of the object — and colorimetry refers to the stimulus. Two different quantities, both reasonably called colour, and most disputes about whether two things are the same colour turn out to be disputes about which one is meant.

A narrower reflectance is the harder case for any discounting, because a narrow band gives the lamp fewer wavelengths to be wrong about.

One reflectance, two illuminants, two colours. A reflectance peaking near 500 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.
Fig. 7 A reflectance peaking near 500 nm with half the usual width, under D65 and under A. The object has not changed and the light has, which is the whole of what it means for colour to be a property of a pair.

How incomplete the discounting is, measured

Constancy is described above as good and imperfect. An appearance model turns the second half into a number, and the number is more interesting than the description.

CIECAM16 carries a degree of adaptation, D, computed from the adapting luminance and the surround. At an ordinary indoor level it comes out at 0.94 — not 1. The remaining 6% is the illuminant that does not get discounted, and its consequence is that a perfectly neutral card seen under a coloured light does not come out neutral in the model. It keeps a trace of the light, which is exactly what everybody’s experience of a tungsten-lit room reports: the room does not look daylit, it looks warm, and the warmth does not go away however long anyone sits in it.

The part worth pausing on is that the residual does not follow how far the white sits from equal energy. D50 is further from equal energy in chromaticity than D65 is, and leaves a smaller residual. The reason is in the second stage rather than the first: almost all of D50’s imbalance is in the channel the opponent recoding nearly discards, since the chromatic axes weight the third cone signal at a ninth and an eleventh of the others. An imbalance in a channel that is barely transmitted barely shows.

That is a fact about how the visual system compresses its own signal, arriving at a question about lamps by an entirely different route, and it is the kind of connection the two-stage structure produces constantly once the stages are computed rather than described.

A very broad reflectance is the opposite case, and the two together bracket what the width of a surface’s own spectrum is worth.

One reflectance, two illuminants, two colours. A reflectance peaking near 560 nm, and the colours it produces under D65 and A. The object has not changed. The light has, and colour is a property of the pair.
Fig. 8 The same construction at 560 nm and nearly twice the usual width. A broad surface samples more of the lamp, so it moves less between the two lights — the same mechanism that makes a broad primary agree between observers.

What the model’s own index is

Sixty to ninety per cent is quoted from the literature, and the same quantity can be computed from this collection’s own machinery.

For each change of light, the index is one minus the residual divided by the whole change — the share of the stimulus difference the gain removes. At the model’s own degree of adaptation of 0.941: D65 to D50, 92.1 per cent; to D40, 91.7; to D100, 90.8; to a blackbody at the same temperature, 87.3; to tungsten, 91.0; to a triphosphor tube, 66.8; to a halophosphate tube, 91.1; to a white LED, 87.7; to three narrow emitters, 87.9; one bounce off a green wall, 91.2; two bounces, 88.2; a red wall, 93.9.

The mean is 88.3 per cent and the range is 66.8 to 93.9. Only five of the twelve fall inside the quoted sixty-to-ninety band; the other seven sit above it.

So the model is more constant than people are. Its diagonal removes more of an illuminant change than psychophysics reports observers removing, on almost every row. That is the expected direction and it is worth stating, because it names what the model leaves out: a real observer has to estimate the illuminant from the scene, and estimates it imperfectly, where the model is handed both whites and told to divide.

Two details make the comparison useful rather than embarrassing. The one row landing in the middle of the quoted band is the triphosphor tube, at 66.8 — which is the case a real observer also finds hardest, so the model and the literature agree about which changes are difficult while disagreeing about the level.

And the incompleteness costs almost nothing to the index. Going from complete adaptation to D = 0.941 moves the mean from 89.8 to 88.3, a point and a half. The six per cent of illuminant left undiscounted is a large addition to a small residual and a small subtraction from a large index, which is one piece of arithmetic seen from its two ends.

Comparing a tungsten lamp against equal energy rather than daylight removes the daylight reconstruction from the argument entirely.

One reflectance, two illuminants, two colours. A reflectance peaking near 520 nm, and the colours it produces under A and E. The object has not changed. The light has, and colour is a property of the pair.
Fig. 9 A 520 nm reflectance under illuminant A and under equal-energy E. Neither light is a daylight, so what remains is the surface’s own spectrum multiplied by two shapes that have nothing in common.

What the pictures cannot show

The central limitation is unavoidable and worth being blunt about. Demonstrating adaptation requires controlling what the observer is adapted to, and this page controls none of it. A reader adapted to their room is looking at a screen inside that room, so every swatch here is already being interpreted against a state the figure knows nothing about.

The two swatches in the hero figure are a particular case. They are the stimuli a camera would record under two illuminants, presented side by side on one screen under a third. That is not the comparison an observer makes when walking from one room to another, and it exaggerates the difference substantially — side-by-side presentation defeats adaptation, which is exactly the point of the figure and also its limitation.

Nothing here can show what the paper looks like to somebody in the tungsten room, because showing it would require putting the reader in that room.

Who found it, and when

Monge described constancy in 1789. Helmholtz gave the standard formulation in the nineteenth century, calling it unconscious inference — the visual system draws conclusions from assumptions without any of it being available to introspection, which remains a fair description.

Johannes von Kries proposed the independent receptor-scaling rule around 1902, and it is still the core of every adaptation transform in use.

Edwin Land’s experiments in the 1950s and 1960s were the most striking demonstrations: scenes lit by two or three narrow-band sources in which surfaces retained their apparent colours far better than the stimulus alone could explain. His retinex theory has not survived as a mechanism, and the demonstrations remain the clearest evidence that constancy is computed from spatial relationships rather than from local measurements.

Where this goes next

The local version of the same machinery is these two patches are identical and brightness is inferred from edges. The physics being discounted is the illuminant is half the answer. And the boundary this places on colorimetry is matching is not appearance.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 28 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AdaptationChromatic adaptationColour constancyIlluminantIlluminant estimationReflectanceThe von Kries transformWhite balance