What a camera does

A camera balances in another basis

White balance is a per-channel gain on raw values, which makes it a von Kries adaptation in whatever axes the filter dyes happen to give. Those axes are not a property of the dyes alone — they move with the light, by up to seventeen degrees across the adaptation census — and the sensor for which they would not move is the one that adapts worst of all.

Assumes Fitted to an eye nobody has and What no adaptation can remove.

Every camera does the same thing about a change of light that an eye does: it multiplies three signals by three numbers. In the eye that is von Kries adaptation; in the camera it is white balance, and it is applied to the raw values before anything else happens.

The census of light changes established that such a gain removes only the part of a change that is diagonal in whatever axes the gain is applied along, so the interesting question about a camera is what its axes are. They are not chosen by anybody. They are three dye transmittances multiplied by a silicon quantum efficiency, arrived at by a supply chain, and they turn out to have a property no eye’s axes have.

The basis a camera balances in is a different basis for every light. A camera's white balance is a per-channel gain on raw values, which is a von Kries adaptation in whatever basis the filter dyes give it. That basis is not a property of the dyes alone: it is the dyes and the light in the room, and it moves when the light does. Each bar is how far the basis has turned, in degrees, from where it sits under D65. A sensor satisfying the Luther condition would have a bar of exactly zero on every row, because for such a sensor the light cancels — which is the one property nobody buys a sensor for.
Fig. 1 How far a silicon sensor’s adaptation basis turns, in degrees, from where it sits under D65, as the light it is balancing changes. A sensor satisfying the Luther condition would have a bar of exactly zero on every row.

What that basis is, and how far it sits from the ones anybody publishes, is the same question asked two more ways.

The three axes, as the coefficients on X, Y and Z that make them. Each row is a basis, and each block of three bars is one of its axes: how much of X, of Y and of Z that direction takes, with bars below the line for negative coefficients. The top row is not a published transform but the one in which a change from D65 to D50 is exactly a gain, taken from the eigenvectors of the change matrix. The four beneath it were fitted to corresponding-colour experiments, and none of them lands on it — the closest is Bradford, whose furthest axis is 37 degrees away.
Fig. 2 The published transforms’ axes as coefficients on X, Y and Z. A camera’s are not on this chart because they are not fixed: they are three dye transmittances multiplied by whatever light is in the room.
What is left of each change after the gain, as a matrix. The residual operator: the change of light written in the adaptation basis, with the gain the observer applies divided out. A change that was a gain leaves the identity — ones down the diagonal and nothing anywhere else — so every mark off the diagonal is something no adaptation can remove. The white is a fixed point of all four of these by construction, which is why an adaptation transform that gets the white right has proved nothing.
Fig. 3 What each change of light leaves after its gain, drawn as a matrix rather than as a number. A basis that suited the change would leave the identity, and the off-diagonal is what a camera’s moving basis adds to what is already there.

Two more views of the same machinery place the camera’s moving basis among the fixed ones the field publishes, which is the comparison the essay has been making in words.

Every published adaptation transform, and one computed from daylight, on every change. What each basis leaves an adapted observer with, row by row. Darker is worse. The last column is not a published transform: it is the basis in which a change from D65 to D50 is exactly diagonal, computed in closed form from the two spectra with nothing fitted. It is far the best on the daylight rows and it is beaten on the discharge lamps, which is the trade the published transforms are sitting in — they were fitted to data containing both kinds of light and are therefore optimal for neither. Over the census as a whole the winner is Bradford at ΔE00 1.14.
Fig. 4 Every basis on every change of light. A camera’s is not a column here because it is not fixed; the columns that are show how much of the spread is decided by axes rather than by the light.

What the axes would have been if anybody had ever chosen them for this particular job is the other end of the same comparison, and arriving at them is a search rather than an adoption of somebody else’s table.

Two routes to the same three axes. A search that minimises a mean colour difference by moving nine numbers, and an eigen-decomposition that never evaluates a colour difference, applied to the same three changes of daylight. They agree on all three axes to within 0.41 degrees and reach the same residual to three figures. Getting there took a restart: from a single simplex the search returned the same answer at nine hundred, three thousand and nine thousand iterations, which reads exactly like convergence, and that answer was twice the closed form's residual.
Fig. 5 And a basis found by searching rather than by adopting. It is what the axes would be if anybody had chosen them for this job, which is the comparison a camera’s dyes have never been put through.

The claim

A camera’s adaptation basis is a property of its dyes and of the light in the room, so it is a different basis for every scene — and the sensor for which it would be constant is the one that balances worst.

  • The map from tristimulus values to raw values is a 3×3 matrix, exactly, on a three-dimensional set of surfaces — the same argument that makes a change of light a matrix.
  • The light appears in that matrix twice and does not cancel. Under a three-emitter source the basis has turned 16.6 degrees from where it sat under D65; under tungsten, 12.6; under two bounces off a green wall, 13.5.
  • For a Luther-condition sensor it cancels exactly, and the drift is 1.5 × 10⁻⁶ degrees across the whole census.
  • And that sensor is worse at white balancing than a real one — ΔE00 2.52 against 1.70, averaged over the census — because satisfying the condition means its channels are the matching functions, and a per-channel gain on the matching functions is the transform this site calls the oldest mistake still shipping.
  • The eye beats both, at 1.41.

The camera has a basis, and here it is

The argument is the census’s, run with a different set of three functions at the end.

A surface is three coefficients on a three-dimensional reflectance basis. Its tristimulus values under a light are a 3×3 matrix times those coefficients. Its raw values under the same light are a different 3×3 times the same coefficients. So

raw = [S diag(E) R] [A diag(E) R]⁻¹ · XYZ

with S the sensor’s three sensitivities and A the matching functions. A camera has axes in exactly the sense an eye has them, and white balance is a diagonal in those axes.

The light is in the expression twice

Look at where E appears. It is inside both brackets, and unless S is a linear combination of A it does not cancel.

That is the whole finding. The eye’s axes are three pigments and do not depend on what is being looked at. A camera’s axes are three dyes and the illuminant, so the basis its white balance is diagonal in is one basis under daylight and a different basis under a tube, and no setting in the camera reflects that.

A silicon sensor's best possible impersonation of the standard observerThe 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 31.7 per cent of the matching functions' own magnitude, worst at 440 nm. Colour reproduction is exact if and only if this is zero.0x̄ ȳ z̄, behind — what a colorimeter needsthe sensor's best linear fit to the matching functionsthe fit goes negative, and has toworst at 440 nmwhat is left over — residual 31.7% overall400450500550600650700750wavelength / nmmodelled silicon sensorLuther–Ives 1927, the fit and its residual
Fig. 6 How far a silicon sensor is from being a linear combination of the matching functions. It is the failure of this fit that puts the light into the expression above — the condition is precisely the condition for the two occurrences of E to cancel.

The drift is not small. Balancing under a three-emitter source is being done in axes 16.6 degrees away from the ones used under daylight. To put that in scale, the five published adaptation transforms for the eye differ from each other by between 3 and 58 degrees, and the whole of the previous essay’s argument is about differences of that order.

The sensor that cannot move is the worst one

The obvious conclusion from the paragraph above is that a camera should be built to satisfy the Luther condition, and then its basis would be fixed and white balance would behave properly. The measurement says otherwise, and the reason is worth following.

Averaged over the census, a silicon sensor’s white balance leaves ΔE00 1.70. A sensor satisfying the condition exactly leaves 2.52. It is half again as bad.

Not at colour: at adapting. This site’s colorimetric sensor mixes the matching functions in proportions close to the identity, so applying a per-channel gain to its raw channels is nearly scaling X, Y and Z directly, which is the worst basis in the entire comparison and comes last on every census row that is not a discharge lamp.

That is a fact about this particular sensor and it was read here, for a long time, as a fact about the condition. It is not: the condition asks for some linear combination and says nothing about which, and which one is the whole of a colorimetric camera’s adaptation basis. Four sensors that all satisfy it exactly leave between 0.97 and 2.46, and the best of them is the best figure any basis reaches.

There is no design with both properties. A sensor whose channels are the matching functions has a fixed basis and a bad one; a sensor whose channels are sharpened and cone-like has a good basis and fails the condition, which means it has metamers of its own that no matrix can repair. The trade is not an engineering compromise anybody chose — it is two properties that cannot both hold.

Which lights move it, and by how much

The drift is largest for exactly the lights that are hardest for an eye.

Three narrow emitters move the basis 16.6 degrees. Two bounces off a green wall move it 13.5, tungsten 12.6, a red wall 11.3. A change from D65 to D50 moves it 3.0, and daylight to a blackbody at the same temperature moves it 2.3.

So a camera’s white balance is worst exactly where an eye’s adaptation is worst, and for a related but not identical reason. The eye’s problem is that narrow structure is not diagonal in any fixed basis; the camera has that problem too, and on top of it the basis it is using has moved.

What the drift actually predicts

The claim that the two effects add is a claim about a correlation, and the census is large enough to check it.

The lights a camera finds hard are the lights an eye finds hard. Ranking the twelve census rows by the residual each device is left with, the camera’s order and the eye’s agree at a Spearman coefficient of 0.944. So the moving basis does not change which changes of light are difficult; whatever makes a change hard to remove with any diagonal makes it hard for both.

What the basis changes is the surcharge. Ranking the rows by the drift angle and by the ratio of the camera’s residual to the eye’s gives 0.741, and against the plain difference between them, 0.664. The turning of the axes is therefore not a separate phenomenon that happens to coincide with the hard rows; it is a penalty on top of them, and it is predictable from an angle computed out of the sensitivities and the lamp before any surface is involved. The drift does also track the eye’s own difficulty, at 0.601 — the lights that turn a camera’s basis furthest are broadly the lights nothing adapts to well — but it tracks the camera’s surcharge more closely than it tracks that, which is what says the two are separable at all.

The camera wins on daylight and only on daylight. All three daylight rows go to it — D65 to D50 at 0.419 against the eye’s 0.458, to D40 at 0.907 against 0.980, to D100 at 0.457 against 0.508 — a margin of eight to ten per cent each time, and it loses the other nine rows. That is a cleaner statement than the mean allows: the sensor’s dyes are a better adaptation basis than CAT16 for the one family of changes they were never designed for, and worse for everything else.

And the nine losses are not one size. They run from 7.5 per cent behind on the three-emitter tube to 40.6 on a bounce off a red wall, with a median of 25. A summary that quotes the top half of that range makes the failure sound more uniform than it is.

Where the colorimetric sensor actually loses

Its mean of 2.52 is a summary of a row-by-row result that has a shape, and the shape is the argument rather than the mean.

It is worse than the silicon sensor on all eight rows that are not a discharge lamp — by a factor of four on the shift to D50, 1.73 against 0.419, and by more than two on the shift to D100. And it splits the four discharge rows evenly, winning the three-emitter tube at 1.55 against 2.50 and the narrow-band trio at 1.11 against 1.38, and losing the halophosphate tube and the white LED.

That pattern is what the argument predicts rather than a curiosity. A gain applied to something close to X, Y and Z is a bad adaptation whenever the change of light is smooth, because a smooth change is nearly diagonal in a cone-like basis and nothing like diagonal in a tristimulus one. When the change is a set of lines, nothing is diagonal in any basis at all, so a bad basis costs nothing to have and a fixed one stops being a liability. Both of the colorimetric sensor’s wins are on discharge sources, and it loses every smooth change on the list.

What a profile does about it, and what it cannot

A camera profile is a matrix fitted to a set of patches under a stated illuminant, and it is the industry’s answer to all of this. It is a fit, and it is a good one within the light it was fitted for.

What it cannot do is fix the basis. The profile is applied after white balance, so the gain has already been applied in the moved axes and the profile is repairing the result. A matrix applied after a diagonal is not the same as a diagonal in different axes, and the residual left over is what this essay measures.

The order could be different. A camera could convert to a fixed cone-like basis first and apply the gain there — that is what several raw converters effectively do when they adapt in a working space rather than in raw — and the measurement here says it should help by about the difference between the two device rows, half a unit of ΔE00 on average.

What it costs a photograph

The numbers so far are means over a census, which is the right way to compare mechanisms and the wrong way to say what a photographer would see. The per-row figures are more useful.

Under a change from D65 to D50 — an ordinary difference between two daylight conditions — a camera’s white balance leaves ΔE00 0.42 and an eye leaves 0.46, so the camera is very slightly better: its sharpened dye channels are closer to a good adaptation basis than CAT16 is, on that change. Under tungsten it leaves 2.04 against the eye’s 1.64. Under two bounces off a green wall, 4.62 against 3.37.

So the shape of the result is: a camera balances daylight about as well as a person adapts to it, and falls behind on everything else by between eight and forty-one per cent. That is a defensible place for a device to be, and it is not what the marketing of white balance implies — the control is presented as a correction that either succeeds or is set wrongly, and the residual it leaves when set perfectly is not a quantity anybody reports.

The largest gap is worth its own sentence, because it is the one a person would actually meet. Photographing a subject in a coloured room, a camera with a perfectly chosen white balance is left with about a third more error than the person standing next to it. The room is the case where a camera’s basis has turned furthest and where the change is sharpest, and the two effects add.

What a raw converter could do about it

The measurement points at a change that is a change to a pipeline rather than to a sensor, which makes it unusually cheap to test.

White balance is applied to raw values because that is where the multipliers are convenient: the raw channels are what the sensor produces, the gains are three numbers, and everything downstream expects balanced raw. Nothing about the arithmetic requires that ordering. A converter could instead map raw to a fixed cone-like space with the profile matrix, apply the gain there, and carry on — which is what happens implicitly whenever adaptation is done in a working space rather than at capture.

The measurement says what that ordering is worth. Applying the gain in a fixed cone basis rather than in the sensor’s own is the difference between the camera row and the eye row of the device comparison: ΔE00 1.70 against 1.41, averaged over the census, and rather more on the rows where the sensor’s basis has turned furthest. On a bounce off a coloured wall it is 4.62 against 3.37.

The catch is that mapping raw to a fixed space needs a matrix, and the matrix is fitted under one illuminant and is wrong under others. So the change trades a gain applied in the wrong basis for a gain applied in the right basis after an approximate mapping, and which is better is not obvious in advance — it depends on how much the profile matrix costs relative to how far the basis has turned.

That is a measurable question with a definite answer for any given sensor, and it is the kind of question a converter’s authors are well placed to settle and nobody appears to have asked. The reason is probably that the two halves belong to different literatures: the matrix is colorimetry and the gain is adaptation, and the ordering between them is nobody’s subject.

Who found it, and when

Luther’s condition is from 1927 and Ives had the same idea before him; the constraint is that a sensor’s sensitivities be a non-singular linear combination of the matching functions, and no practical camera has ever satisfied it.

Von Kries’s diagonal is 1902. White balance as a per-channel gain on raw values is as old as digital colour photography, and its justification in the literature is usually pragmatic — it is what the sensor gives, it is cheap, and it works well enough. Correcting further costs noise, which is the constraint that has kept the operation where it is.

The observation that the two intersect badly does not seem to be standard. The pieces are all well known separately: that a sensor is not colorimetric, that white balance is von Kries, that von Kries is basis-dependent. Putting them together says that a camera is doing an operation whose quality depends on axes that are themselves a function of the scene, and that a sensor built to fix the first problem makes the second one worse.

What was computed, and how

The camera basis is [S diag(E) R][A diag(E) R]⁻¹, computed exactly for each context in the census on the site’s three-dimensional reflectance family. The drift is the worst angle between the rows of that matrix under D65 and under the other light, matched greedily by closest angle rather than by position — two bases can be the same basis with its rows in a different order.

The device comparison applies each device’s own basis to each census row and reports the mean CIEDE2000 left. The colorimetric sensor is this site’s colorimetricSensor, whose channels are the matching functions exactly; its drift is asserted to be under 10⁻⁴ degrees, and it comes out at 1.5 × 10⁻⁶.

The press row in that figure is not a residual at all: a press has no mechanism, so its number is the whole change. That row is asserted to equal the census’s unadapted change exactly, because a table in which four numbers mean “what is left after something” and one means “what happens when nothing is done” is a table that has to say so.

Where it stops

The sensor here is a constructed one: a silicon quantum efficiency curve and three dye bands with a stated infrared leak, not a measurement of any particular camera. The size of the drift is a property of that construction. The direction is not — any sensor failing the Luther condition has a moving basis, and every real sensor fails it.

The comparison assumes each device applies its gain from the ratio of the whites, which is what an idealised white balance does and is not quite what a camera does. Real raw converters estimate the illuminant, apply a gain, and then apply a profile chosen from a small set by estimated colour temperature; the estimate is its own source of error and is not modelled here.

And the whole calculation is about a change of light. A camera photographing one scene under one light is doing something this essay says nothing about; the basis question only arises when the same camera has to be right about two lights, which is what a white balance control is for. A single-illuminant capture has a different set of problems and none of them is this one.

Where the ladder goes next

If a device’s basis is whatever its channels happen to be, then every device on this site has one and none of them chose it. A display’s is its primaries; a press has none at all. The comparison across all four is the essay that follows this one, and its finding is that the ordering by how well each adapts is not the ordering by how good each is at colour.

The other direction is the one the profile section opens. If applying the gain in a fixed cone-like basis rather than in raw would help by half a unit, that is a change to a pipeline rather than to a sensor, and it is testable against every camera that already exists.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 12 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AdaptationAssertionCamera rawChromatic adaptationColour filter arrayColour matrixIlluminantLuther conditionSiliconSpectral sensitivityThe von Kries transformWhite balance