A camera balances in another basis
Assumes Fitted to an eye nobody has and What no adaptation can remove.
Every camera does the same thing about a change of light that an eye does: it multiplies three signals by three numbers. In the eye that is von Kries adaptation; in the camera it is white balance, and it is applied to the raw values before anything else happens.
The census of light changes established that such a gain removes only the part of a change that is diagonal in whatever axes the gain is applied along, so the interesting question about a camera is what its axes are. They are not chosen by anybody. They are three dye transmittances multiplied by a silicon quantum efficiency, arrived at by a supply chain, and they turn out to have a property no eye’s axes have.
What that basis is, and how far it sits from the ones anybody publishes, is the same question asked two more ways.
Two more views of the same machinery place the camera’s moving basis among the fixed ones the field publishes, which is the comparison the essay has been making in words.
What the axes would have been if anybody had ever chosen them for this particular job is the other end of the same comparison, and arriving at them is a search rather than an adoption of somebody else’s table.
The claim
A camera’s adaptation basis is a property of its dyes and of the light in the room, so it is a different basis for every scene — and the sensor for which it would be constant is the one that balances worst.
- The map from tristimulus values to raw values is a 3×3 matrix, exactly, on a three-dimensional set of surfaces — the same argument that makes a change of light a matrix.
- The light appears in that matrix twice and does not cancel. Under a three-emitter source the basis has turned 16.6 degrees from where it sat under D65; under tungsten, 12.6; under two bounces off a green wall, 13.5.
- For a Luther-condition sensor it cancels exactly, and the drift is 1.5 × 10⁻⁶ degrees across the whole census.
- And that sensor is worse at white balancing than a real one — ΔE00 2.52 against 1.70, averaged over the census — because satisfying the condition means its channels are the matching functions, and a per-channel gain on the matching functions is the transform this site calls the oldest mistake still shipping.
- The eye beats both, at 1.41.
The camera has a basis, and here it is
The argument is the census’s, run with a different set of three functions at the end.
A surface is three coefficients on a three-dimensional reflectance basis. Its tristimulus values under a light are a 3×3 matrix times those coefficients. Its raw values under the same light are a different 3×3 times the same coefficients. So
raw = [S diag(E) R] [A diag(E) R]⁻¹ · XYZ
with S the sensor’s three sensitivities and A the matching functions. A camera has axes in exactly the sense an eye has them, and white balance is a diagonal in those axes.
The light is in the expression twice
Look at where E appears. It is inside both brackets, and unless S is a linear combination of A it does not cancel.
That is the whole finding. The eye’s axes are three pigments and do not depend on what is being looked at. A camera’s axes are three dyes and the illuminant, so the basis its white balance is diagonal in is one basis under daylight and a different basis under a tube, and no setting in the camera reflects that.
The drift is not small. Balancing under a three-emitter source is being done in axes 16.6 degrees away from the ones used under daylight. To put that in scale, the five published adaptation transforms for the eye differ from each other by between 3 and 58 degrees, and the whole of the previous essay’s argument is about differences of that order.
The sensor that cannot move is the worst one
The obvious conclusion from the paragraph above is that a camera should be built to satisfy the Luther condition, and then its basis would be fixed and white balance would behave properly. The measurement says otherwise, and the reason is worth following.
Averaged over the census, a silicon sensor’s white balance leaves ΔE00 1.70. A sensor satisfying the condition exactly leaves 2.52. It is half again as bad.
Not at colour: at adapting. This site’s colorimetric sensor mixes the matching functions in proportions close to the identity, so applying a per-channel gain to its raw channels is nearly scaling X, Y and Z directly, which is the worst basis in the entire comparison and comes last on every census row that is not a discharge lamp.
That is a fact about this particular sensor and it was read here, for a long time, as a fact about the condition. It is not: the condition asks for some linear combination and says nothing about which, and which one is the whole of a colorimetric camera’s adaptation basis. Four sensors that all satisfy it exactly leave between 0.97 and 2.46, and the best of them is the best figure any basis reaches.
There is no design with both properties. A sensor whose channels are the matching functions has a fixed basis and a bad one; a sensor whose channels are sharpened and cone-like has a good basis and fails the condition, which means it has metamers of its own that no matrix can repair. The trade is not an engineering compromise anybody chose — it is two properties that cannot both hold.
Which lights move it, and by how much
The drift is largest for exactly the lights that are hardest for an eye.
Three narrow emitters move the basis 16.6 degrees. Two bounces off a green wall move it 13.5, tungsten 12.6, a red wall 11.3. A change from D65 to D50 moves it 3.0, and daylight to a blackbody at the same temperature moves it 2.3.
So a camera’s white balance is worst exactly where an eye’s adaptation is worst, and for a related but not identical reason. The eye’s problem is that narrow structure is not diagonal in any fixed basis; the camera has that problem too, and on top of it the basis it is using has moved.
What the drift actually predicts
The claim that the two effects add is a claim about a correlation, and the census is large enough to check it.
The lights a camera finds hard are the lights an eye finds hard. Ranking the twelve census rows by the residual each device is left with, the camera’s order and the eye’s agree at a Spearman coefficient of 0.944. So the moving basis does not change which changes of light are difficult; whatever makes a change hard to remove with any diagonal makes it hard for both.
What the basis changes is the surcharge. Ranking the rows by the drift angle and by the ratio of the camera’s residual to the eye’s gives 0.741, and against the plain difference between them, 0.664. The turning of the axes is therefore not a separate phenomenon that happens to coincide with the hard rows; it is a penalty on top of them, and it is predictable from an angle computed out of the sensitivities and the lamp before any surface is involved. The drift does also track the eye’s own difficulty, at 0.601 — the lights that turn a camera’s basis furthest are broadly the lights nothing adapts to well — but it tracks the camera’s surcharge more closely than it tracks that, which is what says the two are separable at all.
The camera wins on daylight and only on daylight. All three daylight rows go to it — D65 to D50 at 0.419 against the eye’s 0.458, to D40 at 0.907 against 0.980, to D100 at 0.457 against 0.508 — a margin of eight to ten per cent each time, and it loses the other nine rows. That is a cleaner statement than the mean allows: the sensor’s dyes are a better adaptation basis than CAT16 for the one family of changes they were never designed for, and worse for everything else.
And the nine losses are not one size. They run from 7.5 per cent behind on the three-emitter tube to 40.6 on a bounce off a red wall, with a median of 25. A summary that quotes the top half of that range makes the failure sound more uniform than it is.
Where the colorimetric sensor actually loses
Its mean of 2.52 is a summary of a row-by-row result that has a shape, and the shape is the argument rather than the mean.
It is worse than the silicon sensor on all eight rows that are not a discharge lamp — by a factor of four on the shift to D50, 1.73 against 0.419, and by more than two on the shift to D100. And it splits the four discharge rows evenly, winning the three-emitter tube at 1.55 against 2.50 and the narrow-band trio at 1.11 against 1.38, and losing the halophosphate tube and the white LED.
That pattern is what the argument predicts rather than a curiosity. A gain applied to something close to X, Y and Z is a bad adaptation whenever the change of light is smooth, because a smooth change is nearly diagonal in a cone-like basis and nothing like diagonal in a tristimulus one. When the change is a set of lines, nothing is diagonal in any basis at all, so a bad basis costs nothing to have and a fixed one stops being a liability. Both of the colorimetric sensor’s wins are on discharge sources, and it loses every smooth change on the list.
What a profile does about it, and what it cannot
A camera profile is a matrix fitted to a set of patches under a stated illuminant, and it is the industry’s answer to all of this. It is a fit, and it is a good one within the light it was fitted for.
What it cannot do is fix the basis. The profile is applied after white balance, so the gain has already been applied in the moved axes and the profile is repairing the result. A matrix applied after a diagonal is not the same as a diagonal in different axes, and the residual left over is what this essay measures.
The order could be different. A camera could convert to a fixed cone-like basis first and apply the gain there — that is what several raw converters effectively do when they adapt in a working space rather than in raw — and the measurement here says it should help by about the difference between the two device rows, half a unit of ΔE00 on average.
What it costs a photograph
The numbers so far are means over a census, which is the right way to compare mechanisms and the wrong way to say what a photographer would see. The per-row figures are more useful.
Under a change from D65 to D50 — an ordinary difference between two daylight conditions — a camera’s white balance leaves ΔE00 0.42 and an eye leaves 0.46, so the camera is very slightly better: its sharpened dye channels are closer to a good adaptation basis than CAT16 is, on that change. Under tungsten it leaves 2.04 against the eye’s 1.64. Under two bounces off a green wall, 4.62 against 3.37.
So the shape of the result is: a camera balances daylight about as well as a person adapts to it, and falls behind on everything else by between eight and forty-one per cent. That is a defensible place for a device to be, and it is not what the marketing of white balance implies — the control is presented as a correction that either succeeds or is set wrongly, and the residual it leaves when set perfectly is not a quantity anybody reports.
The largest gap is worth its own sentence, because it is the one a person would actually meet. Photographing a subject in a coloured room, a camera with a perfectly chosen white balance is left with about a third more error than the person standing next to it. The room is the case where a camera’s basis has turned furthest and where the change is sharpest, and the two effects add.
What a raw converter could do about it
The measurement points at a change that is a change to a pipeline rather than to a sensor, which makes it unusually cheap to test.
White balance is applied to raw values because that is where the multipliers are convenient: the raw channels are what the sensor produces, the gains are three numbers, and everything downstream expects balanced raw. Nothing about the arithmetic requires that ordering. A converter could instead map raw to a fixed cone-like space with the profile matrix, apply the gain there, and carry on — which is what happens implicitly whenever adaptation is done in a working space rather than at capture.
The measurement says what that ordering is worth. Applying the gain in a fixed cone basis rather than in the sensor’s own is the difference between the camera row and the eye row of the device comparison: ΔE00 1.70 against 1.41, averaged over the census, and rather more on the rows where the sensor’s basis has turned furthest. On a bounce off a coloured wall it is 4.62 against 3.37.
The catch is that mapping raw to a fixed space needs a matrix, and the matrix is fitted under one illuminant and is wrong under others. So the change trades a gain applied in the wrong basis for a gain applied in the right basis after an approximate mapping, and which is better is not obvious in advance — it depends on how much the profile matrix costs relative to how far the basis has turned.
That is a measurable question with a definite answer for any given sensor, and it is the kind of question a converter’s authors are well placed to settle and nobody appears to have asked. The reason is probably that the two halves belong to different literatures: the matrix is colorimetry and the gain is adaptation, and the ordering between them is nobody’s subject.
Who found it, and when
Luther’s condition is from 1927 and Ives had the same idea before him; the constraint is that a sensor’s sensitivities be a non-singular linear combination of the matching functions, and no practical camera has ever satisfied it.
Von Kries’s diagonal is 1902. White balance as a per-channel gain on raw values is as old as digital colour photography, and its justification in the literature is usually pragmatic — it is what the sensor gives, it is cheap, and it works well enough. Correcting further costs noise, which is the constraint that has kept the operation where it is.
The observation that the two intersect badly does not seem to be standard. The pieces are all well known separately: that a sensor is not colorimetric, that white balance is von Kries, that von Kries is basis-dependent. Putting them together says that a camera is doing an operation whose quality depends on axes that are themselves a function of the scene, and that a sensor built to fix the first problem makes the second one worse.
What was computed, and how
The camera basis is [S diag(E) R][A diag(E) R]⁻¹, computed exactly for each context in the census on the site’s three-dimensional reflectance family. The drift is the worst angle between the rows of that matrix under D65 and under the other light, matched greedily by closest angle rather than by position — two bases can be the same basis with its rows in a different order.
The device comparison applies each device’s own basis to each census row and reports the mean CIEDE2000 left. The colorimetric sensor is this site’s colorimetricSensor, whose channels are the matching functions exactly; its drift is asserted to be under 10⁻⁴ degrees, and it comes out at 1.5 × 10⁻⁶.
The press row in that figure is not a residual at all: a press has no mechanism, so its number is the whole change. That row is asserted to equal the census’s unadapted change exactly, because a table in which four numbers mean “what is left after something” and one means “what happens when nothing is done” is a table that has to say so.
Where it stops
The sensor here is a constructed one: a silicon quantum efficiency curve and three dye bands with a stated infrared leak, not a measurement of any particular camera. The size of the drift is a property of that construction. The direction is not — any sensor failing the Luther condition has a moving basis, and every real sensor fails it.
The comparison assumes each device applies its gain from the ratio of the whites, which is what an idealised white balance does and is not quite what a camera does. Real raw converters estimate the illuminant, apply a gain, and then apply a profile chosen from a small set by estimated colour temperature; the estimate is its own source of error and is not modelled here.
And the whole calculation is about a change of light. A camera photographing one scene under one light is doing something this essay says nothing about; the basis question only arises when the same camera has to be right about two lights, which is what a white balance control is for. A single-illuminant capture has a different set of problems and none of them is this one.
Where the ladder goes next
If a device’s basis is whatever its channels happen to be, then every device on this site has one and none of them chose it. A display’s is its primaries; a press has none at all. The comparison across all four is the essay that follows this one, and its finding is that the ordering by how well each adapts is not the ordering by how good each is at colour.
The other direction is the one the profile section opens. If applying the gain in a fixed cone-like basis rather than in raw would help by half a unit, that is a change to a pipeline rather than to a sensor, and it is testable against every camera that already exists.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The corner of the frame has another filter camera raw · colour filter array · colour matrix · illuminant · silicon · spectral sensitivity · white balance
- A corner is corrected by one row camera raw · colour matrix · illuminant · spectral sensitivity · white balance
- A photograph is not a measurement camera raw · colour matrix · luther condition · spectral sensitivity · white balance
- Constancy is the default adaptation · chromatic adaptation · illuminant · the von kries transform · white balance
- One row for every lamp costs the lamps that lose least camera raw · colour matrix · illuminant · spectral sensitivity · white balance
- The filter that makes colour possible camera raw · colour filter array · illuminant · silicon · spectral sensitivity
What links here
The 8 essays that link to this one and share the most of its objects, of 12 that link here.
The objects this essay names
Each one links to every other essay that touches it.
AdaptationAssertionCamera rawChromatic adaptationColour filter arrayColour matrixIlluminantLuther conditionSiliconSpectral sensitivityThe von Kries transformWhite balance