What a camera does

Luther said when it would work

There is an exact condition under which a fixed three-by-three matrix converts camera raw to XYZ correctly for every spectrum in existence. It was stated in 1927, it is a theorem rather than a guideline, and no camera ever built satisfies it.

Assumes A camera is a fourth observer and Why colour is exactly three-dimensional.

Most engineering constraints are approximate. A tolerance can be tightened, a material improved, a process refined, and the thing that could not be done last decade becomes routine.

This one is not of that kind. It has an if-and-only-if in it, and no amount of refinement moves it.

A silicon sensor's best possible impersonation of the standard observerThe 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 31.7 per cent of the matching functions' own magnitude, worst at 440 nm. Colour reproduction is exact if and only if this is zero.0x̄ ȳ z̄, behind — what a colorimeter needsthe sensor's best linear fit to the matching functionsthe fit goes negative, and has toworst at 440 nmwhat is left over — residual 31.7% overall400450500550600650700750wavelength / nmmodelled silicon sensorLuther–Ives 1927, the fit and its residual
Fig. 1 The colour-matching functions in outline, the best linear combination of a silicon sensor’s three sensitivities laid over them, and what is left over underneath. The residual is 31.7 per cent of the matching functions’ own magnitude. Colour reproduction from this sensor is exact if and only if that lower curve is flat at zero. The handle changes which observer the condition is stated against — the sensor is fixed and the target is not.

The claim

A fixed 3×33 \times 3 matrix converts a camera’s raw response to XYZ correctly for every spectrum if and only if the camera’s spectral sensitivities are a non-singular linear combination of the colour-matching functions.

Necessary and sufficient. Not “approximately correct if approximately a linear combination” — though that is also true and is what every real camera relies on — but an exact equivalence between two exact statements.

Why sufficient

This direction is three lines of linear algebra and is worth doing, because the whole condition falls out of writing the two integrals as matrices.

Let AA be the 3×N3 \times N matrix of colour-matching functions on the wavelength grid and SS the 3×N3 \times N matrix of the camera’s sensitivities. For any spectrum Φ\Phi, the observer records AΦA\Phi and the camera records SΦS\Phi.

Suppose S=MAS = MA for some invertible 3×33 \times 3 matrix MM. Then

SΦ=MAΦS\Phi = MA\Phi

for every Φ\Phi, so M1SΦ=AΦM^{-1}S\Phi = A\Phi: applying M1M^{-1} to the raw triple gives XYZ exactly. Nothing whatever was assumed about Φ\Phi — not that it is smooth, not that it is a reflectance under a common illuminant, not that it resembles anything in a test chart.

That last clause is the part worth dwelling on. A camera satisfying the condition would be a colorimeter, correct on a laser, on a fluorescent tube’s mercury lines, on a metameric pair constructed to be maximally awkward, on anything at all.

Why necessary

The other direction is where the content is, and it is an argument about null spaces rather than about fitting.

Suppose SS is not a linear combination of the rows of AA. Then the row space of SS is not contained in the row space of AA, and by a dimension argument the null space of SS is not contained in the null space of AA: there exists a spectrum bb with Sb=0Sb = 0 and Ab0Ab \neq 0.

Take any spectrum Φ\Phi and consider Φ\Phi and Φ+b\Phi + b. The camera records SΦS\Phi for both — they are identical in the file. The observer records AΦA\Phi and AΦ+AbA\Phi + Ab, which differ. So two stimuli a person can distinguish produce one raw triple, and there is no function of that triple, linear or otherwise, that recovers two different answers from one input.

A merge is not invertible. That is the whole of the necessity argument, and it is why the failure cannot be repaired by a more clever matrix, a bigger lookup table, or a neural network: the information required is not in the file.

Two reflectances the camera records as identicalConstructed by projecting onto the null space of the sensor's own sensitivities, so the two raw triples agree to 0.0000 per cent. To the eye they are ΔE00 15.33 apart, which the swatches show.400450500550600650700wavelength / nmtwo reflectancesas the eye sees themRGBas the camera records themraw gap 0.0000%, ΔE00 15.33null space of the sensor
Fig. 2 That vector bb, constructed. Two reflectances whose raw triples agree to four decimal places and which a person sees as ΔE00 2.4 apart. The swatches are computed through the observer, so the difference in the plate is a difference the camera did not record.

What the residual is, and what it is not

The number quoted throughout this field is 31.7 per cent, and it is worth being exact about what was computed.

Fit the MM minimising MSA\lVert MS - A \rVert over the visible band — the best possible linear redescription of the sensor as an observer — and take the root-mean-square of what is left, relative to the root-mean-square of the matching functions themselves. That is a pure number, comparable between sensors, and it is zero if and only if the condition holds.

Only the visible band is fitted, and that is not a convenience. Outside it the matching functions are zero and the sensitivities need not be, so including the infrared would score a sensor on a region where the target says nothing — and would make a bare photodiode look like a bad observer rather than like something that is not an observer at all. The infrared is a separate problem with a separate component, and mixing the two makes both harder to see.

The control is what makes the number mean anything. Build a sensor whose sensitivities are the matching functions through a fixed mixing matrix, run the identical computation, and the residual comes out at 101610^{-16} — machine precision. A residual with no control is a number; a residual beside a control that reaches machine precision is a measurement.

A sensor built to satisfy the Luther condition, impersonating it exactly. The 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 0.0 per cent of the matching functions' own magnitude, worst at 530 nm. Colour reproduction is exact if and only if this is zero.
Fig. 3 The control. A sensor built as a linear combination of the matching functions, fitted the same way, with a residual flat at zero along the whole band. It cannot be manufactured, and that is what it is for.

Both of those are statements about one observer, and the condition names an observer as well as a sensor.

A silicon sensor's best possible impersonation of the standard observer. The 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 33.9 per cent of the matching functions' own magnitude, worst at 440 nm. Colour reproduction is exact if and only if this is zero.
Fig. 4 The same silicon sensor against the ten-degree observer instead. The residual is a different curve because the target has moved, so a sensor is not near or far from the condition in the abstract — it is near or far from a stated set of functions.

The construction transfers to the second observer as cleanly as it did to the first, which is what says the condition is about a relationship rather than about a sensor.

A sensor built to satisfy the Luther condition, impersonating it exactly. The 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 0.0 per cent of the matching functions' own magnitude, worst at 610 nm. Colour reproduction is exact if and only if this is zero.
Fig. 5 And the control rebuilt for the same observer, which satisfies the condition exactly again. The construction transfers and the failure does not, which is what says the condition is about a relationship rather than about a sensor.

Why it cannot be built

The reason is visible in the plate at the top of this essay and is more specific than “dyes are not ideal”.

The xˉ\bar{x} function has two lobes — a large one around 600 nanometres and a small one around 440. No single filter transmittance has that shape, because a transmittance is a product of absorptions and absorptions are single-peaked; a dye that passes red and violet while blocking green is a dye with two absorption bands and no absorption between them, which is chemically awkward and unstable.

A linear combination of three single-peaked curves can produce a two-lobed function, and the way it does so is by subtraction: the fit is 0.62sR0.31sG+0.62 s_R - 0.31 s_G + \ldots, with negative coefficients. That is fine in a matrix and impossible in a filter, since a filter with negative transmittance would have to emit light.

So the condition is satisfiable in principle — the matrix is allowed to be anything — and the sensor’s own curves must still span the right three-dimensional subspace of functions, and three physically realisable transmittances do not span the subspace the matching functions occupy.

The consequence that is not obvious

The condition is stated against a set of matching functions, and there is more than one set.

A camera fitted to the 1931 observer is not fitted to the 1964 one, and neither is fitted to any particular person — observers differ from one another by more than the two standards differ. So even a hypothetical camera that satisfied Luther’s condition exactly against 1931 would fail it against a real reader of the photograph.

This makes the whole business slightly stranger than it first appears. The target is not “the truth about colour” but “the average of seventeen people in 1927”, and a camera that matched that average perfectly would still disagree with the person looking at the print. The condition is a statement about agreement between two instruments, and one of the instruments is a committee.

How close is close enough

Since nothing satisfies the condition, the operative question is how much the residual costs, and the answer is not proportional in any simple way.

The best matrix this sensor admits, fitted and tested on the same twenty-four surfaces, leaves a worst case of ΔE00 2.85 and a mean well under one. That is the most favourable measurement it is possible to make of a camera, and it is the one usually published. The control sensor under the identical computation reaches 10710^{-7}, which is arithmetic noise.

So a 31.7 per cent residual in the functions becomes a sub-unit mean error in the colours, which sounds like a large reduction and is really a statement about how unrepresentative ordinary surfaces are. Reflectances of real materials are smooth — a few broad features across three hundred nanometres — and a smooth reflectance is exactly the kind of stimulus on which two smooth kernels most nearly agree. The residual lives in the high-frequency part of the function space, and the world mostly does not put anything there.

Mostly. The exceptions are the things that put energy in narrow bands — a saturated dye with a sharp edge, an interference pigment, a metameric pair built against the eye — and the next section separates those from the case they are usually confused with, which is a narrow-band lamp. The two behave in opposite directions, and only one of them is where the residual is spent.

Which narrow band it is

That narrow-band illumination makes the residual worse is the natural reading, and it is the wrong way round.

Fitting the best matrix for each illuminant separately, on the same surfaces, the sensor’s mean error is 0.758 under D65, 0.797 under D50 and 0.816 under illuminant A — and 0.575 under a white LED, 0.410 under a triphosphor tube and 0.224 under a three-emitter source. The three narrow-band lamps average 0.403 against the three smooth ones’ 0.790. The narrow lamps are half as bad, not worse. The control sensor sits at 10⁻¹⁰ throughout, so this is the residual and not the arithmetic.

The reason is that a narrow lamp is a projection. If a source emits in three bands, every stimulus reaching the sensor is a combination of what those three bands do, and a three-by-three matrix has exactly enough freedom to be right about three bands. The residual lives in the parts of the function space the lamp is not illuminating, and a lamp illuminating less of it asks less of the sensor.

What does cost is narrow structure in the sample. Fitting the same sensor to twelve Gaussian surfaces under D65 and narrowing them: at a feature width of 80 nanometres the mean error is 0.399; at 60, it is 1.263; at 40, 3.253; at 25, 5.497; and at 15, 6.010 — fifteen times the broad case. Below that it falls back a little, to 4.846 at six nanometres, as the surfaces approach monochromatic and stop having much shape left for a fit to miss.

So the sentence worth keeping is narrower than the obvious one. The residual is exercised by two stimuli that differ in a narrow band, not by a lamp that emits in one. A photograph of an LED sign is a photograph of a narrow source, and the sensor does comparatively well on it. A photograph of two dyes with sharp absorption edges, or of a pair constructed to be metameric for the eye, is where the thirty-one per cent gets spent.

The trend in illumination therefore runs the other way from the one it is tempting to assert. What has become common is narrow sources, and they help. What would make the condition bite harder is narrow reflectances — interference pigments, sharp-cut filters, structural colour — which have also become more common, in a much smaller way.

How far an actual sensor is from the condition is worth seeing as a residual rather than as a number, since the residual says where the failure is.

A silicon sensor's best possible impersonation of the standard observer. The 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 31.7 per cent of the matching functions' own magnitude, worst at 440 nm. Colour reproduction is exact if and only if this is zero.
Fig. 6 The 1931 matching functions in outline with the closest linear combination of a silicon sensor’s three sensitivities over them, and the leftover underneath. The residual is 31.7 per cent of the functions’ own magnitude and worst at 440 nm, which is where a real camera’s blue channel does not have the shape it would need.

What a fourth channel would and would not buy

The obvious response to a three-channel instrument failing to span the right subspace is to add channels, and cameras with four and six have been built.

It helps, and it does not satisfy the condition. Adding a channel gives a 4×N4 \times N sensitivity matrix and a 3×43 \times 4 conversion, which is more freedom and a better fit — but the condition asks that the row space of the sensor contain the row space of the matching functions, and adding a fourth realisable transmittance adds a fourth non-negative single-peaked function to the span. It moves the residual down; it does not make it zero, because the obstruction was never the number of channels.

What extra channels do buy is more interesting than accuracy. Four channels can distinguish stimuli that three merge, which means a four-channel camera has a smaller null space — so it can, for instance, tell two illuminants apart that a three-channel camera cannot, which is why illuminant estimation improves markedly with a fourth. It is a better spectrometer without being a better colorimeter, and those are different objectives.

Where the model stops

Non-singularity is a real requirement and is usually skipped over. If MM is singular the sensor’s three channels do not span three dimensions — two of them are the same measurement — and the camera has fewer than three effective channels. Real sensors are comfortably non-singular; the requirement bites in designs with four or more channels, where the extra channels are deliberately not independent.

The condition says nothing about noise, and a sensor can satisfy it and be useless. The matching functions go negative when expressed in terms of realisable primaries, and a sensor approximating them closely has channels that are differences of large numbers, so it satisfies Luther and has terrible signal-to-noise. That trade-off is an essay of its own and it is the practical reason nobody pursues the condition to its limit even where they could get closer.

And it is a statement about linearity throughout. A photodiode is linear to a fraction of a per cent over most of its range, which is much better than the eye manages, and it fails at both ends. Near saturation nothing above holds, which is where a hue turns.

The generalisation

The transferable statement is not about cameras and is one of the more useful things in this collection.

Two instruments that reduce the same input to kk numbers agree on every input if and only if one’s kernel is a non-singular linear redescription of the other’s — and otherwise they agree on a subspace and disagree outside it. There is no third case. Agreement is not a matter of degree at the level of the mechanism, though it is measured as one at the level of the output.

The corollary is the one worth carrying: when two instruments disagree, the useful question is not how to calibrate them together but which stimuli they disagree about. Calibration finds the best MM; it cannot remove the residual; and the residual is concentrated on particular inputs rather than spread evenly. For cameras those inputs are the spectrally narrow ones — saturated dyes, LEDs, fluorescent lamps, anything with structure finer than the sensitivities’ own width — which is why the failures a photographer notices are red flowers, stage lighting and screens rather than skin and sky.

This site has met the same shape once before from the other side. A colour rendering index is an average over a sample set, and the same lamps that make a camera fail are the ones that make the index misleading, for the same reason: a narrow spectrum is where two smooth kernels most disagree.

The same theorem, seen from the eye’s side

There is a version of this condition that is not about cameras at all, and noticing it makes the whole thing feel less like an engineering misfortune.

The colour-matching functions are themselves a linear combination of the cone fundamentals — that is what makes XYZ a legitimate space rather than an arbitrary one, and it is why a transform exists between them. In the vocabulary of this essay: the CIE observer satisfies the Luther condition with respect to the cones. The 1931 functions and the 1964 functions are two instruments that agree with the cone system up to an invertible matrix, and that is precisely why they are usable at all.

So the condition is not an unreasonable standard invented to embarrass cameras. It is the property every successful colorimetric system in the discipline has, stated generally, and a camera is the one common instrument that lacks it.

It also explains why the failure feels so unlike other measurement problems. A camera is not noisy, badly calibrated or poorly made; it is a different observer, in the same sense that a person with an unusual pigment complement is a different observer. Observer metamerism between two people and sensitivity metamerism between a camera and a person are the same phenomenon with the same algebra, and the only difference is that one of the observers can be redesigned and the other cannot.

Who noticed, and when

Herbert Ives stated the requirement in 1915 in the context of three-colour photographic separation, and Robert Luther gave the version usually cited in 1927. Neither had an electronic sensor to apply it to; both were describing the taking filters of a colour photographic process, where the same algebra holds and the same negative lobes are unobtainable.

The result was rediscovered several times through the century, which is common for conditions that are simple to state and easy to reach independently. It acquired a practical form in 1995 when the CIE standardised a sensitivity metamerism index, which is a different normalisation of the same residual, computed against a stated set of test colours and reported so that two cameras can be ranked. That standardisation is the interesting modern development: the condition stopped being a remark about an unattainable ideal and became a number in a specification.

What has not changed since 1927 is the exclusivity. Every camera scores above zero, every camera has surfaces it is measurably wrong about, and which surfaces those are is decided by what the manufacturer’s matrix was fitted to rather than by the condition itself.

Where the ladder goes next

Downward, this rung rests on a camera is a fourth observer for the objects it is about, and on why colour is exactly three-dimensional for the collapse the whole argument concerns.

Upward, the condition’s failure has two consequences and each gets a rung. The camera has its own metamers constructs the vector bb this essay only asserted exists, in both directions. A camera profile is a fit takes up what happens when the exact matrix does not exist and one has to be chosen anyway.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 22 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Camera rawColour matrixLeast-squaresLuther conditionMetamerismNull spaceProjectionSpectral sensitivityStandard observerXYZ