What a camera does

A camera is a fourth observer

The 1931 functions, the 1964 functions and a person's own cones are three sets of three curves that collapse a spectrum onto three numbers. A camera is a fourth, built from silicon and dye rather than from pigment and neural wiring, and it agrees with none of them.

Assumes Three numbers and Whose eyes.

The most useful thing anybody can be told about a camera is that it is not a special kind of machine. It is the same kind of machine the reader is using to look at this page.

An eye takes a spectral power distribution — a function of wavelength, with as many degrees of freedom as one likes to give it — and returns three numbers. It does this by integrating that function against three fixed curves, one per cone type, and keeping only the three totals. That is the whole of the collapse, it is why colour is exactly three-dimensional, and it is why two spectra can produce one colour.

A camera does the same thing with different curves.

A camera's spectral sensitivities, after the infrared-cut filter. Silicon quantum efficiency times the colour-filter dye times the infrared-cut filter, per channel, on a grid running to 1100 nm rather than to 780. With the filter removed, 68 per cent of the area under the three curves lies beyond the visible band, and all three curves are the same curve out there.
Fig. 1 The three spectral sensitivities of a camera: silicon’s quantum efficiency multiplied by each colour-filter dye’s transmittance, multiplied by the infrared-cut filter. This is the plot every sensor datasheet prints. The axis runs to 1100 nanometres rather than to 780 because silicon does, and what happens in that right-hand two thirds is the subject of the next four essays in this field.

What the filter is holding back is most of the story, and taking it off says so.

A camera's spectral sensitivities, with the filter removed. Silicon quantum efficiency times the colour-filter dye times nothing else, per channel, on a grid running to 1100 nm rather than to 780. With the filter removed, 68 per cent of the area under the three curves lies beyond the visible band, and all three curves are the same curve out there.
Fig. 2 The same three channels with the infrared cut removed. To the right of the rule the three curves converge on one another, so the fourth observer’s channels are nearly a single channel over a band the eye has no receptor in at all.

Drawing the two states together is the clearest form of the same statement, and the difference between them is a piece of glass.

A camera's spectral sensitivities, after the infrared-cut filter. Silicon quantum efficiency times the colour-filter dye times the infrared-cut filter, per channel, on a grid running to 1100 nm rather than to 780. With the filter removed, 68 per cent of the area under the three curves lies beyond the visible band, and all three curves are the same curve out there.
Fig. 3 And the two drawn together. The difference between the pair of plates is a piece of glass, and it is what makes three numbers out of a device that would otherwise report one.
The green channel, as the three things multiplied to make it. Silicon's quantum efficiency, the colour-filter dye's transmittance, the infrared-cut filter, and their product. The dye is the only stage carrying any colour information and it is clear above about 800 nm — transmittance 0.92 at 900 nm against 0.95 at 550.
Fig. 4 One channel taken apart: silicon’s quantum efficiency, the dye that shapes it, and the filter that ends it. The fourth observer’s three functions are a product of three things, only one of which anybody chose for colour.

The claim

A camera’s spectral sensitivities and the colour-matching functions are objects of the same type — three functions of wavelength, integrated against a stimulus, returning three numbers — and everything colour photography can and cannot do follows from their being different functions.

Not worse functions. Different ones. The word “accurate” does almost no work in this field until it has been said what a sensor is being accurate to, and the answer is a set of curves measured on seventeen people in 1931.

The arithmetic is identical

For an eye, the response is

X=xˉ(λ)Φ(λ)dλ,Y=yˉ(λ)Φ(λ)dλ,Z=zˉ(λ)Φ(λ)dλX = \int \bar{x}(\lambda)\,\Phi(\lambda)\,\mathrm{d}\lambda, \qquad Y = \int \bar{y}(\lambda)\,\Phi(\lambda)\,\mathrm{d}\lambda, \qquad Z = \int \bar{z}(\lambda)\,\Phi(\lambda)\,\mathrm{d}\lambda

For a camera it is

R=sR(λ)Φ(λ)dλ,G=sG(λ)Φ(λ)dλ,B=sB(λ)Φ(λ)dλR = \int s_R(\lambda)\,\Phi(\lambda)\,\mathrm{d}\lambda, \qquad G = \int s_G(\lambda)\,\Phi(\lambda)\,\mathrm{d}\lambda, \qquad B = \int s_B(\lambda)\,\Phi(\lambda)\,\mathrm{d}\lambda

The same equation twice with the kernel renamed. Both are a 3×N3 \times N matrix multiplying an NN-vector, both throw away everything the three integrals do not distinguish, and both therefore have an enormous null space — which is what metamerism is, and which the camera has its own version of.

Nothing about this is metaphorical. It is the same linear algebra, and the site’s own metameric-black machinery does not need modifying to work on a sensor: it projects onto the null space of whatever 3×N3 \times N matrix it is handed.

Where the curves come from

The three colour-matching functions are a measurement. They are what seventeen observers did with three knobs in a laboratory in the nineteen-twenties, tabulated, averaged and standardised, and nothing about them is derivable from anything more fundamental. They are downstream of cone pigments, ocular media, macular pigment and a good deal of neural arithmetic, and the fact that a linear transform relates them to the cone fundamentals at all is a substantive empirical claim rather than a definition.

A camera’s three curves are a product of three things somebody chose.

Silicon’s quantum efficiency is not a design decision at all. It falls away in the blue because short wavelengths are absorbed within a few tens of nanometres of the surface, where the carriers recombine before anything collects them; it stops at 1107 nanometres because that is the energy of the band gap, 1.12 electronvolts, and a photon carrying less than that cannot lift an electron across it. Both ends are physics.

The colour-filter dye is the only stage in the whole stack that has anything to do with colour, and it is an organic pigment chosen from the ones that can be manufactured, deposited in a few micrometres and survive a decade of ultraviolet. There is no dye that is the green cone fundamental.

The infrared-cut filter is a dielectric stack whose whole job is to remove a signal that would otherwise dominate, and it is the subject of its own essay because removing it is not optional and doing so costs something.

So the eye’s curves are a measurement of people and the camera’s are a product of three constraints, none of which is match the eye. That they resemble each other at all is because both are trying to sample a three-hundred-nanometre band in three overlapping places, which does not leave enormous freedom.

How different is different

The question has an exact answer and it is the subject of the next essay but one, so only the number is quoted here. Fit the best possible 3×33 \times 3 matrix taking the sensor’s three curves onto xˉ\bar{x}, yˉ\bar{y} and zˉ\bar{z} — the best linear redescription of the sensor as an observer — and what is left over is 31.7 per cent of the matching functions’ own magnitude.

Thirty-one per cent is not a rounding error and it is not a manufacturing tolerance. It is the size of the gap between what a camera measures and what a colorimeter would have measured, and no amount of processing removes it, for a reason that is a theorem rather than an engineering limitation.

The comparison the site has already made

This is the third time this collection has compared two sets of three curves, and the earlier two are worth having in mind because they set the scale.

The 1931 and 1964 observers differ because one was measured on a two-degree field and the other on ten, so the macular pigment is in the first and largely out of the second. They are two standards for the same species and they disagree enough that a surface-colour laboratory is required to use the second.

Observers differ from one another by more than the two standards differ, which is the more uncomfortable of the two facts and the one that gets left out.

A camera differs from both by considerably more than either of those. The right way to read this is not that cameras are bad but that the family of things that reduce a spectrum to three numbers is larger than the family of eyes, and colour science is a description of one small region of it.

What the difference costs, in the one unit that matters

An abstract residual is easy to shrug at, so it is worth converting into the quantity a person arguing about a photograph would use.

Take twelve surfaces, fit the best possible matrix from this sensor’s raw to XYZ across all of them, and ask how far each one lands from where a colorimeter would have put it. The best achievable worst case is ΔE00 = 2.85 across twenty-four surfaces — against a control sensor built to satisfy Luther’s condition, where the same computation lands at 10710^{-7}, which is arithmetic noise.

To put 2.85 in context: a tolerance written into a paint contract is typically 1.0 and sometimes 0.5, and a unit of ΔE is not a fixed perceptual step anyway. A camera fitted on its own test chart and tested on the same chart is already outside what a supplier would accept for a painted panel, before anything has been asked of it that the chart did not contain.

That is not an indictment of cameras. It is the number that should be in mind whenever a photograph is offered as evidence about a colour, and it is why the closing essay of this field is called a photograph is not a measurement.

Twelve or twenty-four

The 2.85 is quoted twice with two different sample sets. The prose fits the matrix across twelve surfaces and reports the worst case across twenty-four; the figure beside it says the fit and the test used the same twelve. Only one of those can be the number, and which it is changes what the sentence after it means.

If the fit and the test are both on twelve, then 2.85 is what the essay calls it — the most favourable measurement available, a camera marked on its own homework, and the point lands with full force. If the fit is on twelve and the worst case is read off twenty-four, then half the surfaces were never seen by the fit, and 2.85 is not the most favourable measurement; it is a partial held-out result, better than the essay’s framing needs and worse than a same-set figure would be.

The two readings differ in which direction the number is a bound in, which is the property that makes a quoted error useful. A same-set worst case is a floor: nothing a real chart is asked will do better. A mixed-set worst case is neither a floor nor a ceiling, because twelve of its surfaces are floored and twelve are not.

Nothing here decides it, and the point is not that one of two numbers is a typo. It is that a residual without its sample set is not a measurement, which is the argument this essay closes on — the sample set is a hidden argument to every camera profile in existence — arriving one paragraph early and about this essay rather than about a manufacturer.

Why two thirds of the signal is past the visible band

The sixty-eight per cent is the most startling figure here and it is worth taking apart, because most of it is not about silicon at all.

Silicon’s extra band runs 780 to 1107 nanometres — 327 nanometres against the visible band’s 400, so 45 per cent of the width carrying 68 per cent of the area. Per nanometre the uncut sensor is 2.60 times as responsive out there as it is across the visible.

A factor of 2.6 sounds like a claim about quantum efficiency and is mostly a claim about counting. Inside the visible the three dyes divide the band between them, so at any given wavelength the sum of the three channels is roughly one channel’s worth of silicon. Past 780 the dyes have stopped absorbing and the three curves are the same curve, so the sum is three channels’ worth. That alone predicts a per-nanometre ratio near 3, and the measured 2.60 sits just under it — the shortfall being the dyes’ overlap, which lets the visible sum exceed one channel, working against silicon’s own preference for the near infrared.

So the right reading of the plate is not silicon is unexpectedly sensitive in the infrared. It is the infrared is where the camera stops being three instruments and becomes one, and the area follows from that. Which is the same fact the next essay uses to explain desaturation, arrived at from the budget rather than from the ratios.

Thirty-one per cent and two point eight five are not the same measurement

The two headline numbers are easy to read as a large one and its consequence, and they are measured on opposite kinds of stimulus.

The 31.7 per cent is a residual against the matching functions themselves, on a wavelength grid — which is to say it is a statement about how the sensor treats monochromatic light, one wavelength at a time. The 2.85 is a ΔE₀₀ over a set of broad, smooth reflectances. A residual concentrated in narrow spectral features is precisely what a smooth reflectance integrates away, so a sensor can carry a third of the matching functions’ magnitude as residual and still land within three ΔE₀₀ on a chart without any contradiction at all.

That is not a reassurance. It is the reason the chart is the wrong instrument for the question, and the essay’s own spectral strip says so: monochromatic stimuli are where two sets of matching functions disagree most and where a camera is worst. The 31.7 per cent is the honest size of the gap and the 2.85 is what a particular set of two dozen smooth surfaces manages to hide of it.

One small note on the control, since it is offered as the scale the 2.85 is read against. A ΔE₀₀ of 10⁻⁷ on values of order fifty is a relative size of 2 × 10⁻⁹, which is seven orders of magnitude above double precision rather than at it — so it is the accuracy of the fit and the quadrature rather than arithmetic noise. Nothing turns on the distinction here, because both are far below any threshold that means anything; it is worth stating only because a number described as machine noise stops being looked at, and this one is a measurement.

What raw is, given all this

A raw file contains the three integrals and nothing else. It is not a picture, it is not a colour, and it has no gamut, because a gamut is a statement about a display and these numbers have not met one.

This matters because a raw triple is routinely spoken of as though it were an RGB colour. It has three components and two of them are called R and G, and the temptation to treat it as sRGB is very strong. But an sRGB triple is a set of instructions to a display with named primaries, a named white and a named transfer function; a raw triple is three integrals against three curves that were never standardised and are usually not published. The two are the same shape and mean unrelated things, which is precisely the confusion a hex code invites one level further down the pipeline.

The one thing that would make it exact

There is a condition under which a fixed matrix takes camera raw to XYZ correctly for every spectrum whatever, and it was written down twice in the nineteen-tens and twenties. It is the subject of the essay named after it, and it says: the camera’s sensitivities must be a non-singular linear combination of the colour-matching functions.

It is worth noticing why that is the condition, because it follows from what has already been said here. If S=MAS = MA for some invertible 3×33 \times 3 matrix MM, where SS is the sensor’s three curves and AA is the observer’s, then for any spectrum Φ\Phi

SΦ=MAΦS\Phi = MA\Phi

so the raw triple is MM times the XYZ triple, and M1M^{-1} recovers XYZ exactly. Nothing about Φ\Phi was assumed. Conversely if SS is not in the row space of AA then there is a spectrum in the null space of SS that is not in the null space of AA — two lights the camera merges and the eye separates — and no linear map can undo a merge.

A difference of null spaces is not a difference a matrix can absorb. That single sentence is the whole field.

The three-hundred-nanometre omission

There is one respect in which the analogy between an eye and a camera fails completely, and the plate at the top of this essay is drawn on an axis running to 1100 nanometres so that it cannot be missed.

The colour-matching functions are zero above about 780 nanometres. Not small — zero, because the pigments do not absorb there and the question of what the eye does with an 850-nanometre photon has the answer nothing. Every illuminant on this site, every reflectance, every integral, stops at 780 for that reason, and stopping there costs the eye nothing whatever.

Silicon does not stop until 1107.

A camera's spectral sensitivities, with the filter removed. Silicon quantum efficiency times the colour-filter dye times nothing else, per channel, on a grid running to 1100 nm rather than to 780. With the filter removed, 68 per cent of the area under the three curves lies beyond the visible band, and all three curves are the same curve out there.
Fig. 5 The same sensor with the infrared-cut filter removed, which is what the sensor is. Sixty-eight per cent of the area under the three curves lies past 780 nanometres, and out there the three curves are the same curve — the dyes have stopped absorbing, so all three channels are measuring the same thing.

Two consequences follow and both get an essay. The first is that all three channels see the same thing out there, so every ratio is pushed towards one and the picture desaturates — which is what the infrared-cut filter exists to prevent. The second is methodological and is about this site rather than about cameras: the 380–780 nanometre grid every figure here has been drawn on is sufficient for an eye and insufficient for a sensor, and the phase that opened this field had to decide what to do about that.

Where the model stops

Three things this essay has quietly assumed and should not be read as establishing.

The sensor here is a model, not a datasheet. Its silicon curve is the right shape for the right physical reasons, its dyes are Gaussians with a stated infrared leak, and its filter is a smooth shoulder. Real sensitivity curves are measured with a monochromator and differ between manufacturers by more than the model differs from any of them. Everything below about the shape of the problem survives that; nothing about the exact 31.7 per cent does, and the figures name the model rather than a camera for that reason.

Linearity is assumed throughout, and it is nearly true. A photodiode is linear in photon count to a fraction of a per cent over most of its range, which is a much better linearity than the eye has, and it fails at both ends — at the noise floor and at saturation. Saturation is where a hue turns, and it is the subject of its own essay because the failure is not gradual.

Nothing here is about focus, resolution or noise. A camera is also an optical system and a sampling system, and the second of those has a colour consequence that is entirely independent of everything above: a black-and-white edge comes off the mosaic coloured.

The generalisation

The useful form of this essay is not about cameras.

Any instrument that reduces a spectrum to a small number of numbers has a null space, and the null space is the instrument. A three-channel camera, a three-cone eye, a four-filter satellite radiometer, a photometer with one channel, a spectrophotometer with thirty-one — each is a projection, and what each cannot see is exactly the orthogonal complement of what it can.

The reason this framing pays is that it makes the question does this instrument agree with that one into a question with an answer. Two instruments agree on every stimulus if and only if one’s kernel is a linear redescription of the other’s. Otherwise they agree on a subspace and disagree elsewhere, and the interesting engineering question is not whether they can be made to agree but which stimuli one is prepared to be wrong about.

That question has an answer for cameras and it is a commercial one: the sample set a manufacturer fits its matrix on. It is a hidden argument to every camera profile in existence, and it is chosen by someone who will never meet the photograph.

Who noticed, and when

Robert Luther stated the condition in 1927 and Herbert Ives had it in 1915, both in the context of colour photography and photographic three-colour separation rather than of electronic sensors, which did not exist. The mathematics did not need updating when they did.

What has changed is that the condition used to be a remark about an unattainable ideal and is now a design target that manufacturers publish a number against. The measure usually quoted is not the residual computed here but the sensitivity metamerism index, standardised by the CIE in 1995 — a different normalisation of the same quantity, reported so that two cameras can be ranked. That it can be ranked at all is the useful modern development; that no camera scores zero is Luther’s result, unchanged.

Where the ladder goes next

Downward, this rung sits on three numbers, which is where the collapse itself is set out, and on whose eyes, which is where the observer stops being a single object.

Upward, the field divides in two. One branch follows the curves: where they come from, what the filter is for, and the exact condition under which they would be enough. The other follows the file: what raw is, what the matrix does to it, and eventually what a photograph can be used to establish.

Both branches meet at the same place, which is that a photograph is a measurement made by an instrument whose kernel nobody published.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 16 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Camera rawColour filter arrayCone fundamentalsIntegrationProjectionSiliconSpectral sensitivityStandard observerTrichromacyXYZ