What a camera does

Fitted to an eye nobody has

A camera profile's residual against the observer it was fitted to is under a unit, and that is the number everybody quotes. Against a person drawn from a population it is 3.35 at the ninety-fifth percentile — and a camera with no spectral error whatever, reporting the standard observer's tristimulus values exactly, would carry 3.47. The camera is not the problem.

Assumes A camera profile is a fit and Luther said when it would work.

A camera profile is a 3×3 matrix fitted by least squares so that a sensor’s three raw numbers land as near as possible to the tristimulus values a standard observer would have reported. What is left over is the Luther residual: the part of the sensor’s spectral sensitivities that no linear map can turn into colour-matching functions.

That residual is the honest measurement of a camera. It is a property of the silicon and the dye filters, it does not depend on the scene, and it is what every serious comparison of sensors reports. For the sensor modelled here it is a mean of ΔE00 0.92 over a set of test surfaces.

It is also not the number anybody wants, because nobody is the standard observer.

How wrong a camera profile is, and who it is wrong for. A colour matrix fitted against the 1931 observer, evaluated four ways. Its residual against that observer is ΔE00 0.92 — the Luther failure, which is a property of the sensor and is the honest measurement of the camera. Against a person drawn from a population of 160, the worst-off twentieth report 3.35. And a camera with no spectral error whatever, reporting the standard observer's own tristimulus values exactly, would leave 3.47. The camera is not the problem. It was fitted to somebody who does not exist, and so is the standard it was fitted against.
Fig. 1 The same profile evaluated four ways. Its residual against the observer it was fitted to is under a unit. Against a person, the worst-off twentieth report 3.35 — and a camera with no spectral error at all would leave them 3.47.

What the same profile is worth against one observer is a different question from what it is worth against a chart, and both are computed here.

A camera profile is a fit, and the sample set is a hidden argument to it. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.2 and tested twice: on those same surfaces (mean 0.65) and on 12 at chroma 0.9 (mean 1.75). Both bars come from the same matrix; only the surfaces differ.
Fig. 2 Where the residual sits across the space rather than surface by surface. A fit spreads what it cannot remove over whatever it was shown, so this map is a picture of the chart as much as of the sensor.
The best a camera profile can do on the surfaces it was fitted to. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.5 and tested on those same 12 surfaces — mean 1.13, worst 2.04. This is the most favourable measurement it is possible to make of a camera and it is the one usually published.
Fig. 3 And the self-report at a moderate chart: fitted and tested on the same twelve surfaces. Every number a manufacturer publishes is one of these, and none of them contains an observer at all.

The two extremes of the chart argument are worth having beside the population one, because a reader has to be told which of the two terms is larger.

The best a camera profile can do on the surfaces it was fitted to. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.2 and tested on those same 12 surfaces — mean 0.65, worst 1.12. This is the most favourable measurement it is possible to make of a camera and it is the one usually published.
Fig. 4 The flattering self-report: fitted on a narrow chart and asked about that same chart. This is the number a profiling package prints, and it contains neither an observer nor a surface anybody photographs.
The best a camera profile can do on the surfaces it was fitted to. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 24 surfaces at chroma 0.7 and tested on those same 24 surfaces — mean 1.38, worst 2.85. This is the most favourable measurement it is possible to make of a camera and it is the one usually published.
Fig. 5 And the same self-report on twenty-four saturated surfaces. Enlarging and saturating the chart moves the number by less than choosing a different observer does, which is the comparison this essay exists to make.

The claim

A camera’s colour error against a person is mostly not the camera’s.

  • The profile’s residual against the 1931 observer is ΔE00 0.92, and against a person drawn from a population of a hundred and sixty it is 1.18 at the median and 3.35 at the ninety-fifth percentile — nearly four times larger.
  • A camera with no Luther failure whatever carries 3.47 at the same percentile. Removing the entire spectral error of the sensor makes the answer very slightly worse, not better.
  • So the ratio between a real camera and a perfect one is 0.97. Whatever a manufacturer does to the filters, the error a viewer reports is dominated by a choice made in 1931.
  • And that is not a fault in the standard observer. It is what an average is: the mean of seventeen people is close to nobody, and a device fitted to it inherits the distance.

What is being compared

Four quantities, and keeping them apart is most of the work.

The Luther residual. Fit the matrix on a set of surfaces, evaluate it on the same set, measure the difference between what the camera reports and what the standard observer would have. This is the sensor’s own error and it is what a camera profile is a fit already computes.

The camera against a person. For each member of a population, compute what that member sees of each surface, and compare it with what the camera reports. This includes the Luther residual and the observer difference together.

A colorimetric camera against a person. Replace the camera with a device that reports the 1931 observer’s tristimulus values exactly — no spectral error at all — and ask the same question. This is the observer difference alone.

And the population against itself, which this site already reports: two members picked at random disagree by 4.6 units at the median.

The third of those is the one that has no equivalent in the imaging literature, and it is the one that makes the first interpretable. A residual is small or large relative to something, and the something has always been zero.

The arithmetic that makes them comparable

The camera reports a tristimulus value for the standard observer and a member reports one for themselves. Comparing them needs both in the same place, and the choice of place decides part of the answer.

Both are carried into CIELAB through an adaptation from their own white to D65, using the same transform the population machinery uses everywhere. That is the construction which makes the numbers here directly comparable with the disagreement between two people that the same file reports, rather than a second scale nobody can set against anything.

It also has a consequence worth stating: an overall difference in white balance between a member and the standard observer is largely removed, because both are adapted to their own view of the illuminant. What survives is the part of the disagreement that is not a white-point shift — which is the honest thing to measure, since a camera’s white balance is chosen per shot and its matrix is not.

Why removing the camera’s error changes almost nothing

The two population numbers at the ninety-fifth percentile — 3.35 for the real camera and 3.47 for a perfect one — are close enough that the difference is inside the noise of the population sample, and the ordering is the wrong way round from the obvious expectation. At the median the same inversion is a quarter rather than a twenty-eighth, which is a separate matter and has its own section below.

The reason is that the two errors are not aligned. The Luther residual points in whatever direction the sensor’s sensitivities fail to be a linear combination of the colour-matching functions; the observer difference points in whatever direction that member’s cones differ from the standard’s. There is no reason for those to be parallel, and when two errors of similar size are combined at random angles the result is dominated by whichever is larger — which here is the observer.

The residual is 0.92 and the observer difference is 3.47, so the camera contributes about a quarter as much in magnitude and, added in quadrature at a random angle, less than a tenth as much to the total. A profile ten times better would move the number a reader reports by a fraction of a per cent.

The inversion is a quarter, not a whisker, and it needs explaining

The claim above is that removing the camera’s spectral error makes things “very slightly worse”, and at the ninety-fifth percentile that is fair — 3.47 against 3.35 is three and a half per cent, which is inside the noise of a hundred and sixty draws.

At the median it is twenty-five per cent. The real camera leaves 1.18 and the colorimetrically perfect one leaves 1.47, and a quarter is not a whisker.

It is also not something two independent errors can do. If the camera’s residual and the observer difference were unrelated vectors, adding the first to the second could only push the median magnitude up — a fixed non-zero offset added to a cloud of displacements increases the typical distance from the origin, whatever direction it points. The measured median is twenty per cent below what a perfect device manages, so the two errors are not unrelated: the camera’s residual must be partly anti-aligned with the population median’s own displacement from the 1931 observer.

That is a coincidence unless something makes them share a direction, and there is a candidate. The population’s members are pigment templates seen through a modelled lens and macular pigment, and both of those absorb in the short wavelengths — so a typical member is, relative to the 1931 functions, deficient in the blue. The camera is silicon behind dye filters, and silicon’s quantum efficiency falls away at the short end too. Two devices that are both a little weak in the blue relative to the standard are, in that respect, a little like each other, and a matrix fitted to carry one of them onto the standard lands closer to the other than an exact standard device would.

That is a hypothesis and it is stated as one, because nothing here measures it. The test is cheap and it is the obvious next computation: project both displacements — the camera’s residual and the population median’s offset — onto the same axes and take the angle between them. Anti-alignment predicts an obtuse angle, and the size of the cancellation predicts how obtuse. If the angle comes out near ninety degrees the explanation fails and the twenty-five per cent is something else.

The percentile behaviour is consistent with the hypothesis and is the reason it is worth testing. A directional cancellation helps the observers who lie along the direction it cancels and does nothing for the ones who lie across it. The median member is near the population’s centre, which is where the displacement is, so they get the full benefit — twenty-five per cent. The ninety-fifth-percentile member is unusual in some direction the cancellation was not aimed at, so they get almost none — three and a half per cent. A cancellation that fades as the observer becomes unusual is what a fixed anti-aligned offset produces, and it is not what a chance agreement in magnitude would produce.

None of that disturbs the essay’s conclusion and it sharpens what the conclusion is about. The camera is still not the binding term: 3.35 against 3.47 at the percentile a specification would be written against, and 0.97 as the ratio. What changes is the reading of the median row. The real camera is not merely no worse than a perfect one for the typical viewer — it is measurably better, by an amount that comes from an accident of which direction two unrelated devices happen to fail in. An accident is a poor thing to rely on, and a manufacturer who improved the sensor’s blue response would be removing it.

What this does not say

It does not say camera profiles are pointless. A profile with no matrix at all — raw values interpreted as though they were tristimulus values — is wrong by an enormous margin, and the fit removes almost all of that. What the fit cannot remove is the residual, and the argument here is only about the residual.

It does not say sensor design is pointless either. The comparison is made at one point in the space of sensors, and a genuinely bad sensor — a machine-vision camera with no colour filters worth the name, or one whose infrared cut is missing — has a residual far larger than any observer difference. The claim applies to sensors already good enough that their residual is under a unit, which is most of what is sold for photography.

And it says nothing about accuracy in the sense a photographer means. Skin tones, memory colours and the tone curve decide whether a photograph looks right, and none of them is a colorimetric residual. That is a different subject and this site is careful to keep them apart.

A camera profile is a fit, and the sample set is a hidden argument to it. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.2 and tested twice: on those same surfaces (mean 0.65) and on 12 at chroma 0.9 (mean 1.75). Both bars come from the same matrix; only the surfaces differ.
Fig. 6 The measurement a profile’s residual is most sensitive to, which is not the sensor: which surfaces it was fitted on. Fitting on pale samples and testing on saturated ones multiplies the error, and the number quoted for a camera is usually the first of those.

Whose eye should it be fitted to

If the residual is not the binding constraint, the interesting question is what the matrix should be fitted against instead of the 1931 observer.

Three answers are available and none is obviously right. Fitting against the population’s median observer moves the error from being centred on nobody to being centred on the median of a model — an improvement if the model is right and a different error if it is not. Fitting against the 1964 ten-degree observer is what a surface-colour laboratory would do and would be a defensible change; it is a different filter with a position, and for a photograph of a scene most of which subtends more than two degrees it is arguably the correct one.

Fitting to minimise the population’s ninety-fifth percentile is the third, and it is the same optimisation the display essay runs from the other end. It is computable with what is already here, it would produce a matrix nobody’s standard specifies, and it would be better for the ninety-fifth percentile by construction and worse against the standard observer by construction.

That last trade is the one nobody can make, because every acceptance test for a camera profile is written against the standard observer. A manufacturer who improved the number a viewer reports would fail the number a laboratory measures.

The four numbers, in one table

what is being measured median ΔE00 95th percentile
the camera against the observer it was fitted to 0.92
the camera against a person 1.18 3.35
a colorimetrically perfect camera against a person 1.47 3.47
two people against each other 4.6

The last row is not this essay’s measurement — it is the population file’s, quoted here because it is the scale everything else has to be read against. Two people picked at random disagree about a colour by more than a real camera disagrees with either of them.

Which means the honest way to describe a camera profile’s residual is not the error but a term, and not the largest one. A pipeline that reported a colour with an uncertainty attached would put the sensor’s contribution third, behind the observer and behind whatever the display is doing.

What was computed, and how

The sensor is built from its parts, as everywhere in this family: a silicon quantum efficiency curve, three colour-filter dyes with a stated infrared leak, and an infrared-cut filter of stated cut and width. Nothing about it is fitted to a real camera and the residual it produces is representative rather than particular.

The matrix is ordinary least squares from raw triples to tristimulus values over the fit set. The fit and test sets are the same here, deliberately — a profile evaluated on its own fit set gives the smallest residual it can, which is the conservative choice when the argument is that the residual is not what matters.

The population is a hundred and sixty eyes from five measured variates, each of which was already an argument to the retinal model. Twelve test surfaces at moderate chroma, which is where a camera profile is usually characterised.

Where it stops

The population is a model with independently sampled variates that are not independent in people, so the spread is an upper estimate. If the real population is narrower than the model, the camera’s share of the error rises — and it would have to be four times narrower before the residual became the larger term.

One sensor, one matrix, one set of surfaces, one illuminant. Every one of those could be varied and the essay’s number would move; what would not move is the structure, which is that the two error terms are compared at all.

And the standard observer is treated as a reference rather than as a claim. Nothing here says the 1931 functions are wrong — they are an average and they are an accurate average. The point is that a device fitted to an average is accurate for the average, and a viewer is not one.

Who found it, and when

Luther’s condition is from 1927 and Ives’ from a little earlier: a camera reproduces colour correctly if and only if its sensitivities are a non-singular linear transformation of the colour-matching functions. Everything since is a measurement of how badly that condition is failed and what to do about the failure.

The literature on evaluating camera profiles is large and almost entirely single-observer. The standard figures of merit — the sensitivity metamerism index, the various forms of quality factor, and plain fitting residuals — all measure a distance from the standard observer, because that is what a colorimetric device is supposed to agree with.

Observer variability entered the display world first, driven by narrowband backlights, and has not crossed to capture in the same way. There is a reason: a display is looked at directly and a photograph is looked at on a display, so the observer variation in a photographic pipeline enters at the viewing end where it is already being discussed. What this essay adds is that it also enters at the capture end, that the two are the same size, and that the capture end is where the effort is currently going.

Where the disagreement comes from, one variate at a time. The same measurement with one source of variation live and the other three held at their medians. The largest is the lens — entered as an age, because that is what it is a function of — at 7.7 ΔE00 against lens density, entered as age 20–70. So the biggest single reason two people disagree about a colour is the one thing about an observer that is written on their passport, and no colour specification has a field for it. The shares do not sum to the whole and are not expected to: the variates enter a nonlinear function of the spectrum, so this is a ranking rather than a decomposition.
Fig. 7 Where the observer term comes from: five variates, ranked by how much of the disagreement each accounts for. The largest is the lens, which is age — the one thing about an observer written on their passport, and the one no camera profile has a field for.

The same argument, one step downstream

There is a version of this that is easier to see and has been visible for years, which is worth naming because it makes the unfamiliar version familiar.

A photograph taken on any camera is looked at on a display, and a display’s primaries are narrow. Two people looking at the same screen disagree about its white by a few units, and the narrower the primaries the more they disagree — which is measured, published, and the reason cinema colourists argue about laser projectors.

Nobody responds to that by improving the camera. The observer term at the viewing end is understood to be an observer term, because there is no device between the display and the eye to blame it on.

At the capture end there is a device, and the device gets the blame. The arithmetic is identical and the term is the same size; what differs is that a camera has a specification, a residual and a competitor, so the quantity with a number attached is the one that receives attention.

Where the ladder goes next

The optimisation is the obvious next thing and it is one function call away: fit the matrix against a percentile of the population instead of against a point, and report what it costs against the standard observer. That produces a matrix with a stated trade rather than an implicit one.

The sharper question is upstream of the matrix. A profile is three numbers per pixel, and everything this essay is about happens because three numbers are not enough to reconstruct a spectrum. A camera with four or five channels can be inverted toward a spectrum rather than toward one observer’s tristimulus values — and a spectrum can be rendered for whichever observer is looking. That is a real product category already, sold for measurement rather than for photography, and the argument for it in photography is exactly the number in this essay.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Camera rawColour managementColour matrixIndividual variationLeast-squaresLuther conditionObserver metamerismQuality controlSpecificationSpectral sensitivityStandard observerWhite balance