What a camera does

A camera profile is a fit

Since no matrix is exact, the one a manufacturer ships is a least-squares compromise over a set of surfaces somebody chose. Fitted on desaturated patches it is excellent on desaturated patches — and the sample set is a hidden argument to every camera profile in existence.

Assumes Luther said when it would work and A tolerance is a shape.

The previous rung established that no fixed matrix converts this camera’s raw to XYZ correctly, and that this is a theorem rather than an engineering shortfall.

A camera nevertheless has to ship with a matrix. Somebody has to choose it, and the choice is made the only way it can be: by picking a set of surfaces, measuring what a colorimeter says about them, measuring what the camera says, and solving for the matrix that makes those two agree as closely as possible.

A camera profile is a fit, and the sample set is a hidden argument to it. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.2 and tested twice: on those same surfaces (mean 0.65) and on 12 at chroma 0.9 (mean 1.75). Both bars come from the same matrix; only the surfaces differ.
Fig. 1 The same matrix, evaluated twice. On the twelve desaturated surfaces it was fitted to, mean ΔE00 0.65; on twelve saturated ones it never saw, 1.75. Both bars come from one matrix and one camera. Only the surfaces differ, and neither set is more legitimate than the other.

The claim

A camera profile is a least-squares fit, its error depends on which surfaces it was fitted to, and the sample set is a hidden argument to every one in existence.

Hidden in a specific and consequential sense: the matrix is published in the raw metadata or the converter’s tables, and the surfaces it was derived from are not. What survives into the file is the answer, without the question.

How a matrix is actually derived

The procedure is short enough to state completely.

Photograph a chart of nn patches under a stated illuminant. For each patch, record the raw triple rir_i and obtain the reference XYZi\mathrm{XYZ}_i from a spectrophotometer. Stack them as 3×n3 \times n matrices RR and XX, and solve

M=XRT(RRT)1M = XR^{\mathsf{T}}(RR^{\mathsf{T}})^{-1}

which is the MM minimising MRX\lVert MR - X \rVert in the Frobenius norm. Three lines of linear algebra and a chart.

Every decision in the procedure is in the words “a chart”, “a stated illuminant” and “the Frobenius norm”, and each is worth pulling apart.

The chart

The industry standard is a colour rendition chart of twenty-four patches, in production since 1976, and it was designed for a purpose that is not this one: it was made so that a photographer could judge a film’s rendering by eye, and its patches are chosen to be memorable colours — skin, foliage, sky, and a grey ramp — rather than to span the space a fit needs to be constrained over.

Six of its twenty-four patches are neutral. That is a quarter of the sample set carrying no chromatic information at all, and it pulls the fit hard towards getting neutrals right, which is the most visible thing to get wrong and the easiest.

The remaining eighteen are moderately saturated at best. Nothing on the chart approaches the chroma of a saturated dye, an LED, or a flower — which is exactly where the Luther residual is concentrated, because the residual lives in the spectrally narrow part of the function space and the chart is made of broad, smooth, paint-like reflectances.

So the standard sample set is systematically drawn from the region where the fit is easiest. That is not a criticism of the chart, which was built for a different job in a different decade. It is a description of what a profile derived from it is a statement about.

The best a camera profile can do on the surfaces it was fitted to. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.35 and tested on those same 12 surfaces — mean 0.92, worst 1.55. This is the most favourable measurement it is possible to make of a camera and it is the one usually published.
Fig. 2 The comfortable case: fitted and tested on the same moderately saturated set, which is how a profile is usually evaluated. Every bar is small and the mean is well under one. This is a measurement of the fit rather than of the camera.
The best a camera profile can do on the surfaces it was fitted to. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.2 and tested on those same 12 surfaces — mean 0.65, worst 1.12. This is the most favourable measurement it is possible to make of a camera and it is the one usually published.
Fig. 3 The flattering case, which is the one a profiling package reports: fitted on a narrow chart and asked about that same chart. Nothing here is wrong and nothing here is a prediction about a photograph.

Where in the space the residual actually sits is the other half of what a single number hides, and it is a map of the chart as much as of the camera.

A camera profile is a fit, and the sample set is a hidden argument to it. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.2 and tested twice: on those same surfaces (mean 0.65) and on 12 at chroma 0.9 (mean 1.75). Both bars come from the same matrix; only the surfaces differ.
Fig. 4 And where the error sits across the space rather than per surface. A fit spreads its residual over whatever it was shown, so the map of where a profile is wrong is a map of the chart as much as of the camera.

Two more corners of the same square say whether the size of the gap is a property of the chart or of the fit.

The best a camera profile can do on the surfaces it was fitted to. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 24 surfaces at chroma 0.7 and tested on those same 24 surfaces — mean 1.38, worst 2.85. This is the most favourable measurement it is possible to make of a camera and it is the one usually published.
Fig. 5 Twenty-four saturated surfaces, fitted and tested on themselves. Doubling the chart and saturating it does not make the fit exact, because what is left over is not a sampling error.
How wrong a camera profile is, and who it is wrong for. A colour matrix fitted against the 1931 observer, evaluated four ways. Its residual against that observer is ΔE00 0.92 — the Luther failure, which is a property of the sensor and is the honest measurement of the camera. Against a person drawn from a population of 160, the worst-off twentieth report 3.35. And a camera with no spectral error whatever, reporting the standard observer's own tristimulus values exactly, would leave 3.47. The camera is not the problem. It was fitted to somebody who does not exist, and so is the standard it was fitted against.
Fig. 6 And the same profile scored against a population of observers rather than against one. The fit was made to a standard nobody is, which is a term no chart can remove and no report contains.

The illuminant

A matrix derived under one illuminant is not correct under another, and the reason is worth stating because it is not the obvious one.

If the sensor satisfied Luther’s condition, one matrix would serve every illuminant — the condition is about the functions, not about any particular light, and M1M^{-1} recovers XYZ from raw regardless of what is illuminating the scene. The illuminant-dependence of a real profile is therefore a symptom of the condition failing, not an independent problem.

What happens in practice is that the fit trades off residuals across the sample set, and the weighting of that trade-off depends on which parts of each reflectance the illuminant emphasises. A tungsten lamp puts most of its visible power in the red; a daylight source spreads it. The same eighteen patches under the two lamps present different problems to the fit, and the two solutions differ.

The standard response is to ship two matrices — one derived near 2856 K and one near 6500 K — and interpolate between them according to the estimated illuminant. Adobe’s DNG specification has fields for exactly this pair. It is a good engineering answer to a problem that has no exact one, and it is worth noticing what it concedes: the profile is a function of the light, and the light was guessed by an algorithm making an assumption about the scene.

The reference

There is a fourth argument hidden inside the second half of the procedure, and it is the one this site is best placed to notice.

The reference XYZi\mathrm{XYZ}_i is not a fact about the patch. It is the patch’s reflectance integrated against a chosen observer under a chosen illuminant, and both choices are made by whoever operates the spectrophotometer. Almost universally that is the 1931 two-degree observer, because it is the discipline’s default — and it is the wrong one for surface colour, where the 1964 ten-degree functions are what a laboratory is required to use.

So a camera profile is typically fitted to make the camera agree with a two-degree observer looking at a chart, and is then used on photographs of scenes subtending far more than two degrees, viewed by people using something nearer the ten-degree functions.

This is not a large effect compared to the ones above and it is a real one, and it is invisible in every specification. The target of the fit is unrecorded, in the same way the sample set is: the matrix is published, and what it was fitted to agree with is not.

The norm

Minimising squared error in XYZ is the convenient choice and it is not the right one.

XYZ is not perceptually uniform — this site measures how far from uniform its derivatives are — so equal squared error in XYZ is not equal visible error. A fit that minimises XYZ residual spends its freedom where the eye is least sensitive as readily as where it is most, and the resulting profile is worse in ΔE terms than one fitted in a perceptual space would be.

Fitting in CIELAB or with a ΔE2000 objective is better and is harder: the problem stops being linear least squares with a closed form and becomes an iterative optimisation, which is why the closed form is what gets used.

The norm is a third hidden argument, and unlike the chart and the illuminant it is rarely even mentioned. Two profiles for one camera, from one chart, under one illuminant, fitted in two different spaces, are two different matrices with different error distributions.

What the norm is worth, measured

Calling the norm a third hidden argument invites the reading that the three are comparable in size. They are not, and the difference is worth having because it says which of the three is worth arguing about.

Take this page’s own sensor under D65, fit it two ways on the same twelve surfaces at chroma 0.25, and evaluate both on the same sets. The first fit is the closed-form least squares in XYZ. The second starts from that matrix and runs a simplex search over all nine entries against mean ΔE₀₀ directly, which is the fit the section above says is better and harder.

fitted against on the fit set on twelve at chroma 0.9
squared error in XYZ mean 0.758, worst 1.28 mean 1.757, worst 3.61
mean ΔE₀₀ mean 0.725, worst 1.81 mean 1.679, worst 2.84

The right norm is worth 4.4 per cent on the fit set and 4.4 per cent on the saturated one. The sample set, on the same sensor, is worth a factor of 2.3. So of the three hidden arguments, the two that are occasionally disclosed dominate, and the one that is almost never mentioned is the smallest by a factor of forty. The search was run out to eight restarts and six thousand steps and moved the third decimal place, so this is the converged answer rather than an under-searched one.

The interesting part is the worst column rather than the mean. Fitting against ΔE₀₀ makes the worst patch on the fit set worse, 1.28 to 1.81, while pulling the worst saturated patch from 3.61 to 2.84. A perceptual objective spends its freedom differently rather than uniformly better: it stops protecting the patches XYZ over-weighted, which are the pale ones the eye discriminates least well, and hands what it recovers to the saturated end.

That is the honest form of the recommendation. Fitting in a perceptual space is right, and a converter that does it is doing something defensible; but a photographer choosing between two profiles for one camera should ask which chart they came from long before asking which norm, and only the answer to the second is ever printed.

What the sample set costs, measured

The library’s assertion is the clean statement. Fit on twelve surfaces at chroma 0.2 and test on the same twelve: mean ΔE00 0.65. Fit on those and test on twelve at chroma 0.9: mean 1.75.

A factor of 2.7, from nothing but the choice of what to evaluate on.

Neither number is dishonest and both are routinely quoted. The first is what a manufacturer reports; the second is what a photographer of a red car experiences; and the gap between them is not a measurement error but a statement about which surfaces the profile was asked to care about.

This is the same shape as a finding the lamps essay reached from the other end of the pipeline. A colour rendering index is an average of shifts on a stated sample set, and how saturated that set is decides the verdict — two lamps within two points of each other on the index differ by 2.36 against 1.47 on how much they distort saturated surfaces. A sample-set-dependent average, quoted as a property of the device, is a recurring error in this discipline rather than a local one.

What the fit does with its freedom

A 3×33 \times 3 matrix has nine numbers and the fit has 3n3n equations, so with a twenty-four-patch chart it is heavily overdetermined and has to compromise. Where it puts the compromise is decided by the norm and by the sample set, and the result has a characteristic shape.

The diagonal entries come out near the white-balance multipliers, which is what one would expect: most of the job is scaling three channels so a neutral comes out neutral. The off-diagonal entries are large and of both signs, and they are where the interesting behaviour is. A typical row looks like 0.9R0.3G+0.4B0.9 R - 0.3 G + 0.4 B: the matrix is subtracting one channel from another to synthesise a response the sensor does not have, which is the same manoeuvre the Luther fit made when it went negative.

Two things follow, and each gets its own essay in this field.

The first is that the correction has a price in noise. Differences of large numbers amplify uncorrelated errors, so a matrix that corrects colour well has worse signal-to-noise than one that does not, and the trade-off can be swept out and measured rather than argued about.

The second is that the compromise is not evenly distributed over the space. The fit minimises a sum, so it will accept a large error on one patch to reduce small errors on many, and which patch takes the hit depends on the sample set again.

Correcting colour costs noise, and the two cannot both be least. Sweeping the colour matrix from its own diagonal — a white balance with no cross terms — to the full least-squares fit. Mean ΔE00 falls from 18.61 to 1.29; photon noise rises by a factor of 1.03. At 10000 photons per pixel.
Fig. 7 The price of the off-diagonal terms, swept from a matrix with none to the full fit. Mean colour error falls and photon noise rises along the same sweep, and no setting minimises both. This is a consequence of the sensor failing Luther’s condition rather than a manufacturing compromise.

The white balance is inside the matrix, and that is a decision too

A subtlety that trips up anybody reading a camera’s metadata for the first time: the shipped matrix and the white balance are not cleanly separable.

The formal decomposition is that white balance is a diagonal matrix applied in the sensor’s own space and the colour matrix is applied after it, so the whole conversion is MDrM D r with DD diagonal. In practice the two are usually folded into one 3×33 \times 3, and different converters split the product differently — which means two converters can apply the same total transform and disagree completely about what the white balance “was”.

It matters because the two are estimated differently. The matrix is fitted once, in a laboratory, against a chart. The white balance is estimated per photograph, by an algorithm, from the photograph itself. Folding them together mixes a calibration with a guess, and the folded product inherits the uncertainty of the guess.

Which is why the diagonal is not where a profile’s interesting content is. The nine numbers are three that were measured, six that were fitted, and a per-shot correction that was inferred, and only the middle six are a statement about the camera.

Fitting at a middling chroma and testing at a high one is the case a manufacturer’s own chart most resembles.

A camera profile is a fit, and the sample set is a hidden argument to it. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.5 and tested twice: on those same surfaces (mean 1.13) and on 12 at chroma 0.9 (mean 1.77). Both bars come from the same matrix; only the surfaces differ.
Fig. 8 Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on twelve surfaces at chroma 0.5 and tested both on those and on twelve at chroma 0.9 — mean 1.13 against 1.77. Both bars come from the same matrix and only the surfaces differ.

Where the model stops

A real profile is not just a matrix. ICC camera profiles and DNG’s own format both allow a lookup table after the matrix, mapping the residual error out on a grid. That helps considerably on the surfaces the table was built from and has the identical structure of failure: a table is a fit too, with a denser sample set and the same hidden argument.

The forward direction is not the whole job. A profile also encodes a rendering intent — what to do about colours outside the output gamut — which is a separate decision with its own literature and is not a measurement at all.

And nothing here is about the exposure or the tone curve, which come after the matrix and dominate what a picture looks like without affecting whether it is colorimetrically right.

The generalisation

The transferable claim is about the difference between a model’s error on its training set and its error anywhere else, and it is old and endlessly rediscovered.

A fit reports its own residual, and its residual on the data it was fitted to is not an estimate of its error. Everyone knows this in the abstract and the field of camera profiling is a good example of how thoroughly it survives being known: profiles are routinely characterised by their error on the chart they were derived from, published that way, and compared between manufacturers that way.

The specific protection is not complicated. Report the error on a set the fit did not see, and state both sets. That is what the figures in this essay do — the caption names the fit chroma and the test chroma every time — and it is why the two bars are drawn together rather than the good one alone.

The version of this with the sharpest teeth is about where the extrapolation goes. A fit is interpolating inside the convex hull of its sample set and extrapolating outside it, and the error grows much faster outside. A chart of moderately saturated patches puts every saturated surface outside the hull, which is why the failures are concentrated exactly where they are most visible: on the reds, the deep blues and the fluorescent things that people photograph and then complain about.

Fitting and testing on the same saturated set is the most favourable measurement a camera can be given, and it is the one most often published.

The best a camera profile can do on the surfaces it was fitted to. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 24 surfaces at chroma 0.9 and tested on those same 24 surfaces — mean 1.73, worst 3.77. This is the most favourable measurement it is possible to make of a camera and it is the one usually published.
Fig. 9 Twenty-four surfaces at chroma 0.9, fitted and tested on themselves: mean 1.73, worst 3.77. That is the best case, and it is the only case in which the number quoted is the number a reader is entitled to.

What a photographer can do about it

The argument so far is entirely about what the manufacturer decided, and it is fair to ask what is left to anybody downstream.

Fitting a profile on the surfaces that matter is available and rarely done. A photographer who will spend a week photographing one class of object — textiles, paintings, botanical specimens, printed proofs — can photograph a chart made of those materials under the actual lighting and fit a matrix to it. The result is better on that class and worse on everything else, which is exactly the trade this essay says a profile is, taken deliberately rather than inherited.

Photographing a reference in the frame converts the problem. A chart beside the subject makes the question relative rather than absolute, which is the kind of question a camera is good at. It does not repair a null-space mismatch and it removes almost everything else.

And knowing which surfaces are outside the hull is worth more than any correction. Deep reds, saturated blues, anything fluorescent, anything lit by a narrow-band source: these are the cases where the extrapolation is furthest and the published error is least relevant. A photographer who knows the list can treat those photographs as unreliable evidence about colour and everything else as reasonably reliable, which is a far better position than trusting all of it equally.

Who noticed, and when

The colour rendition chart is McCamy, Marcus and Davidson, 1976, published in the Journal of Applied Photographic Engineering. Its patches were chosen to have spectra resembling natural objects rather than to span a space, and the paper says so; the use of it as a fitting set came later and from a different community.

Camera profiling as a formal practice arrived with ICC’s colour management architecture in the mid nineteen-nineties, which gave input profiles a standard container and a standard set of transform types. DNG in 2004 added the two-illuminant matrix pair and the field structure this essay describes.

The most useful modern development is the publication of measured spectral sensitivities for large numbers of cameras by academic groups from around 2012 onwards — which, for the cameras covered, removes the hidden argument entirely: with the sensitivities in hand a matrix can be fitted for any illuminant and any sample set the user cares about, and the residual can be computed rather than trusted. That is a small literature and it covers a small fraction of the cameras in use.

Where the ladder goes next

Downward, this rung sits on Luther said when it would work, which is why a fit is needed at all, and on a tolerance is a shape, which is where an error in ΔE stops being a number and becomes a region.

Upward, no matrix is right everywhere asks what the best possible matrix leaves behind, which is a different question from what a particular fit leaves — and answers it with the same machinery run against a control that reaches machine precision.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 18 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Camera rawCIEDE2000Colour matrixΔEThe ICC profileIlluminantLeast-squaresLuther conditionSaturationWhite balance