What a camera does

The chart decides the profile

A camera's colour matrix is nine numbers fitted to a set of patches somebody chose, and the number that comes with it is the error on those patches. On a chart with no chromatic range that number is 0.19 ΔE00 and the matrix delivers 1.65 — and a second matrix, indistinguishable on the chart, delivers 2.03.

Assumes A camera profile is a fit, No matrix is right everywhere and An image does not determine the light.

A camera profile is a nine-number matrix fitted by least squares to a chart of patches. It arrives with an error figure, and the error figure is the residual on the patches it was fitted to.

There are two separate things wrong with that, and only one of them is the familiar one.

What a camera matrix reports on its own chart, and what it delivers off it. The same camera fitted on charts of increasing chromatic range. The left bar of each pair is the mean error on the chart the matrix was fitted to, which is the number a profile comes with; the right bar is the error on a saturated set it never saw. At the thinnest chart the fit reports 0.19 ΔE00 and delivers 1.65, a factor of 8.6. The gap closes as the chart widens, and it closes because the chart improves rather than because the camera does.
Fig. 1 The same camera, fitted on charts of increasing chromatic range. The left bar of each pair is the error on the chart the matrix was fitted to; the right bar is the error on a saturated set it never saw.

What a chart cannot reach is as measurable as what it can, and it is the same thing whatever chart is used — the chart decides what the camera scores and not how good it is.

The part of a sensor a chart reaches, and the part it does not. Each channel's sensitivity, and the residue left after projecting it onto the span of the 12 illuminated patch spectra — the part no measurement of that chart constrains. A second camera differing by that residue produces identical raw values on every patch, to 2.9e-6 relative, and differs by 1.1% on a narrow light at 546 nm. The residue is 3.5% of each curve, which is small — patches that differ in where they are dark are a good probe, and that is the design rule.
Fig. 2 The residue. Each of the three channel sensitivities is drawn with the part of it that no measurement of any chart constrains, at six times scale — a second camera differing from this one by exactly that residue would fit every patch identically and photograph a room differently.

The same three columns applied to a scene and to a lamp say what a determined fit is worth against an undetermined one.

How short of determining the light a photograph is, as the scene grows. Each cell is the number of unknowns left over after every equation the image supplies: three sensors, a three-dimensional illuminant, and reflectances confined to a linear model of the dimension on the left. At one and two dimensions more surfaces close the gap. At three the gap never closes, because each further surface adds three equations and three unknowns; at four it widens. The count is arithmetic and has no algorithm in it.
Fig. 3 How many unknowns are left over once every equation an image supplies has been used. A camera profile has none left over; a scene has many, and the two failures look nothing alike.
A camera matrix fitted under each light, used under each light. Mean ΔE00 over the same surfaces, with the matrix fitted under the row's light and the scene under the column's. The diagonal is what a profile's data sheet quotes and is between 1.0 and 1.2 everywhere. Off it the numbers rise steeply: the matrix fitted under illuminant A reports 1.17 there and delivers 9.34 under a 9000 K daylight, a factor of 8.0. Nothing about the camera changes between cells.
Fig. 4 And the same camera and surfaces with the matrix fitted under one light and the scene under another. The chart is one hidden argument to the fit; the illuminant it was measured under is a second.

The claim

A colour matrix has nine parameters and a chart determines all nine, so the fit is never underdetermined. What varies is how well it determines them, and a chart with no chromatic range determines them badly while reporting an excellent number — 0.19 ΔE00 against 1.65 delivered, a factor of 8.7, with a design conditioned at 178 to 1.

  • The familiar failure is overfitting, and it is here: the fit residual and the held-out residual diverge as the chart narrows.
  • The less familiar one is a family. A second matrix that differs from the fitted one along the design’s weakest direction is indistinguishable on the chart and delivers 2.03 ΔE00 off it.
  • The conditioning is the diagnostic and it is free. The ratio of largest to smallest singular value of the patch responses runs from 178:1 on a near-neutral chart to 12:1 on a saturated one.
  • Widening the chart makes the reported number worse and the matrix better, which is the shape that makes this hard to sell: the honest chart looks like the worse instrument.
  • And none of it is about the camera. The same sensor is used throughout; only the patches change.

Two failures, kept apart

The first is the one every account of fitting mentions. A model with nine parameters fitted to twelve patches will follow those patches more closely than it follows anything else, and the residual it reports is optimistic. The repair is a held-out set, and this collection has had one for a long time: fit on one set of surfaces, report on another.

The second is about directions rather than about optimism, and a held-out set does not by itself diagnose it.

Least squares finds the matrix minimising error over the patches. If the patches happen to span the three sensor channels unevenly — if, say, they are all near-neutral, so that the three raw responses rise and fall together — then some combinations of the matrix’s entries barely affect the residual at all. Those combinations are not determined by the data in any useful sense, even though the fit returns a unique answer for them, and the answer it returns is decided by whatever tiny structure the near-neutral patches happen to have.

A camera profile is a fit, and the sample set is a hidden argument to it. Per-surface ΔE00 after the best 3 × 3 from raw to XYZ, fitted on 12 surfaces at chroma 0.2 and tested twice: on those same surfaces (mean 0.65) and on 12 at chroma 0.9 (mean 1.75). Both bars come from the same matrix; only the surfaces differ.
Fig. 5 And the same matrix asked about surfaces it was not fitted to. The gap is the first failure; the second one is invisible in both pictures and needs the conditioning number to see.

The slack, measured

The way to make the second failure concrete is to move the matrix along the weak direction and ask what it costs.

Take the fitted matrix, add a rank-one perturbation along the design’s weakest right singular direction, and increase it until the chart error rises by five hundredths of a ΔE — a change no measurement of the chart could distinguish. Then evaluate the perturbed matrix on the saturated set.

On the near-neutral chart the perturbed matrix reports 0.24 ΔE00 on the chart and delivers 2.03 off it, against 0.19 and 1.65 for the original. A profile builder given the two matrices and the chart has no basis whatever for preferring one, and one of them is a fifth worse in use.

On the saturated chart the same construction moves nothing: the design is well conditioned, every direction of the matrix is pinned by the data, and there is no slack to exploit.

What a near-neutral chart looks like to the fit

It is worth picturing what “the patches do not separate the channels” means, because the phrase is abstract and the situation is not.

A near-neutral patch reflects roughly equally across the visible band, so the camera’s three raw responses to it are roughly proportional to the three channel sensitivities integrated against the light — the same three numbers, scaled, for every patch. Plot the twelve patches as points in the three-dimensional raw space and they lie almost along a single line through the origin.

A line is one direction, and a 3×3 needs three. The fit is not literally rank deficient — the patches are not exactly neutral and the line has thickness — but the two directions perpendicular to it are represented by the thickness, which is small, and it is the thickness that decides six of the nine matrix entries.

A saturated chart is a cloud rather than a line. Its three principal directions are comparable in size, every entry of the matrix is pinned by data that genuinely vary along the direction that entry controls, and the condition number falls accordingly.

Why the conditioning is the number to quote

The condition number of the patch responses is computable before any fitting is done, from the chart alone, and it is the quantity that predicts everything above.

At chroma 0.05 it is 178 to 1. At 0.85 it is 12 to 1. In between it falls monotonically. A profile builder can compute it in a second and almost none of them do, because the reported quantity is a residual and a residual is what a customer asks about.

The interpretation is standard and is worth stating in these units: a condition number of 178 means that a one per cent error in the measured patch data can become nearly a two-fold error in the weakest combination of the matrix entries. The reported residual does not go up when that happens, because the weakest combination is by definition the one that barely affects the residual.

This is the general form of the rule the essay on measurement conditions arrived at from a different direction — a 3×3 fitted on six near-white sheets appeared to succeed and was measuring the dimensionality of its test set. Count the dimensions the samples span before believing a linear fit, and the count is a condition number.

How many numbers a surface takes, on the friendliest set available. The cumulative share of the variance accounted for by the first few principal components of this collection's 240 constructed reflectances. Three components reach 99.94 per cent and two reach 97.99. The counting result allows two, and these are deliberately smooth curves built from a handful of Gaussians — the friendliest possible case, and it already needs 3.
Fig. 6 The same question about surfaces rather than about patches. A chart is a sample of a space; how many directions it actually visits is what decides everything a fit on it can say.

What was computed, and how

The camera is this site’s silicon sensor with an infrared-cut filter, and the reflectances are the site’s own test surfaces at a stated chroma. The chart and the held-out set are drawn from the same generator with different chroma, so the comparison isolates chromatic range and holds everything else fixed.

The design matrix is the raw responses to the chart patches, on the extended grid, through the illuminant, exactly as the shipped colourMatrix computes them. Its singular values come from the eigenvalues of the response Gram matrix by power iteration with deflation; the weakest direction is the dominant direction of the inverse Gram matrix, which is the same eigenvector and is cheaper to reach.

The perturbation is rank one and asymmetric across the matrix’s rows, applied more strongly to the row carrying luminance. That is a choice: a perturbation orthogonal to everything would be a smaller effect and a symmetric one would be a larger one, and the family of indistinguishable matrices is nine-dimensional rather than one-dimensional in any case. What is reported is therefore a lower bound on the slack rather than the size of the family.

The stopping band of five hundredths of a ΔE is stated rather than tuned. It is about the repeatability of a good spectrophotometer on a patch, so a matrix inside that band genuinely cannot be distinguished from the fitted one by measuring the chart again.

The conditioning is one over the chart’s chroma

The two condition numbers are given at the two ends of the sweep and the relation between them is simple enough to be a design rule.

178 to 1 at chroma 0.05 and 12 to 1 at 0.85 is an exponent of 0.95 — the conditioning is inversely proportional to the chart’s chroma, to within five per cent. Fitting the constant gives κ ≈ 9.6 / chroma, with the two endpoints implying 8.9 and 10.2.

The mechanism is visible in the essay’s own picture of the patches as a line in raw space. The smallest singular value is the line’s thickness and the thickness is the chroma; the largest is the line’s length, which is the neutral direction and does not depend on the chroma at all. A ratio of a constant to something proportional to the chroma is one over the chroma.

That turns use a chart with chromatic range into a number. A chart needs a mean chroma above about 0.5 to be conditioned better than 20 to 1, and 20 to 1 is where a one per cent measurement error becomes a twenty per cent error in the weakest matrix combination rather than a two-fold one.

chart chroma predicted conditioning
0.05 190
0.20 48
0.50 19
0.85 11

The rule also says why the standard chart is where it is. A target designed to sample the colours of skin, foliage and sky is sampling a population whose mean chroma is well under 0.3, so its conditioning is over thirty to one before anybody adds a grey scale — and the grey scale, which is the part everybody adds, drives it up.

What the slack is worth per unit of invisibility

The perturbation experiment has a leverage in it that the essay reports as two pairs of numbers.

A change of 0.05 ΔE₀₀ on the chart — chosen to be below a spectrophotometer’s repeatability, so genuinely undetectable — buys a change of 0.38 ΔE₀₀ off it, from 1.65 to 2.03. The leverage is 7.6: one unit of invisible movement on the chart is seven and a half units of visible movement in use.

And the ratio between the two failures is worth having. The fitted matrix reports 0.19 and delivers 1.65, a factor of 8.68; the design’s conditioning is 178, whose square root is 13.3. So the reported-to-delivered gap is 65 per cent of what the conditioning permits — the bound is not merely respected, it is nearly saturated, which is what says the gap really is the conditioning rather than something else.

That comparison has a use at the other end of the sweep. At chroma 0.85 the conditioning is 12 and its square root is 3.46, so on a saturated chart the conditioning can account for at most a factor of three and a half between what a fit reports and what it delivers. Anything larger than that on a well-conditioned chart is the held-out set being harder rather than the fit being loose, and the two have different remedies: a better chart fixes the first and nothing fixes the second except deciding which surfaces to be right about.

Two matrices, two per cents

The perturbed matrix’s two numbers are worth putting in the same units, because they say what the tolerance band is buying.

It is 26 per cent worse on the chart — 0.19 to 0.24 — and 23 per cent worse in use — 1.65 to 2.03. Those are nearly the same fraction, which sounds like the perturbation being harmless and is the opposite.

The point is the absolute sizes. Twenty-six per cent of 0.19 is five hundredths of a ΔE, which is below the repeatability of the measurement the chart was made with. Twenty-three per cent of 1.65 is nearly four tenths, which is a third of a tight delivery tolerance. The same proportional change is noise at one end of the fit and a real error at the other, and the reason is entirely that the two ends are two orders of magnitude apart in scale.

Which is the compact statement of the whole essay. A residual reported on a near-neutral chart is a small number, and a small number has small differences in it, so every distinction that matters downstream has been compressed into a range the measurement cannot resolve. Widening the chart does not merely give a better answer; it puts the answer on a scale where the differences between candidate answers are larger than the noise.

Where the model stops

No matrix is right everywhere in any case. A camera is a fourth observer whose sensitivities are not a linear transformation of the colour-matching functions, so there is no 3×3 that is correct on all spectra and the best possible one has a floor. Everything in this essay is about how well a chart locates the best available matrix, not about how good the best available one is.

A real profile is not a 3×3. Commercial camera profiles carry two matrices for two illuminants, a look-up table, and a tone curve, and the extra machinery changes the parameter count without changing the argument — a table has more parameters and the same question about which of them the chart determines.

The held-out set is also constructed, and is a hard one. Chroma 0.85 surfaces are more saturated than most real objects, so the delivered numbers here are pessimistic in absolute terms. What they are not is pessimistic in ratio: the comparison between a narrow chart and a wide one uses the same held-out set for both, so the factor of 8.7 stands whatever the held-out set is, as long as it is the same one throughout. Choosing what to be right about is the underlying decision, and this essay is about how well any given choice is executed.

And the surfaces are constructed. The chroma parameter is a knob on a generator rather than a property of a real chart, so the numbers belong to this construction. What does not belong to it is the mechanism: a set of patches that does not separate the channels cannot pin the combinations that separate them, whatever the patches are made of.

Who found it, and when

The conditioning of least-squares design matrices is old and general — it is the subject of the numerical-analysis literature from the 1960s onward, and the singular-value diagnosis is standard everywhere that regression is taken seriously.

Its absence from colour calibration practice is the interesting part. The standard 24-patch chart was designed in 1976 to sample the colours of common natural objects — skin, foliage, sky — which is exactly the right criterion for evaluating a camera on realistic scenes and exactly the wrong one for determining a matrix, because natural objects are clustered. The chart’s design goal and the fit’s requirement are in tension, and the tension is not usually noticed because both are called “a colour chart”.

Charts with more and more saturated patches exist and are used by people who profile for a living. The argument for them is normally phrased as covering more of the gamut, which is the held-out argument. The conditioning argument is a second and independent reason for the same recommendation.

The two arguments recommend different things at the margin, which is how they can be told apart. Covering the gamut says to add patches wherever the gamut is not yet covered, including near the neutral axis if that region is sparse. Conditioning says to add patches that are unlike the ones already there in the three raw channels, which means saturated patches specifically and says nothing in favour of more neutrals. A chart designed for the first reason and evaluated by the second usually turns out to have too many greys — and greys are the patches everybody adds, because they are the ones a tone curve needs.

The generalisation

The pattern is a fit that is determined but badly determined, reporting a residual that is insensitive to exactly the directions the data failed to pin.

The two diagnostics are cheap and neither is a residual. The conditioning of the design says how much of the parameter space the data actually visited. And the slack — how far the answer can move without the reported number changing — says what that costs, in the units the answer is used in.

Both are available before any held-out data exist, which matters because a held-out set is often the expensive thing. A chart that is badly conditioned can be identified and improved without printing a second one.

The trap that makes this hard to act on is the direction of the effect. Widening the chart raises the reported error and lowers the delivered one. A profile built on a narrow chart wins every comparison that quotes a residual, and quoting a residual is what everybody does.

What each fitted thing in these essays carries, what its data fix, and what is left. Three columns per row: how many numbers the model has, how many the stated data determine, and the difference — the dimension of the family that fits equally well. The third column is the one nobody publishes. A zero there does not mean the model is right; it means it is determined, which is a much weaker property and is compatible with being determined badly, as the camera row is.
Fig. 7 The camera row in this collection’s ledger. Its last column is zero — the fit is determined — and the note beside it is the whole of this essay, because determined is not the same as determined well.

What a profile ought to come with

Three numbers rather than one, and all three are already computed by any code that fits a matrix.

The residual on the fit set, which is what is quoted now and is not useless — it catches gross errors in the measurement of the chart.

The condition number of the design, which costs one small eigendecomposition and says whether the residual can be believed.

The residual on a held-out set, which is the number the profile will actually deliver and is the one a user wants.

None of that is a research programme. It is three lines added to a fitting routine, and the reason it is absent is the same reason a whiteness figure is quoted without its measurement condition: the extra numbers are unflattering, they invite questions, and the single number has been the convention long enough that quoting more looks like special pleading.

Why the effect is invisible in practice

A defect that has been in every profiling workflow for thirty years needs an explanation of why nobody trips over it, and there is one.

Everybody uses similar charts. The standard targets are all built on the same principle — sample the colours of ordinary objects — so every profile is conditioned about equally badly, every reported residual is optimistic by about the same factor, and every comparison between two cameras or two profiling packages is fair. A systematic bias shared by every measurement is invisible to every comparison anybody makes.

And the delivered errors are small enough to live with. A mean of 1.65 ΔE00 on saturated surfaces is not good and it is not catastrophic, and a photographer will attribute it to the camera, the lens, the light or their own grading long before attributing it to the patch set. There is no moment at which the failure announces itself.

That combination — a bias shared by everybody and a consequence that has other plausible causes — is the standing recipe for a defect that survives indefinitely. It is the same recipe that kept a measurement condition out of a whiteness figure until two instruments started to disagree.

Where the ladder goes next

If the chart leaves directions unpinned, there is a family of sensor descriptions consistent with it, and the chart determines the sensitivities themselves rather better than it determines the matrix — which is the same arithmetic asked a different question, and it comes out the other way.

The same counting question runs the other way in delivery. A printer profile is a table exact at its nodes and interpolated everywhere anybody prints, and there the unmeasured directions are literally the space between the patches.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 16 that link here.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationCamera rawColour matrixDegrees of freedomHeld-out validationIdentifiabilityLeast-squaresLuther conditionOverfittingRankSpectral sensitivity