No matrix is right everywhere
Assumes A camera profile is a fit and How far apart are two colours.
The previous rung was about a choice: which surfaces a profile is fitted to, and what that choice costs on the ones it was not.
This one removes the choice. Suppose the sample set were perfect — every surface anyone will ever photograph, weighted correctly, fitted in the right space with unlimited computation. What is left?
The claim
The best possible from this sensor’s raw to XYZ leaves a worst case of ΔE00 2.85 across twenty-four surfaces, and the same computation against a Luther-satisfying control reaches .
The control is the load-bearing half. A residual quoted alone is a number that could mean anything — a limit of the model, a limit of the arithmetic, a badly conditioned solve. A residual beside a control that reaches machine precision under the identical code is a measurement of the sensor.
Why the best case is not the interesting case, except once
Fitting and testing on the same set is normally a mistake, and this essay does it deliberately, which is worth explaining.
The previous rung is about generalisation: a fit’s error on its own data is not an estimate of its error elsewhere, and quoting the first as the second is the standard failure. That argument stands.
But there is one question the self-test answers and nothing else does: is the residual a property of the fit or of the sensor? Give the fit the test set in advance, let it optimise directly against exactly what it will be scored on, and whatever is left cannot be blamed on generalisation, on the chart, on the illuminant or on the norm. It is what the sensor’s kernel makes impossible.
That is why the figure above is drawn with one bar series rather than two, and why the generator asserts the equality rather than the inequality when the sets are the same. A figure that quietly asserted nothing in its degenerate case would be worse than one that throws — which is how this figure was found: three placements asked for the self-test case and the build stopped.
Where the error lands
It is not spread evenly, and the pattern is the same one the whole field has been circling.
The largest residuals are on the most saturated surfaces, and the reason is spectral rather than geometric. A saturated reflectance is one with a lot of structure — a steep edge, a narrow band, a deep trough — and structure is exactly what distinguishes two smooth kernels. A pale, broad, paint-like reflectance is nearly in the subspace both kernels agree on; a saturated dye is not.
Both of those are one matrix on two sets. A third and fourth setting say whether the size of the gap is a property of the sets or of the fitting.
A moderately saturated chart fitted and tested on itself is the case a manufacturer publishes, and it sits between the two above.
Two more settings of the same pair of arguments say where the boundary between “fitted” and “asked about” actually falls, and neither of them is a boundary a report mentions.
The observer is the argument none of those four settings touches, and it is the one that decides whether “right” was ever available.
So the practical form of the theorem is a list. Deep reds, saturated blues and violets, anything under a narrow-band lamp, anything fluorescent, and any photograph of a screen. Those are the cases where a camera’s colour is least trustworthy, and the list is derivable rather than empirical: it is the set of stimuli with energy where the two kernels differ.
How many surfaces it takes to see it
The residual is quoted over twenty-four surfaces, and almost all of it is visible in four.
Fitting and testing on the same set with the set’s size varied: three surfaces give exactly 0.0000. Nine coefficients against nine equations — a 3×3 interpolates three surfaces and leaves nothing over. The control stays at 10⁻¹² throughout, as it must.
The fourth surface takes the worst case from zero to 2.822. Adding twenty more takes it to 2.847, a further nine tenths of a per cent; adding forty-five more takes it to 2.989. From four surfaces upward the sequence runs 2.822, 2.564, 2.570, 2.536, 2.579, 2.710, 2.847, 2.989.
So the sensor’s failure is fully exposed by the first surface a 3×3 cannot interpolate. A twenty-four-patch chart says nothing about the worst case that the fourth patch had not already said — which is a stronger claim than the one above and is exactly what the self-test is for: if generalisation is not the issue, sample size is not either.
The mean behaves differently, and the difference is instructive. It is 0.874 at four surfaces, 1.756 at five, and settles between 1.38 and 1.45 from eight upward. So the mean needs a dozen surfaces before it means anything and the worst case needs four — two statistics over one measurement, converging at very different rates, and the one a specification is written against is the slower of them.
It also says what a larger chart buys. Between twenty-four surfaces and forty-eight the worst case moves by five per cent and the mean by two, so a chart’s size is buying precision on a mean rather than any new information about what the sensor cannot do.
Why a lookup table does not fix it
The obvious escalation is to stop insisting on a linear map. Fit a lookup table on a dense grid, interpolate between its nodes, and the residual on the grid can be driven to zero.
It helps and it does not repair anything, for a reason that follows directly from the null-space argument.
A lookup table is still a function of the raw triple. Two spectra with the same raw triple get the same table entry, the same output and the same colour, however dense the grid is. The pairs the camera merges stay merged. And the pairs the camera invents stay invented, because a table cannot know that two different inputs should map to one output any more than a matrix can.
A table increases the accuracy of the map and does not change its domain. The domain is three numbers, the information was destroyed at the sensor, and no post-processing recovers it — which is the same sentence as the necessity half of Luther’s condition, arriving as a practical remark about profiles.
What a table does buy is a better fit to the smooth part of the relationship, which is real and is why high-end profiles use one. It converts a residual dominated by the linear model’s inadequacy into a residual dominated by the sensor’s kernel, and the second is smaller. It is a floor, not a fix.
The control, in more detail
The Luther-satisfying sensor is worth describing because it is the only thing in this field that behaves the way everybody assumes cameras behave.
Its sensitivities are the colour-matching functions through a fixed non-singular mixing matrix, so by construction . Run the fitting procedure and comes back out. Test on twenty-four surfaces at any chroma and the worst error is , which is not a good fit but an exact one measured through floating-point arithmetic.
Test it for metamers and there are none: the largest raw gap found over the whole search is . It merges what the eye merges and separates what the eye separates.
And it cannot be built, because two of its three curves go negative and a filter with negative transmittance would have to emit light. The control exists in the library and nowhere else, which is exactly what makes it a good control: it isolates the sensor’s failure from every other source of error in the computation, including the ones nobody thought of.
How large is 2.85
A number in ΔE00 means nothing until it is put beside the tolerances people actually write, and this site has already established that those are not one number either.
A paint or coatings tolerance is commonly 1.0 and sometimes 0.5 for a critical match. Textile and automotive interior matching runs tighter than that on adjacent parts. So a worst case of 2.85 is several times a commercial tolerance, from an instrument given every advantage.
But a unit of ΔE is not a fixed perceptual step, and that cuts both ways here. MacAdam’s discrimination ellipses differ in area by a factor of 74 across the diagram, so a difference of 2.85 in one region is unmistakable and in another is at the edge of noticing. Where a camera’s worst residuals land — the saturated regions — is where the ellipses are largest, which means the visible consequence is smaller than the number suggests.
That is a genuine mitigation and it is not an exoneration, for two reasons. The first is that saturated colours are precisely what people care about getting right in a photograph, whatever the discrimination data says about threshold. The second is that a suprathreshold difference is not a scaled threshold difference — the site measures the two disagreeing by a factor of nearly five in scale and by up to seventy degrees in direction — so an argument from ellipse size to noticeability is weaker than it looks.
What happens if the sensor is allowed to change
Since the obstruction is the sensor rather than the arithmetic, it is worth asking what a differently designed one would buy, and the answer is a well-defined optimisation with a disappointing conclusion.
The design space is three non-negative transmittance curves, subject to being manufacturable, stable and reasonably efficient. Within it, the residual can be reduced by making the curves narrower and more separated — which improves the fit to the matching functions’ shapes and simultaneously throws away light, since a narrow filter passes less of it. A sensor with the best achievable Luther residual is a sensor with poor sensitivity, and low sensitivity costs noise, which costs colour accuracy through a different route.
So the residual is not a number anybody is trying to minimise. It is one term in an objective that also contains quantum efficiency, noise, cost and manufacturability, and the shipped answer is a point on a surface rather than a failure to reach a target. That reframing matters because it explains the otherwise odd fact that camera colour accuracy has barely moved in two decades while every other imaging quantity has improved enormously.
Where the model stops
The 2.85 is this modelled sensor’s number. A different dye set has a different residual, and published sensitivity-metamerism indices for real cameras span a considerable range. What survives is the structure: a non-zero best case, concentrated on saturated and spectrally narrow surfaces, with a control reaching machine precision.
Twenty-four surfaces is a small test set and a larger one drawn from the same family would give a similar mean and a worse maximum, because the worst case is a maximum over whatever was sampled. The mean is the more stable number and the maximum is the more honest one.
And the fit is in XYZ rather than in a perceptual space, for the reason the previous rung gives: the closed form exists in XYZ and does not in CIELAB. A perceptually-fitted matrix would report a lower ΔE and a higher XYZ residual, and would not change the sign of anything.
The generalisation
The useful abstraction is about the difference between a model that is under-fitted and one that is structurally incapable, and about how to tell them apart.
An under-fitted model improves when given more capacity: more parameters, a denser grid, a nonlinear form. A structurally incapable one does not, because the information it needs is not in its input. The two look identical from the outside — both show a residual that does not go to zero — and they call for opposite responses. The first wants a bigger model; the second wants a different measurement.
The way to tell them apart is a control that the model can fit exactly, and this essay is an extended argument for building one. If a synthetic case in which the answer is known produces machine precision under the identical code, then whatever the real case leaves is the real case’s own. If the synthetic case does not, the residual is the model’s or the code’s and nothing has been learnt about the world.
That is a habit worth more than any of the specific numbers here, and it recurs throughout this collection. The published CIECAM16 test vector plays the same role; so does the round trip through an inverse, and so does the closed form a quadrature is checked against. In each case a claim about the world is made trustworthy by a claim about arithmetic that had to come out exactly right first.
The residual is not the same object as the error
One distinction has been running under this whole field and is worth stating explicitly before the ladder moves on, because conflating the two makes the numbers incomparable.
The Luther residual is a property of the functions. It is 31.7 per cent, it is computed wavelength by wavelength against the matching functions, it makes no reference to any illuminant or any surface, and it is the quantity that is zero if and only if the condition holds.
The colour error is a property of a set of stimuli. It is 2.85 ΔE00 worst case, it is computed on twenty-four specific surfaces under one illuminant through one matrix, and it would be a different number for any other choice of those things.
The first is the cause and the second is one of infinitely many measurable consequences. A large residual guarantees that some stimuli have large errors; it says nothing about whether the ones anybody photographs are among them. That is why both numbers appear throughout — the residual to establish that the failure is structural, the error to establish that it reaches into practice.
The gap between them is also the reason the two literatures on this subject have been slow to connect. Sensor design papers optimise the residual, because it is the intrinsic quantity. Photographic and colour-management papers report the error, because it is what anybody experiences. Neither number converts into the other without stating the illuminant, the surfaces and the norm, and those are exactly the things that go unstated.
Who noticed, and when
The impossibility is Luther’s and is 1927. What is more recent is the sustained effort to characterise how much it costs, which begins in earnest with the CIE’s 1995 sensitivity metamerism index and continues with a substantial literature on optimal filter design: given that the condition cannot be met, which three realisable transmittances get closest, and closest in what norm.
That literature has produced two useful results and one negative one. The useful ones are that the optimum depends strongly on the norm and on the illuminant, and that a fourth channel buys a genuine improvement without ever reaching zero. The negative one is that the achievable residual has not moved much in three decades, because the obstruction is chemical rather than computational: the set of dyes that can be manufactured, deposited and kept stable has not changed as fast as anything else in imaging.
The commercial consequence is that camera colour has improved almost entirely through better fitting and better lookup tables — that is, through reducing the model’s contribution to the residual — while the sensor’s own contribution has stayed roughly where it was. There is a floor and the industry has been approaching it.
What this means for a photograph offered as evidence
The whole field converges on a practical question and this rung supplies its numeric half, so it is worth stating the answer here even though the closing essay develops it.
A photograph’s colour is trustworthy to about a ΔE of one on ordinary surfaces under ordinary light, given a good profile and a correctly estimated white point, and to no better than about three on saturated ones — before any of the other decisions in the pipeline are counted. Add an uncertain white balance and the figure becomes much worse and much more variable, because a white-balance error moves everything at once and can reach tens of degrees of angular error on an awkward scene.
So the honest summary is that a photograph supports claims of the form these two things in this frame differ considerably better than claims of the form this thing is that colour. The first cancels most of the pipeline; the second inherits all of it.
That is not a small conclusion for a technology that is routinely used the other way round — to establish, from a supplied image, that a delivered item does or does not match a specification. The instrument in that transaction has an unpublished kernel, a guessed illuminant, a fitted matrix and a tone curve chosen for pleasingness, and none of those is recorded in what arrives.
Where the ladder goes next
Downward, this rung sits on a camera profile is a fit, which is about the choice, and on how far apart are two colours, which is where ΔE became a quantity with an argument behind it.
Upward, the field turns to the stages after the matrix. Correcting colour costs noise is the price of the off-diagonal terms this essay has been optimising, and a photograph is not a measurement collects the whole field’s answer to what a photograph establishes.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- One row for every lamp costs the lamps that lose least camera raw · colour matrix · least-squares · spectral sensitivity
- The chart was measured by an observer too camera raw · colour matrix · least-squares · luther condition
- A black level is multiplied by the balance camera raw · colour matrix · spectral sensitivity
- A camera cannot record the excitation camera raw · colour matrix · spectral sensitivity
- A colour has a name ciede2000 · δe · gamut
- A matrix is fitted under one light camera raw · colour matrix · the icc profile
What links here
The 8 essays that link to this one and share the most of its objects, of 19 that link here.
The objects this essay names
Each one links to every other essay that touches it.
Camera rawCIEDE2000Colour matrixΔEGamutThe ICC profileLeast-squaresLuther conditionSaturationSpectral sensitivity