What a camera does

The objective nobody chose

Every camera profile here, and as far as can be told everywhere, is a linear least-squares solve in tristimulus space. That is an objective and it is on nobody's menu — it weights an error by how bright the patch is. Refitting the same matrix to minimise a real colour-difference formula improves the fit in all six, and moves the matrix, which means different pixels rather than a different report.

Assumes A camera profile is a fit, The coincidence was a mechanism and A choice with no magnitude.

A camera profile is a 3×3 matrix, and there is exactly one way anybody computes it: least squares from raw to tristimulus. That is a choice of objective, and it is a choice nobody defends because nobody notices making it.

A camera matrix refitted to minimise each unit, rather than solved in XYZ. Every camera profile here, and as far as can be told every camera profile anybody ships, is a linear least-squares solve in XYZ. That is an objective and it is on nobody's menu: it weights a difference by how large the tristimulus values are. Each row here refits the same 3×3 by direct search to minimise one of the six units instead. The upper bar is how much better the fit gets in that unit; the lower is how far the matrix itself moves, as a relative Frobenius norm. Both matter and they do not agree: CAM16-UCS moves the matrix least, at 0.34 per cent, for the largest improvement of the six, while ΔE*94 moves it 6.8 times as far for less. A score that changes is a report changing; a matrix that changes is the camera rendering different pixels.
Fig. 1 The same camera matrix refitted by direct search to minimise each of six colour-difference formulae, rather than solved in tristimulus space. The upper bar is how much better the fit gets in that unit; the lower is how far the matrix itself moves.

The claim

A camera profile is fitted in an objective that is on nobody’s menu, and fitting it in one that is improves the result in every unit and moves the matrix.

  • Least squares in tristimulus space is a weighting. Minimising the sum of squared XYZ errors weights each patch by the size of its tristimulus values, so a bright yellow counts for several times a dark blue.
  • Nobody chose it. It is on no list of colour-difference formulae, no standard recommends it for this, and it is used because it has a closed-form solution.
  • Refitting improves every unit, by between 4.3 and 10.8 per cent on the set it is fitted to and by between 1.4 and 5.2 per cent on a held-out set.
  • And the matrix moves, by between 0.34 and 2.33 per cent in relative Frobenius norm. A score that changes is a report changing; a matrix that changes is the camera rendering different pixels.
  • The two do not agree about which unit is the outlier. CAM16-UCS moves the matrix least and gains most; ΔE*94 moves it seven times as far for half the gain.

What the standard construction is

The arithmetic is short enough to state completely, which is part of why it is universal.

A sensor gives three numbers per patch. A colorimeter gives three numbers per patch. Stack the sensor’s readings for n patches into a 3 × n matrix X and the tristimulus values into a 3 × n matrix Y, and solve M = Y Xᵀ (X Xᵀ)⁻¹. That is the matrix minimising Σ ‖M xᵢ − yᵢ‖², it is one line of linear algebra, and it is what every profile in this collection uses.

The objective is the sum of squared tristimulus errors. Which is a weighting, and an odd one:

It weights by brightness. A patch at Y = 80 contributes errors sixteen times the size of an otherwise identical patch at Y = 20, because the errors scale with the values. So a light patch dominates the fit.

It weights the three axes equally. X, Y and Z have comparable magnitudes but very different perceptual significance; an error of one unit in Y is a lightness error and one unit in Z is a blue-yellow error, and nothing about a person makes those equivalent.

And it is linear. No cube root, no chroma weighting, no compression of any kind — every one of the operations that the last twenty-five years of colour difference has been about is absent.

Nobody would propose it as a colour-difference formula. It is used because (X Xᵀ)⁻¹ exists and a nine-parameter search does not have to be run.

Refitting

Replace the objective with each of the six units in turn and search over the nine matrix entries from the least-squares answer, using the collection’s own simplex. The fit set is twelve desaturated reflectances and the test set is twelve saturated ones, which is the split that makes the profile a measurement of the camera rather than of the fit.

unit least squares, fit refitted, fit least squares, test refitted, test matrix moves
ΔE*ab 0.405 0.381 0.954 0.928 1.17%
ΔE*94 0.550 0.518 1.035 1.011 2.33%
ΔE2000 0.758 0.725 1.184 1.123 1.29%
ΔE*uv 0.424 0.399 0.997 0.970 0.60%
Oklab 0.394 0.374 0.947 0.933 1.52%
CAM16-UCS 1.516 1.353 2.180 2.122 0.34%

Three readings, in order of how surprising each is.

The least-squares matrix is beaten in every unit, which is not surprising at all — it is what fitting means, and a matrix fitted to minimise something will minimise it better than one fitted to minimise something else.

The improvement survives on the held-out set, which is less obvious. A nine-parameter fit on twelve patches has room to overfit, and the held-out improvement being smaller than the in-sample one (1.4 to 5.2 per cent against 4.3 to 10.8) is the signature of exactly that, at a modest level. The refit is a real improvement and about half of what it looks like.

And the matrix moves by up to 2.3 per cent, which is the reading that matters.

Why a moving matrix is a different fact from a moving score

Everything else in this audit measures a report changing. A census residual, a tolerance reading, an observer gap — the underlying object is fixed and the number attached to it moves.

A camera profile is not like that. The matrix is applied to every pixel of every photograph the camera takes. A profile fitted to minimise ΔE2000 and one fitted to minimise ΔE*ab produce different images, not different opinions about the same image, and the difference is out in the world rather than in a table.

Two point three per cent of a matrix is not large and is not nothing. Applied to a saturated patch it moves the rendered colour by an amount comparable to the profile’s own error, which is to say: the choice of fitting objective is the same size as the thing the fitting is trying to remove.

A silicon sensor's best possible impersonation of the standard observerThe 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 31.7 per cent of the matching functions' own magnitude, worst at 440 nm. Colour reproduction is exact if and only if this is zero.0x̄ ȳ z̄, behind — what a colorimeter needsthe sensor's best linear fit to the matching functionsthe fit goes negative, and has toworst at 440 nmwhat is left over — residual 31.7% overall400450500550600650700750wavelength / nmmodelled silicon sensorLuther–Ives 1927, the fit and its residual
Fig. 2 The fit the refits are measured against: a silicon sensor’s raw responses carried to tristimulus by one 3×3, on the surfaces it was fitted to and on surfaces it was not. Every point off the diagonal is an error the objective was weighting by brightness.

The two rankings disagree

The last column and the improvement columns rank the units differently, and the disagreement is not decoration.

CAM16-UCS moves the matrix least — 0.34 per cent — and gains most, 10.8 per cent on the fit set. So its optimum is very close to the least-squares matrix in matrix space and very far from it in score space.

ΔE*94 moves the matrix furthest — 2.33 per cent — for a 5.7 per cent gain. Its optimum is a long way off in matrix space for half the improvement.

The two quantities answer different questions and both are needed. How much better could this profile be? is the score. How much does the objective decide the profile? is the matrix movement. A search reporting only the first would say the appearance unit is the outlier; reporting only the second would say ΔE*94 is. Neither alone is wrong and neither alone is the finding.

The mechanism is the local shape of each objective around the least-squares point. A steep, well-curved objective has an optimum close by and a large drop; a shallow, poorly-curved one has an optimum far off in a direction the score barely cares about. A large parameter movement for a small score gain is the signature of a flat valley, which is the same diagnosis this collection reached about basis searches and about the shape of a fitted quadratic near its minimum.

Which means ΔE*94’s refit should be trusted least: a matrix 2.3 per cent away for a modest gain is a matrix the objective is not confident about, and a slightly different fit set would put it somewhere else.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 3 The six objectives by how far each is from a rescaling of the published one. A profile fitted in one and reported in another is exposed to exactly this disagreement, and the standard construction is fitted in a seventh that is not on the chart.

One unit generalises and the others do not

The two improvement columns are quoted as ranges, and dividing one by the other row by row finds something a range hides. How much of each in-sample gain survives on the held-out set:

unit fit gain held-out gain retained
ΔE2000 4.35% 5.15% 1.18
ΔE*ab 5.93% 2.73% 0.46
ΔE*uv 5.90% 2.71% 0.46
ΔE*94 5.82% 2.32% 0.40
Oklab 5.08% 1.48% 0.29
CAM16-UCS 10.75% 2.66% 0.25

Five units retain between a quarter and a half of their gain, which is the ordinary picture of a nine-parameter fit on twelve patches. ΔE2000 retains more than all of it — its refit improves the unseen set by more than it improves the set it was fitted on, which is not what overfitting looks like and is not what any of the other five does.

It also inverts the ranking. On the fit set ΔE2000 is the worst of the six, gaining 4.35 per cent against the appearance unit’s 10.75; on the held-out set it is the best by a factor of two. A reader taking the in-sample column at face value would conclude that fitting in the published unit buys least of anything on the menu, and the opposite is true.

The mechanism the split suggests is the one the split was designed to expose. The fit set is twelve desaturated reflectances and the test set twelve saturated ones, so transferring between them means predicting behaviour at chroma from a fit made near neutral. ΔE2000 is the formula whose weighting depends most elaborately on where in the space a pair sits, so an optimum found under it is an optimum that already accounts for chroma — while a Euclidean objective fitted on pale patches has nothing to say about saturated ones and, on this evidence, mostly does not transfer.

And it is twelve patches, so this is a hint rather than a result. One unit out of six behaving differently on a twelve-point held-out set is well within what chance produces, and the honest form is that the ranking of these six by generalisation is not established by this table and is not the same ranking as the one it reports.

The units that move the matrix least gain the most

The two rankings are said above to disagree, and they do more than disagree — they run backwards. Taking the rank correlation between how far each unit moves the matrix and how much it gains on the fit set:

ρ=0.71\rho = -0.71

A larger movement in the parameters buys a smaller improvement in the score, across six units, monotonically enough to be worth a number.

That is the flat-valley diagnosis stated as a measurement rather than as an interpretation. If the objectives differed mainly in where their optima sat, a unit whose optimum is further away would have further to fall and would gain more. What is measured is the reverse, which means the distance is not being travelled towards anything: the objectives that move the matrix furthest are the ones whose valleys are shallowest in the direction they move it, so the search wanders a long way for very little.

The two extremes make it concrete. CAM16-UCS moves the matrix by 0.34 per cent and gains 10.75, which is a steep well with its bottom nearly where the least-squares answer already is. ΔE*94 moves it by 2.33 per cent and gains 5.82, seven times the distance for half the return.

So the movement column is not a measure of how much the objective matters. It is closer to a measure of how badly conditioned each objective is around the least-squares point, and a large number in it is a reason to distrust that row’s refit rather than a reason to take it seriously. The essay’s own recommendation is unaffected — fitting in any real formula beats fitting in none — but the ranking of the six by how far they move the matrix should be read as an inverse measure of confidence.

Where the seventh objective would sit

Least squares in tristimulus space is not on the menu, and the natural question is where it would go if it were.

It cannot be placed on the same axis, because it is not a distance between two colours at all — it is a distance between two vectors, with no reference to a white, no lightness compression, and no dependence on where in the space the pair sits. Every unit on the menu is at least a function of a pair and a white; this is a function of a pair.

What can be measured is the consequence, and it is in the table. The least-squares matrix is beaten by every one of the six on its own terms, and beaten by more than the six differ from each other. In ΔE2000 the gap between the least-squares matrix and the ΔE2000-optimal matrix is 0.033 on the fit set, which is larger than the gap between the ΔE2000-optimal and ΔE*94-optimal matrices scored in ΔE2000.

So the practical ordering is clear even if the placement is not: fitting in any real colour-difference formula beats fitting in none by more than the choice among them costs. That is the one recommendation this essay makes, and it costs a nine-parameter search per profile — a fraction of a second — against a closed-form solve.

What the brightness weighting does to a chart

The objective’s oddest property — weighting by tristimulus magnitude — has a consequence for chart design that is worth separating out, because it interacts with the chart choice rather than being independent of it.

A standard test chart has patches at a range of lightnesses, and the dark ones are usually the ones a camera gets worst: at low signal the sensor’s own noise and any black-level error matter most, and a small tristimulus error there is a large perceptual one because lightness is a cube root of luminance. So the patches the objective weights least are the patches the camera fails on worst.

Measured on the collection’s own chart, the least-squares fit’s error on the darkest third of the patches is 2.04 times its error on the lightest third when read in ΔE2000, and 1.12 times when read as a raw tristimulus distance. So the objective sees the two groups as almost equally well fitted and the report sees a factor of two between them. The objective and the report disagree about which patches the profile is failing on, which is a more specific complaint than “the objective is wrong”.

The refit under ΔE2000 moves weight accordingly, and the movement is visible in the matrix rather than only in the score: the Z row changes by 1.40 per cent against the Y row’s 0.84, and the green and blue columns by 1.46 and 1.32 per cent against the red column’s 0.34. A closed-form solve cannot make that trade, because it has no way to express it.

Why nobody does it

Three reasons, and only the first is good.

The closed form is exact and instantaneous, and a direct search is neither. A profile computed a thousand times in a factory calibration line is worth having in closed form.

The improvement is small in absolute terms. 0.033 ΔE2000 on a fit whose error is 0.758, against a camera whose errors on real scenes are dominated by things a matrix cannot fix at all. A supplier optimising the wrong four per cent is a familiar shape.

And the objective is invisible. A closed-form solution does not present itself as a choice. There is no parameter to set, no line in a configuration file, nothing to name in a specification — which is the definition of the kind of decision this collection exists to find. The chart is chosen, the illuminant is chosen, the reporting unit is chosen and argued about; the objective is a property of the algebra somebody wrote in 1980.

Two sensitivities from two libraries, under every unit. Two quantities that share no code, no test set and no physical question: how much the adaptation census's residual depends on how saturated its surfaces are, and how much a camera profile's reported error depends on how saturated its test chart is. The first is a mean over fourteen changes of light built from cosine combinations; the second is one number about one silicon sensor scored on Gaussian bumps. Under the published unit they sit at 0.687 and 0.656. Across the whole menu they move together, from about 0.5 under the appearance unit to about 1.15 under plain CIELAB, staying within 12 per cent of each other at the worst point. Two numbers agreeing once is a coincidence; two curves agreeing at six points across a factor of two and a half is a shared mechanism, and the mechanism is the compression the unit applies to a chroma difference.
Fig. 4 The same camera profile’s sensitivity to its chart, across the menu, tracking the adaptation census’s sensitivity to its test set. The objective is a third thing the profile depends on and the only one with no name.
Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu.
Fig. 5 Where on the scale the six units disagree with the published one. A camera profile’s errors sit around one to two units, in the band where the proportional disagreement is still 15 to 57 per cent — so the objective and the report are being taken at the noisiest part of the curve.

Both of those are statements about the scale the objective is chosen on. The third is about its size relative to everything else the profile rests on, and it is the one that decides whether any of this is worth a reader’s attention.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 6 The camera profile’s error is the second row here and has the largest spread across the menu of the six quantities, at a factor of 2.30 — before the objective it was fitted under is varied at all.

A single ellipse is where the objective’s choice of unit stops being an average and becomes a shape somebody could disagree with.

One of MacAdam's ellipses as each unit sees it, at x 0.596, y 0.283. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 1.15; the least round is ΔE00 at 2.87. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 7 One measured ellipse in each of the six units, every outline scaled to its own mean radius so only shape is compared. Which unit the objective is stated in decides which of these outlines the answer is about.

Where the model stops

Twelve patches is a small fit set and nine parameters is a lot to fit on it, which is why the held-out column is in the table and why the honest improvement is the held-out one. A real profile is fitted on twenty-four or on a hundred and forty, and the overfitting would be correspondingly less.

The search is a simplex from the least-squares point with three restarts, and a colour-difference objective over nine parameters is not convex. Nothing here proves the optima found are global, and the matrix movements are lower bounds on the distance to the true optimum. A search that stopped early would understate the movement and overstate nothing, so the direction of the finding is safe.

And a real profile is not a 3×3. It is a matrix, a tone curve, and usually a lookup table, and an interpolation between the table’s nodes whose error is a separate question. The objective question applies to each stage and this essay reaches one of them.

Who found it, and when

The least-squares camera matrix is old enough that no attribution is available; it is what anybody does with an overdetermined linear system, and it predates colour management.

That fitting in a perceptual metric would be better has been proposed periodically — the phrase in the literature is perceptual optimisation of colour correction matrices, and there are papers going back to the 1990s, mostly reporting improvements of the size found here and mostly not adopted. What is not usually reported is the matrix movement, and the reason is the same one that makes it worth reporting: a paper on profile optimisation reports the score it optimised, because that is the result. The distance the parameters travelled is a diagnostic rather than a result, and it is the number that says whether the answer is stable.

A sensor built to satisfy the Luther condition, impersonating it exactly. The 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 0.0 per cent of the matching functions' own magnitude, worst at 530 nm. Colour reproduction is exact if and only if this is zero.
Fig. 8 The control: a sensor whose sensitivities are a linear combination of the colour-matching functions. For it the fit is exact in every objective and every unit, and every refit moves the matrix by nothing — which is what says the movements above are the mismatch rather than the search.

Where the ladder goes next

One quantity in the audit’s inventory is native to the appearance unit rather than to the published one: an observer part-way through adapting to a new room, which comes out of a model that defines its own difference formula. Measuring it in a matching unit means asking what tristimulus a settled observer would need to be shown to give the same report — which requires pretending an appearance was a stimulus, and that step is the whole subject.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationCamera profileColour differenceFittingLeast-squaresLuther conditionObjectiveOptimisationTest setTristimulus