The objective nobody chose
Assumes A camera profile is a fit, The coincidence was a mechanism and A choice with no magnitude.
A camera profile is a 3×3 matrix, and there is exactly one way anybody computes it: least squares from raw to tristimulus. That is a choice of objective, and it is a choice nobody defends because nobody notices making it.
The claim
A camera profile is fitted in an objective that is on nobody’s menu, and fitting it in one that is improves the result in every unit and moves the matrix.
- Least squares in tristimulus space is a weighting. Minimising the sum of squared XYZ errors weights each patch by the size of its tristimulus values, so a bright yellow counts for several times a dark blue.
- Nobody chose it. It is on no list of colour-difference formulae, no standard recommends it for this, and it is used because it has a closed-form solution.
- Refitting improves every unit, by between 4.3 and 10.8 per cent on the set it is fitted to and by between 1.4 and 5.2 per cent on a held-out set.
- And the matrix moves, by between 0.34 and 2.33 per cent in relative Frobenius norm. A score that changes is a report changing; a matrix that changes is the camera rendering different pixels.
- The two do not agree about which unit is the outlier. CAM16-UCS moves the matrix least and gains most; ΔE*94 moves it seven times as far for half the gain.
What the standard construction is
The arithmetic is short enough to state completely, which is part of why it is universal.
A sensor gives three numbers per patch. A colorimeter gives three numbers per patch. Stack the sensor’s readings for n patches into a 3 × n matrix X and the tristimulus values into a 3 × n matrix Y, and solve M = Y Xᵀ (X Xᵀ)⁻¹. That is the matrix minimising Σ ‖M xᵢ − yᵢ‖², it is one line of linear algebra, and it is what every profile in this collection uses.
The objective is the sum of squared tristimulus errors. Which is a weighting, and an odd one:
It weights by brightness. A patch at Y = 80 contributes errors sixteen times the size of an otherwise identical patch at Y = 20, because the errors scale with the values. So a light patch dominates the fit.
It weights the three axes equally. X, Y and Z have comparable magnitudes but very different perceptual significance; an error of one unit in Y is a lightness error and one unit in Z is a blue-yellow error, and nothing about a person makes those equivalent.
And it is linear. No cube root, no chroma weighting, no compression of any kind — every one of the operations that the last twenty-five years of colour difference has been about is absent.
Nobody would propose it as a colour-difference formula. It is used because (X Xᵀ)⁻¹ exists and a nine-parameter search does not have to be run.
Refitting
Replace the objective with each of the six units in turn and search over the nine matrix entries from the least-squares answer, using the collection’s own simplex. The fit set is twelve desaturated reflectances and the test set is twelve saturated ones, which is the split that makes the profile a measurement of the camera rather than of the fit.
| unit | least squares, fit | refitted, fit | least squares, test | refitted, test | matrix moves |
|---|---|---|---|---|---|
| ΔE*ab | 0.405 | 0.381 | 0.954 | 0.928 | 1.17% |
| ΔE*94 | 0.550 | 0.518 | 1.035 | 1.011 | 2.33% |
| ΔE2000 | 0.758 | 0.725 | 1.184 | 1.123 | 1.29% |
| ΔE*uv | 0.424 | 0.399 | 0.997 | 0.970 | 0.60% |
| Oklab | 0.394 | 0.374 | 0.947 | 0.933 | 1.52% |
| CAM16-UCS | 1.516 | 1.353 | 2.180 | 2.122 | 0.34% |
Three readings, in order of how surprising each is.
The least-squares matrix is beaten in every unit, which is not surprising at all — it is what fitting means, and a matrix fitted to minimise something will minimise it better than one fitted to minimise something else.
The improvement survives on the held-out set, which is less obvious. A nine-parameter fit on twelve patches has room to overfit, and the held-out improvement being smaller than the in-sample one (1.4 to 5.2 per cent against 4.3 to 10.8) is the signature of exactly that, at a modest level. The refit is a real improvement and about half of what it looks like.
And the matrix moves by up to 2.3 per cent, which is the reading that matters.
Why a moving matrix is a different fact from a moving score
Everything else in this audit measures a report changing. A census residual, a tolerance reading, an observer gap — the underlying object is fixed and the number attached to it moves.
A camera profile is not like that. The matrix is applied to every pixel of every photograph the camera takes. A profile fitted to minimise ΔE2000 and one fitted to minimise ΔE*ab produce different images, not different opinions about the same image, and the difference is out in the world rather than in a table.
Two point three per cent of a matrix is not large and is not nothing. Applied to a saturated patch it moves the rendered colour by an amount comparable to the profile’s own error, which is to say: the choice of fitting objective is the same size as the thing the fitting is trying to remove.
The two rankings disagree
The last column and the improvement columns rank the units differently, and the disagreement is not decoration.
CAM16-UCS moves the matrix least — 0.34 per cent — and gains most, 10.8 per cent on the fit set. So its optimum is very close to the least-squares matrix in matrix space and very far from it in score space.
ΔE*94 moves the matrix furthest — 2.33 per cent — for a 5.7 per cent gain. Its optimum is a long way off in matrix space for half the improvement.
The two quantities answer different questions and both are needed. How much better could this profile be? is the score. How much does the objective decide the profile? is the matrix movement. A search reporting only the first would say the appearance unit is the outlier; reporting only the second would say ΔE*94 is. Neither alone is wrong and neither alone is the finding.
The mechanism is the local shape of each objective around the least-squares point. A steep, well-curved objective has an optimum close by and a large drop; a shallow, poorly-curved one has an optimum far off in a direction the score barely cares about. A large parameter movement for a small score gain is the signature of a flat valley, which is the same diagnosis this collection reached about basis searches and about the shape of a fitted quadratic near its minimum.
Which means ΔE*94’s refit should be trusted least: a matrix 2.3 per cent away for a modest gain is a matrix the objective is not confident about, and a slightly different fit set would put it somewhere else.
One unit generalises and the others do not
The two improvement columns are quoted as ranges, and dividing one by the other row by row finds something a range hides. How much of each in-sample gain survives on the held-out set:
| unit | fit gain | held-out gain | retained |
|---|---|---|---|
| ΔE2000 | 4.35% | 5.15% | 1.18 |
| ΔE*ab | 5.93% | 2.73% | 0.46 |
| ΔE*uv | 5.90% | 2.71% | 0.46 |
| ΔE*94 | 5.82% | 2.32% | 0.40 |
| Oklab | 5.08% | 1.48% | 0.29 |
| CAM16-UCS | 10.75% | 2.66% | 0.25 |
Five units retain between a quarter and a half of their gain, which is the ordinary picture of a nine-parameter fit on twelve patches. ΔE2000 retains more than all of it — its refit improves the unseen set by more than it improves the set it was fitted on, which is not what overfitting looks like and is not what any of the other five does.
It also inverts the ranking. On the fit set ΔE2000 is the worst of the six, gaining 4.35 per cent against the appearance unit’s 10.75; on the held-out set it is the best by a factor of two. A reader taking the in-sample column at face value would conclude that fitting in the published unit buys least of anything on the menu, and the opposite is true.
The mechanism the split suggests is the one the split was designed to expose. The fit set is twelve desaturated reflectances and the test set twelve saturated ones, so transferring between them means predicting behaviour at chroma from a fit made near neutral. ΔE2000 is the formula whose weighting depends most elaborately on where in the space a pair sits, so an optimum found under it is an optimum that already accounts for chroma — while a Euclidean objective fitted on pale patches has nothing to say about saturated ones and, on this evidence, mostly does not transfer.
And it is twelve patches, so this is a hint rather than a result. One unit out of six behaving differently on a twelve-point held-out set is well within what chance produces, and the honest form is that the ranking of these six by generalisation is not established by this table and is not the same ranking as the one it reports.
The units that move the matrix least gain the most
The two rankings are said above to disagree, and they do more than disagree — they run backwards. Taking the rank correlation between how far each unit moves the matrix and how much it gains on the fit set:
A larger movement in the parameters buys a smaller improvement in the score, across six units, monotonically enough to be worth a number.
That is the flat-valley diagnosis stated as a measurement rather than as an interpretation. If the objectives differed mainly in where their optima sat, a unit whose optimum is further away would have further to fall and would gain more. What is measured is the reverse, which means the distance is not being travelled towards anything: the objectives that move the matrix furthest are the ones whose valleys are shallowest in the direction they move it, so the search wanders a long way for very little.
The two extremes make it concrete. CAM16-UCS moves the matrix by 0.34 per cent and gains 10.75, which is a steep well with its bottom nearly where the least-squares answer already is. ΔE*94 moves it by 2.33 per cent and gains 5.82, seven times the distance for half the return.
So the movement column is not a measure of how much the objective matters. It is closer to a measure of how badly conditioned each objective is around the least-squares point, and a large number in it is a reason to distrust that row’s refit rather than a reason to take it seriously. The essay’s own recommendation is unaffected — fitting in any real formula beats fitting in none — but the ranking of the six by how far they move the matrix should be read as an inverse measure of confidence.
Where the seventh objective would sit
Least squares in tristimulus space is not on the menu, and the natural question is where it would go if it were.
It cannot be placed on the same axis, because it is not a distance between two colours at all — it is a distance between two vectors, with no reference to a white, no lightness compression, and no dependence on where in the space the pair sits. Every unit on the menu is at least a function of a pair and a white; this is a function of a pair.
What can be measured is the consequence, and it is in the table. The least-squares matrix is beaten by every one of the six on its own terms, and beaten by more than the six differ from each other. In ΔE2000 the gap between the least-squares matrix and the ΔE2000-optimal matrix is 0.033 on the fit set, which is larger than the gap between the ΔE2000-optimal and ΔE*94-optimal matrices scored in ΔE2000.
So the practical ordering is clear even if the placement is not: fitting in any real colour-difference formula beats fitting in none by more than the choice among them costs. That is the one recommendation this essay makes, and it costs a nine-parameter search per profile — a fraction of a second — against a closed-form solve.
What the brightness weighting does to a chart
The objective’s oddest property — weighting by tristimulus magnitude — has a consequence for chart design that is worth separating out, because it interacts with the chart choice rather than being independent of it.
A standard test chart has patches at a range of lightnesses, and the dark ones are usually the ones a camera gets worst: at low signal the sensor’s own noise and any black-level error matter most, and a small tristimulus error there is a large perceptual one because lightness is a cube root of luminance. So the patches the objective weights least are the patches the camera fails on worst.
Measured on the collection’s own chart, the least-squares fit’s error on the darkest third of the patches is 2.04 times its error on the lightest third when read in ΔE2000, and 1.12 times when read as a raw tristimulus distance. So the objective sees the two groups as almost equally well fitted and the report sees a factor of two between them. The objective and the report disagree about which patches the profile is failing on, which is a more specific complaint than “the objective is wrong”.
The refit under ΔE2000 moves weight accordingly, and the movement is visible in the matrix rather than only in the score: the Z row changes by 1.40 per cent against the Y row’s 0.84, and the green and blue columns by 1.46 and 1.32 per cent against the red column’s 0.34. A closed-form solve cannot make that trade, because it has no way to express it.
Why nobody does it
Three reasons, and only the first is good.
The closed form is exact and instantaneous, and a direct search is neither. A profile computed a thousand times in a factory calibration line is worth having in closed form.
The improvement is small in absolute terms. 0.033 ΔE2000 on a fit whose error is 0.758, against a camera whose errors on real scenes are dominated by things a matrix cannot fix at all. A supplier optimising the wrong four per cent is a familiar shape.
And the objective is invisible. A closed-form solution does not present itself as a choice. There is no parameter to set, no line in a configuration file, nothing to name in a specification — which is the definition of the kind of decision this collection exists to find. The chart is chosen, the illuminant is chosen, the reporting unit is chosen and argued about; the objective is a property of the algebra somebody wrote in 1980.
Both of those are statements about the scale the objective is chosen on. The third is about its size relative to everything else the profile rests on, and it is the one that decides whether any of this is worth a reader’s attention.
A single ellipse is where the objective’s choice of unit stops being an average and becomes a shape somebody could disagree with.
Where the model stops
Twelve patches is a small fit set and nine parameters is a lot to fit on it, which is why the held-out column is in the table and why the honest improvement is the held-out one. A real profile is fitted on twenty-four or on a hundred and forty, and the overfitting would be correspondingly less.
The search is a simplex from the least-squares point with three restarts, and a colour-difference objective over nine parameters is not convex. Nothing here proves the optima found are global, and the matrix movements are lower bounds on the distance to the true optimum. A search that stopped early would understate the movement and overstate nothing, so the direction of the finding is safe.
And a real profile is not a 3×3. It is a matrix, a tone curve, and usually a lookup table, and an interpolation between the table’s nodes whose error is a separate question. The objective question applies to each stage and this essay reaches one of them.
Who found it, and when
The least-squares camera matrix is old enough that no attribution is available; it is what anybody does with an overdetermined linear system, and it predates colour management.
That fitting in a perceptual metric would be better has been proposed periodically — the phrase in the literature is perceptual optimisation of colour correction matrices, and there are papers going back to the 1990s, mostly reporting improvements of the size found here and mostly not adopted. What is not usually reported is the matrix movement, and the reason is the same one that makes it worth reporting: a paper on profile optimisation reports the score it optimised, because that is the result. The distance the parameters travelled is a diagnostic rather than a result, and it is the number that says whether the answer is stable.
Where the ladder goes next
One quantity in the audit’s inventory is native to the appearance unit rather than to the published one: an observer part-way through adapting to a new room, which comes out of a model that defines its own difference formula. Measuring it in a matching unit means asking what tristimulus a settled observer would need to be shown to give the same report — which requires pretending an appearance was a stimulus, and that step is the whole subject.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A chart decides what a camera scores camera profile · colour difference · least-squares · luther condition · test set
- The chart was measured by an observer too calibration · camera profile · least-squares · luther condition
- The observers differ by a unit's worth calibration · colour difference · test set · tristimulus
- A dial through a discrete menu calibration · colour difference · test set
- A sensor has no lens calibration · camera profile · luther condition
- Saturation is nearly everything camera profile · colour difference · test set
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CalibrationCamera profileColour differenceFittingLeast-squaresLuther conditionObjectiveOptimisationTest setTristimulus