What the brain does

One matrix doing two jobs

CIECAM16 adapts in CAT16 and then applies its response compression in the same axes, so a single matrix decides both how well the model handles a change of light and how uniform the space it produces is. The two jobs have different best answers, and the matrix was chosen against only one of them.

Assumes No basis is good at both, The cones an appearance model uses and A difference needs a basis too.

The reason CAT16 and CIECAM16 share a number is that they share a matrix. The model’s first step is a von Kries gain in three cone-like axes, and those axes are CAT16’s; the model’s next step is a compressive nonlinearity applied in the same three axes, and everything downstream — lightness, chroma, hue, and the uniform space differences are measured in — is computed after it. One 3×3 therefore decides two things the model is judged on, and the committee that chose it was looking at one.

How far from circles every basis leaves the ellipsesEight bases ranked on the mean ratio of the long to the short axis of MacAdam's twenty-five discrimination ellipses, measured in a lightness–chroma space built on that basis. The range runs from 1.61 for best for discrimination to 7.70 for best for adaptation. The ordering is not the ordering on the other objective and is nearly its reverse.best for discrimination1.61from the confusion points2.60CAT162.71Hunt–Pointer–Estévez2.81XYZ scaling3.44Bradford3.68CAT024.29best for adaptation7.70shorter is a better space to measure a difference inCIE 1931 2° observer · two objectives, one basis
Fig. 1 Eight bases ranked by how nearly the discrimination ellipses come out as circles in a space built on each. CAT16 is best of the published set here — which is not what it was selected for.

The claim

CAT16 is doing two jobs and was scored on one. It leaves 1.31 ΔE00 after a gain, against a floor of 0.97, and an ellipse axis ratio of 2.71, against a floor of 1.61. Freed to use its own matrix, each job improves by 26 and 40 per cent respectively — and the two matrices that do so are nothing like each other.

  • The adaptation job: CAT16 is second-best of the published transforms, behind Bradford’s 1.14.
  • The uniformity job: CAT16 is best of the published transforms, ahead of the receptors’ 2.60 by a whisker and ahead of Bradford’s 3.68 by a great deal.
  • Being good at both was not the criterion. CAT16 was fitted to corresponding-colour data, which measures the first job only.
  • So the second result is luck, and it is worth saying so, because a model that is accidentally well placed is a model that could as easily have been badly placed.
  • Bradford, in the same slot, would have been. A pipeline that adapts with Bradford and then measures differences in Bradford’s axes gets 3.68.

The two steps, and the one matrix

CIECAM16’s forward path is a sequence and the order matters. Tristimulus values enter, are carried into the CAT16 axes, and are multiplied there by a per-channel gain that depends on the adopting white and on the degree of adaptation. That is the adaptation step and it is exactly the CAT16 transform, which is why the two objects have one name.

What happens immediately afterwards is a compressive nonlinearity — a saturating function of each channel, with the adapting luminance in it — and then the achromatic and opponent signals are formed from the compressed values. Lightness, brightness, chroma, colourfulness, saturation and hue are all read out of those. The uniform space that colour-difference formulae are built on is a rescaling of three of them.

So the same three directions that decide whether a white balance behaves also decide what a small step in colour costs, because a nonlinearity applied per channel in a basis is a different function for every basis. Nothing commutes past a nonlinearity, which is the whole reason the second job has a basis at all.

Every basis against both objectives at once. A scatter with the mean adaptation residual across the illumination census on the horizontal axis and the mean axis ratio of MacAdam's ellipses in a lightness–chroma space on the vertical. Lower is better on both. The two winners sit at the two ends of an empty diagonal: the basis that adapts best leaves 7.70 on the vertical and the basis that discriminates best leaves 1.79 on the horizontal, each worse on the other objective than every published transform. The basis built from the dichromat confusion points is at (1.65, 2.60) — best at neither and within a factor of two of both floors, which no other entry in the picture manages.
Fig. 2 Both jobs at once. CAT16 sits close to the middle-left of the picture, which is the best position any published transform occupies — and no candidate is near the bottom-left corner.

How much of the model the choice reaches

It is worth being concrete about how far downstream the matrix’s influence runs, because the answer is all the way.

The compression is applied per channel in the CAT16 axes. The achromatic signal is a weighted sum of the three compressed values; the two opponent signals are differences of them. Lightness is a power of the achromatic signal over its value for the white. Chroma is built from the opponent pair and lightness. Hue is the angle of the opponent pair. Colourfulness and saturation are chroma rescaled by the adapting luminance and by the achromatic signal.

So every correlate the model reports is a function of three numbers that were compressed in these axes, and there is no later stage that could undo the choice — a nonlinearity has already happened. The uniform space CAM16-UCS is a further rescaling of lightness and chroma, which changes the radial geometry and cannot change the angular geometry at all.

That is why the ellipse measurement is a fair proxy despite using a cube root instead of the model’s own compression. What is being varied is the thing the whole downstream chain is built on, and the exponent turns out to be nearly irrelevant beside it.

What each job would prefer

Minimise the adaptation residual over all nine numbers and the answer is a set of narrow, nearly independent bands at 0.974. Minimise the ellipse anisotropy and the answer is a set of broad overlapping ones at 1.620. Neither is anything like the other, and each is a disaster on the other’s measure.

CAT16’s position between them is what makes it usable, and the size of the gaps is what makes the position a decision rather than a detail:

  • On adaptation, CAT16 is 35 per cent above the floor.
  • On uniformity, CAT16 is 66 per cent above the floor.

The second number is larger, which is the mildly uncomfortable finding. The job the model was not optimised for is the one where more is being left on the table — and it is the job with by far the wider public consequences, because a difference formula is what a print tolerance, a quality-control limit and a gamut-mapping algorithm are all written in terms of.

How badly the ellipses fail to be circles, and how badly they must. Two failures per diagram: the mean axis ratio, which is whether a step of a given size means the same thing in every direction, and the size spread, which is whether it means the same thing everywhere. The last two rows are not published diagrams — they are the best any choice of primaries can do, found by search over all nine coefficients. The shape failure cannot be taken below 2.02, so it is a property of the eye rather than of anybody's coordinates.
Fig. 3 The bases this collection measures uniformity on, tabulated. CAT16’s row is the best of the published set and was chosen against a different criterion.

Why the accident is a happy one

CAT16 came out of a re-derivation intended to fix a specific defect: CAT02 could produce negative channel values on saturated colours near the gamut boundary, which broke the inverse model and made the transform unusable in exactly the cases gamut mapping cares about. The remedy was a less aggressively sharpened matrix — entries closer to the receptors, less extreme in the off-diagonal terms.

That remedy is why CAT16 is the best published transform on the uniformity job. Sharpening is what the adaptation objective rewards and what the uniformity objective punishes, so backing off the sharpening to fix a negativity problem moved the matrix towards the second objective’s optimum without anybody scoring it there.

How cone-like a basis is, against how well it adapts. Each basis placed by how far its own implied deuteranope confusion point falls from the measured one (horizontal) and by how much an adapted observer is left with in it (vertical). The construction from the confusion points sits at zero on the horizontal by definition and near the top on the vertical. Nothing near the left of the picture is near the bottom: the closer a basis is to the receptors, the more a von Kries gain leaves behind. The unconstrained winner sits at 1.63 on the horizontal, further from the measurement than any published transform except CAT02 and Bradford.
Fig. 4 Distance from the measured deuteranope confusion point against the adaptation residual. CAT02 is the most sharpened entry and the furthest from the receptors; CAT16 sits between it and Hunt–Pointer–Estévez.

Two more views of the same eight bases say what the trade looks like when it is walked rather than plotted.

How much every basis leaves an adapted observer. Eight bases ranked on the mean ΔE00 an adapted observer is left with after the gain, averaged over every change of illumination this collection models. The range runs from 0.97 for best for adaptation to 2.37 for XYZ scaling. The ordering is not the ordering on the other objective and is nearly its reverse.
Fig. 5 The eight ranked on the first job. The transform that wins here is not the one that wins the second, which is the whole of what “one matrix doing two jobs” costs.
Walking from the receptors to the best adaptation basis. Two curves along a straight path in the nine coefficients, from the basis built out of the dichromat confusion points at the left to the basis that minimises the adaptation residual at the right, with every row renormalised on the white so that each stop is a legitimate basis rather than a blend of two pictures. The adaptation residual falls from 1.65 to 0.97 and the ellipse anisotropy rises from 2.60 to 7.70. There is no stop where both are good and no kink where a compromise would sit.
Fig. 6 And a straight path in the nine coefficients from the receptor basis to the adaptation optimum. Every point on it is a candidate matrix, and the second job’s score changes along it too.

The measurement bears that out. CAT02 implies a deuteranope confusion point 4.09 away from the measured one; CAT16, 1.45. On the uniformity job CAT02 leaves 4.29 and CAT16 leaves 2.71. The change that was made for numerical robustness bought a 37 per cent improvement in a quantity nobody was measuring.

What a pipeline built the other way would look like

The counterfactual is not hypothetical, because a colour-managed pipeline routinely uses Bradford for white-point adaptation and then computes differences in CIELAB. Those are two different matrices — Bradford’s and XYZ — so the pipeline is not making the mistake this essay describes; it is making a different one, of using a well-adapted intermediate and then measuring in axes chosen in 1931 for non-negativity.

But a pipeline that did what CIECAM16 does with Bradford’s matrix in the slot would score 1.14 on the first job and 3.68 on the second — worse at measuring differences than CIELAB’s own basis at 3.44, and more than twice the floor. Nothing would announce it. Every white balance would be slightly better and every tolerance would be measured in a space where a step in one direction is 3.7 times as visible as a step in another.

That is the shape of the hazard: the second job is silent. A poor adaptation basis shows up as a visible cast in a rendered image. A poor uniformity basis shows up as tolerances that are too tight in some directions and too loose in others, which looks like a difficult product rather than a bad space.

What was computed, and how

Both scores are computed by machinery this collection already had, and neither was written for CAT16.

The adaptation residual takes the exact 3×3 relating tristimulus values under two lights, applies the diagonal an adapted observer would use in the candidate basis, and reports the mean CIEDE2000 remaining over a hundred and twenty-five constructed surfaces and fourteen changes of light.

The uniformity score carries the boundary points of MacAdam’s twenty-five measured ellipses through the candidate basis, applies CIELAB’s arithmetic in those axes, and reports the mean of the longest radius over the shortest. Using CIELAB’s cube root rather than CIECAM16’s own saturating compression is a simplification and a deliberate one: the compression’s shape is a second variable, and changing it turns out to matter very little compared with changing the basis, so holding it fixed isolates the thing under test.

Both floors are Nelder–Mead searches over the same nine numbers from the same start with restarts, and both are checked at twice the budget.

What the second job is worth, in units somebody uses

An axis ratio is an abstraction until it is priced, and it can be.

A colour tolerance is a statement that two things are close enough, and it is written as a single number — one ΔE00, or two, or five, depending on the trade. What a single number means geometrically is a sphere: everything within that distance passes. If the space the distance is computed in has an axis ratio of 2.71, then the set of stimuli that actually look equally different is an ellipsoid nearly three times longer one way than another, and the sphere is either inscribed in it or circumscribed about it or somewhere between.

Either way the tolerance is wrong in a direction. Inscribed, it rejects things nobody could tell apart; circumscribed, it accepts things anybody can. A tolerance is a shape before it is a number, and the axis ratio is how far from round that shape is.

Going from 2.71 to 1.61 would not make the shape round — nothing does — but it would halve the departure. On a production line running a one-unit tolerance, that is the difference between a criterion that is between two and three times too strict in its worst direction and one that is around 1.6 times too strict.

The order of the two steps, and whether it could be otherwise

A reader might reasonably ask why the two jobs have to share a matrix at all. Nothing in the physics requires it: a model could adapt in one basis, come back to tristimulus values, and apply its compression in another.

Two reasons it does not, and only one of them is good.

The good reason is that the model is meant to describe a mechanism. The story CIECAM16 tells is that three receptor classes apply a gain and then respond nonlinearly, and in that story there is one set of channels because there is one set of cells. Splitting the basis would be admitting that the model is a curve fit with two free 3×3s in it, which it would then have to fit — doubling the parameters against the same corresponding-colour data.

The less good reason is inertia. The chain from Hunt’s models through CIECAM97s, CIECAM02 and CIECAM16 has always had this shape, each revision changing constants rather than structure, and what a second model changes is rarely its skeleton. Nobody chose to couple the two jobs; the coupling arrived with the mechanism story and stayed after the model became a fit.

The measurement here does not argue for splitting them. It argues for knowing the price: about a quarter on one job and two thirds on the other, at present, with a matrix that happens to be well placed.

Who chose it, and against what

CAT16’s entries were published in 2016 in the revision that also produced CIECAM16, and the case made for them was explicit and narrow: recover CAT02’s accuracy on corresponding-colour data while removing the negative-value failure. Both halves of that are about the first job or about numerical behaviour, and neither is about the geometry of small differences.

The corresponding-colour data sets themselves deserve a word, because they are what fitted means here. An observer adapts to one condition, matches a stimulus, adapts to a second, and matches again; the pairs are the data, and a transform is scored by how well it predicts the second member from the first. The protocols are decades old, the sets are small, and the adaptation period they specify is a moment on a clock the model does not have.

What no scoring procedure in that tradition can see is the uniformity of the resulting space, because uniformity is measured against threshold data — ellipses, difference judgements — that corresponding-colour experiments do not produce. The two evidence bases are disjoint. That is the whole mechanism of the oversight, and it is not anybody’s error.

Where the model stops

This is not a re-derivation of CIECAM16. Nothing here runs the appearance model with a different adaptation matrix and reports what the correlates do; that would need the whole forward and inverse path rebuilt on a parameterised basis, and every figure in this collection that uses the model would move with it. What is measured is the two objectives the matrix sits between, not the model’s own output.

The uniformity measure is a cube root, not the model’s compression. The two are similar in shape over the range that matters and they are not the same function. The ranking of bases is what is being claimed, and it is insensitive to the exponent over a range from a square root to a tenth root.

And a matrix is not a mechanism. CIECAM16’s axes are a parameter in a model that predicts judgements; calling them cone-like is a description of their shape, and what happens when the actual receptors are imposed is a different measurement with a different answer.

The generalisation

When one parameter serves two stages of a pipeline, it acquires an objective from each, and the fit that produced it saw one of them. That is a general hazard of models built in layers, and it is invisible from inside the fitting procedure because the second objective is defined downstream of the thing being fitted.

The check is cheap and is the one performed here: hold the pipeline fixed, vary the shared parameter, and score the other stage. If the second score moves a lot, the parameter is carrying an undeclared decision, and whether it happens to be a good one is a matter of what the fit’s own criterion happened to correlate with.

Here it moved by a factor of 2.5 across the published candidates, and the one in the standard is the best of them by an accident of numerical hygiene.

The receptors are ahead, not behind

CAT16 is best of the published transforms, ahead of the receptors’ 2.60 by a whisker has the comparison the wrong way round. On this measure a smaller axis ratio is better, and 2.60 is smaller than 2.71: the receptor construction beats CAT16 by four per cent, and CAT16 is best of the published transforms while being second to a construction nobody publishes as a transform.

Which strengthens the essay’s own mechanism rather than weakening it. The account offered is that sharpening is what the adaptation objective rewards and the uniformity objective punishes, so a matrix backed away from CAT02’s sharpening moves towards the receptors — and if that is right, the receptors should be the best of all, being the thing being moved towards. They are.

The ordering across four transforms bears it out. Ranked by how nearly cone-like each is, using its implied deuteranope confusion point’s distance from the measured one, and ranked by uniformity:

transform distance from the measured point uniformity
Hunt–Pointer–Estévez 1.29 2.81
CAT16 1.47 2.71
Bradford 2.40 3.68
CAT02 4.07 4.29

The two orderings agree at a Spearman of +0.80, with a single inverted pair — HPE and CAT16, which swap. So being cone-like predicts making circles across the published set, which is the essay’s explanation turned into a measurement rather than a story, and the one exception is the two entries that are closest together on both axes.

The uniformity job is twice as far from its floor

The two improvements are quoted as 26 and 40 per cent, and there is a second pair of numbers that says the same thing more sharply.

Measured as a share of what is removable, CAT16 gives up 26 per cent on adaptation and 40 on uniformity. Measured as distance above the floor, it sits 34.5 per cent above 0.974 and 67.3 per cent above 1.620.

The uniformity job is 1.95 times as far above its own floor. That is the number the essay’s uncomfortable finding deserves, because a proportional distance is the comparable quantity when the two objectives are in different units — one is a colour difference in ΔE₀₀ and the other a dimensionless axis ratio, so nothing but a ratio to their own floors can be set side by side.

And it says which way to read the second result is luck. Luck placed CAT16 first among the published transforms on a job it was not scored on, and the same luck left it twice as far from that job’s floor as from the other’s. Being the best of a bad set is what an accident produces, and the proportional distances are what distinguishes that from being good.

Two point five spans more than the candidates

It moved by a factor of 2.5 across the published candidates is the generalisation’s supporting figure, and among the transforms that could actually occupy CIECAM16’s slot the spread is smaller.

The five adaptation bases run 2.71 to 4.29 — a factor of 1.58. Reaching 2.5 needs the comparison to include something that is not an adaptation transform at all: the receptor construction at 2.60 and a set of display primaries at 6.40 span 2.46, which is the figure quoted.

That is not a flaw in the check the section recommends — hold the pipeline fixed, vary the shared parameter, and score the other stage is exactly right, and 1.58 is still a large undeclared variation for a parameter chosen against a different criterion. It is a flaw in the number offered for it. The decision CAT16 carries is worth 1.6 across the things it could have been, not 2.5, and 1.6 across five serious candidates is a more alarming statement than 2.5 across a set containing a display’s primaries, because nobody was ever going to put a display’s primaries there.

Where the ladder goes next

The uniformity job’s floor of 1.61 has been quoted three times now without being examined. It sits below a number this collection has called a property of the eye, and reconciling those two sentences is the next rung.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

BasisCAT16Chromatic adaptationCIECAM16Colour appearanceMacAdam's ellipsesPerceptual uniformitySpecificationTrade-offThe von Kries transform