What the brain does

The trade only runs one way

Standing at the basis that adapts best, one per cent of adaptation buys forty-four per cent of the way to the discrimination floor. Standing at the basis that discriminates best, the same one per cent buys under two. The scatter that shows two objectives pulling apart looks symmetric and is not, and the asymmetry is what a committee choosing between them would most want to know.

Assumes No basis is good at both, How long is the bowl and One matrix doing two jobs.

Two objectives written over the same nine numbers pull in different directions, which is a shape. How much the first step costs is a number, and it is the only part of the curve anybody is ever standing on.

The cheapest direction to give ground in is the flattest one. Six bars, one per direction the adaptation objective can see, showing how much of the other objective a fixed budget of adaptation buys if it is spent along that direction. The rate is the slope of the second objective divided by the square root of the first's curvature, so it rewards a direction the second objective wants and punishes one the first is stiff in. The flattest direction wins at 10.68 against 3.29 for the next best and 0.54 for the stiffest — a factor of 20. Spending 1 per cent of the adaptation optimum there moves the anisotropy from 7.70 to 5.02.
Fig. 1 How much of the second objective a fixed budget of the first buys, direction by direction. The flattest direction wins by a factor of twenty.

The claim

From the basis that minimises the adaptation residual, spending one per cent of that residual moves the ellipse anisotropy forty-four per cent of the way to its floor. From the basis that minimises the anisotropy, spending one per cent of it moves the adaptation residual under two per cent of the way to its floor. The trade is real in one direction and nearly absent in the other.

  • The rate is |gₖ|/√λₖ. At an optimum the first objective rises quadratically and the second changes linearly, so a budget Δ spent along direction k buys gₖ√(2Δ/λₖ) of the second.
  • The flattest direction wins, at a rate of 10.68 against 3.29 for the next best and 0.54 for the stiffest — a factor of twenty between the ends.
  • Measured rather than extrapolated, one per cent takes the anisotropy from 7.70 to 5.02 against a floor of 1.61. Five per cent takes it to 3.28.
  • The other way round, one per cent takes the adaptation residual from 1.793 to 1.778 against a floor of 0.974 — 1.8 per cent of the gap.
  • And nothing is converted into anything. A residual is a ΔE00 and an anisotropy is a ratio of two lengths. What is computed is how much of one a stated amount of the other buys, which is a fact about two surfaces rather than an exchange rate anybody invented.

Where the asymmetry comes from

At the optimum of one objective its own gradient is zero, so its cost rises as ½λt²: the first step is nearly free. The other objective has no reason for its gradient to vanish there, so its value changes as gt: the first step is full price. That asymmetry between quadratic and linear is what makes any first step a bargain, and it is symmetric between the two — it applies at either end.

What is not symmetric is the size of g and the size of λ at each end.

At the adaptation optimum, the anisotropy is 7.70, which is a long way above its floor of 1.61, so there is a great deal of room to move; and the flattest direction the adaptation objective has is very flat, λ = 0.760, so a small budget buys a long step. The product is large.

At the discrimination optimum, the adaptation residual is 1.793 against a floor of 0.974 — a gap of 0.82 rather than 6.09 — and the flattest direction there is stiffer relative to the objective’s own scale. A small budget buys a step, and the step arrives somewhere the first objective barely cares about.

So the pair is not a symmetric trade-off between two equally particular objectives. One of them is standing on a steep hillside with a long way to fall, and the other is standing in a shallow basin.

Why this is not the slope of a front

The natural object here is the Pareto front — the curve of bases that cannot be improved on one objective without losing on the other — and the natural quantity is its slope, which is the exchange rate in the usual sense. That quantity does not exist at the point this essay starts from, and the reason is worth a paragraph because it is why the arithmetic is phrased as a budget instead.

At the adaptation optimum the first objective rises as and the second falls as t, so d(second)/d(first) behaves as 1/t and diverges as the step goes to zero. The front is vertical there. Any attempt to quote a rate as a ratio of derivatives returns infinity, correctly, and says nothing useful: an infinite rate is the statement that the very first sliver of adaptation is free, which is true and is not a decision anybody can act on.

Fixing a budget and asking what it buys is finite, is what a decision actually looks like, and carries the square root that makes the flat directions valuable. √(2Δ/λ) is where the factor of twenty between the ends comes from — the reach goes as the inverse square root of the curvature, so a direction thirty times flatter is only five and a half times longer, and the rest of the twenty comes from the gradient.

A vertical front is the normal situation at any interior optimum, not a peculiarity of this pair, and it is the reason multi-objective results are usually reported as fronts rather than as rates. Reporting a budget is the middle course: less than a whole curve, more than a number that is infinite.

Which direction to spend in

The budget could be spent in any of the six directions the first objective can see, and the rate differs by twenty times between them.

The rate is |gₖ|/√λₖ, which rewards two things at once — and which needs the objective to be differentiable at all, something only one of the two measurements was until recently: a direction the second objective wants — large |gₖ| — and a direction the first objective does not mind — small λₖ. Neither alone is enough, and the winner is not the one that would be picked by looking at either.

Measured at the adaptation optimum, the six rates are 0.54, 0.78, 3.29, 2.93, 2.75, 10.68, stiffest first. The flattest direction is the winner by more than a factor of three over the next and twenty over the stiffest.

That answers a question this collection left standing when it first measured the shape of these optima: what the flat directions are for. They are not merely places where the answer is undetermined. They are the cheapest currency the objective has, and the only thing that decides their value is what else is on offer along them.

The bowl the eigenvalues describe and the bowl a sample found. Six points on a logarithmic vertical axis — the distance from the optimum of the adaptation residual to a 5 per cent rise along each of the six directions the objective can see — with a shaded band behind them showing the whole range 24 random directions reported. The eigen-radii run from 1.2e-2 to 3.6e-1, a factor of 29.8. The band runs from 2.2e-2 to 1.8e-1, a factor of 8.0, and sits entirely inside the ends of the true range: a random direction in nine dimensions carries a share of every eigenvector and so reports the middle of the bowl, never an end of it.
Fig. 2 The six reaches. The rightmost is where a budget buys most, and it is the direction the objective has the least opinion about.

The measurement, not the extrapolation

A linear term extrapolated over a finite step is exactly the arithmetic that reads well and is wrong, so the essay’s headline number is walked rather than predicted.

Predicted: one per cent of the adaptation optimum, spent along the best direction, should take the anisotropy from 7.697 to 6.207. Measured, by moving the basis and recomputing: 5.023. The step does better than the linear term promises, because the second objective is also curving downward along that direction — the walk is going towards its optimum, not away from anything.

At five per cent the prediction is 4.364 and the measurement is 3.283.

So the honest statement is that a linear rate identifies which direction to spend in and understates how much the spending is worth, and both numbers are reported rather than one.

What that means for a matrix that does both jobs

CIECAM16 adapts and compresses in the same axes: one matrix, chosen once, doing two jobs. What that costs has been measured — 35 per cent above one floor and 68 above the other — and this essay adds the shape of the choice it faced.

Starting from the adaptation optimum and spending one per cent lands at an anisotropy of 5.02, which is still well above CAT16’s 2.71. Spending five per cent lands at 3.28, which is above it. So the single matrix is not simply a point on the cheap part of this exchange — it is far better on the second objective than a small budget from the first optimum can reach, and it pays 35 per cent on the first to be there.

That is a defensible place to stand, and it is a choice about weights that nothing in the standard records. The exchange rate makes the choice legible: at the adaptation end the first per cent is worth an enormous amount of the second objective, so a design that refuses to spend anything is leaving a great deal on the table, whatever the eye is doing about the same problem; and by the time enough has been spent to reach CAT16’s anisotropy, the price is no longer small.

Where the walk stops

Two points on a curve do not say whether the curve has an end. It has one.

Walked out along the winning direction, the anisotropy falls 7.697, then 5.023 at one per cent, 4.259 at two, 3.283 at five, 2.802 at ten — and reaches 2.594 at a twenty-five per cent budget, after which it climbs again: 2.603 at thirty, 2.629 at thirty-five, 2.801 at fifty. The bottom is 83.7 per cent of the way to the floor of 1.611, and it is bought with about a quarter of the adaptation residual.

The linear prediction has no such bottom, which is the point of walking. It falls straight through the floor and out the other side — −1.12 at a thirty-five per cent budget, −2.84 at fifty — which is a negative ratio of two lengths. That is what a first-order term does when it is asked about a finite step.

The ratio between measurement and prediction is informative in its own right. At half a per cent the walk beats the extrapolation by 2.11 times; at one per cent by 1.79; at two by 1.63; at five by 1.32; at ten they draw level at 1.04; and past that the prediction is running away downwards. The bonus for walking rather than extrapolating is largest exactly where the budget is smallest, which is the region the whole argument is about.

The scatter can be read at a larger rise as well, which tests whether the ordering of the constraints is a local accident of where the level set was drawn.

What a constraint costs is how far it pushes, in the directions that are seen. A scatter of every constraint imposed here on the nine free numbers. The horizontal axis is the length of the displacement from the optimum measured only in the six directions the objective can see; the vertical, on a logarithmic scale, is the excess cost that displacement actually carries. Requiring the basis to be the inverse of three realisable display primaries sits at the bottom left, at 0.068 and 0.022 ΔE00 — it removes three degrees of freedom and moves the answer almost nowhere. Requiring it to hit the three dichromat confusion points removes six and pushes 13 times as far, for 0.68. The vertical spread at similar horizontal positions is the part a count of parameters cannot predict.
Fig. 3 Every constraint this collection has imposed, at a five per cent rise rather than one. The display-primaries constraint still sits at the bottom left and the dichromat-confusion-point one still pushes about thirteen times further, so the ordering survives a fivefold change in where the boundary is put.

How long the reverse trade lasts

The other direction is not merely weaker. It is shorter, and its length is the sharper statistic.

From the discrimination optimum the adaptation residual falls from 1.7930 to a minimum of 1.7634 at a budget of 0.35 per cent, and rises from there: 1.7785 at one per cent, 1.8243 at two, 1.8832 at three. It is worse than where it started before two per cent has been spent. Its entire lifetime yield is 3.6 per cent of the gap to the adaptation floor.

So the asymmetry in the claim understates itself. The forward trade has a runway a quarter of its objective long and takes five sixths of the gap available to it. The reverse trade has a runway a third of one per cent long, takes a thirtieth of its gap, and then begins charging for the privilege.

Where the walk lands among the published transforms

The consequence worth following is that a walk along one eigenvector is not a poor imitation of a designed compromise. It beats most of them.

At a fifteen per cent budget the walk sits at an adaptation residual of 1.1202 and an anisotropy of 2.6638. That is better than XYZ, the Hunt–Pointer–Estévez axes, Bradford, CAT02 and CAT16 on both objectives at once — five of the six published bases this collection carries, dominated by moving one optimum along one direction.

CAT16 deserves its own sentence, being the current recommendation. It stands at 1.3118 and 2.7060. The walk reaches CAT16’s anisotropy at a budget of 12.9 per cent, where its adaptation residual is 1.1001 — the same discrimination performance, for thirteen per cent where CAT16 pays thirty-five. So the reading above, that a small budget cannot reach the single matrix’s position, is right at one and five per cent and reverses by thirteen.

Two cautions belong immediately after that, and neither withdraws it. Both objectives are constructions over this collection’s own censuses, so a different census moves every number here. And CAT16 was fitted against corresponding-colour data rather than against these two surfaces, so it is being measured by this comparison rather than competing in it.

What survives both is narrower and still worth having. On these two objectives, in this parameterisation, the published transforms are not on the front — and the direction that would carry them to it is the one the first objective minds least about, which is the direction nothing in a fitting procedure has any reason to explore.

How much of the disagreement is structural rather than chosen can be read off the principal angles, which is the same measurement taken at a third of the earlier rise.

Where the two objectives agree and where they do not. Six horizontal bars: the principal angles between the six-dimensional subspaces the two objectives can see. Three of them are zero to within a rounding error — which is forced, because two six-dimensional subspaces of a nine-dimensional space must share at least three directions, and their being exactly zero is the check that the angles are being computed correctly. The other three are 52.6°, 41.5°, 8.1°. The two objectives are neither the same question asked twice nor two independent questions; they overlap in half of what they can see.
Fig. 4 The six principal angles between the two objectives’ visible subspaces. Three are zero because two six-dimensional subspaces of a nine-dimensional space must share three directions, and the other three are 52.6°, 41.5° and 8.1° — so the two objectives overlap in half of what they can see and differ in the rest.

What a committee is actually choosing

The exchange makes a decision legible that is usually made by not making it.

A standard that recommends an adaptation transform is choosing a point on a curve between two objectives, and the choice is recorded as a matrix rather than as a weight. Reading the matrix back through the exchange says what weight it implies. The receptor construction sits at 1.65 and 2.60; CAT16 at 1.31 and 2.71; Bradford at 1.14 and 3.68. Bradford’s extra quarter of a ΔE00 of adaptation performance over CAT16 is bought at a full unit of anisotropy, which at the rates measured here is an expensive purchase — and nothing in the record of either transform says that trade was considered, because each was fitted against one criterion alone.

A discipline that reports one of two objectives has made the decision without recording it. That sentence has been in this collection since the two objectives were first put on one picture. The exchange rate is what turns it from an accusation into an arithmetic: the implied weight can be computed, and it is different for every transform in the table.

The uncomfortable half is that the implied weights are not stable across the census. Every rate here depends on both objectives’ gradients, and both gradients depend on a set of illuminants and a set of contours. So the statement “this standard implicitly weights adaptation over discrimination” is robust; the number attached to it is not, and quoting one without the census it was computed on would be quoting a choice as though it were a measurement.

What was computed, and how

The Hessian of the first objective at its own optimum is 154 central differences at a measured step; its eigendecomposition is a one-sided Jacobi singular value decomposition. The gradient of the second objective — the anisotropy of the discrimination contours — at that same point is eighteen central differences at the same step, and it is not zero — that non-vanishing is the whole mechanism.

The rate for each direction is |gₖ|/√λₖ, and the step for a budget Δ is √(2Δ/λₖ). Both directions along an eigenvector are available, so the gain is the size of the linear term rather than its signed value: a direction the second objective dislikes is a direction to walk backwards along.

The assertion in the build has four parts. The measured gain must be at least twenty per cent of the gap to the second floor; the winning direction must be the flattest of the six rather than any other; it must win by at least a factor of three over the next; and the same budget spent the other way round must buy at most six per cent. The third condition is there because a rate table with two near-equal leaders would not support the sentence about flat directions.

Where the model stops

The rate is local and the walk is not. |gₖ|/√λₖ is exact at the optimum and the step it prescribes is finite; the measurement is what is quoted precisely because the extrapolation is not to be trusted at a per cent, let alone at five.

Both objectives are constructions, and their relative scales are what an exchange is sensitive to. The adaptation residual is a mean ΔE00 over fourteen changes of illumination; the anisotropy is a mean ratio over twenty-five ellipses, measured from their own quadratic forms rather than by walking round them, measured on one observer in 1942. Changing either census changes the gradient, the curvature and therefore the rate. Nothing here converts one into the other, and that restraint is deliberate: an exchange rate between two incommensurable quantities is exactly the decision this collection keeps saying nobody has made explicitly.

And the asymmetry is a statement about two particular optima. It says the adaptation optimum is a bad place to stand if the other objective matters at all, and that the discrimination optimum is not correspondingly bad in the other direction. It does not say one objective is more important; nothing here could.

The generalisation

At an optimum, the first step is free in one currency and full price in the other, so the first step is always the best deal — and how good a deal depends on where the other objective’s floor is relative to where it currently stands.

The practical form is a habit. Before defending a design that optimises one objective, price one per cent of it in the others. The arithmetic needs a gradient and a Hessian, both of which are cheap, and it produces the one number a decision actually turns on: not how far apart are the two optima, which is a shape, but what does the first step cost, which is a choice.

The second half is about where to spend. The cheapest direction is the one with the best ratio of what the other objective wants to what this one minds, and it is generally neither the flattest direction nor the direction of steepest improvement in the other objective. Picking either of those alone is a common shortcut and it loses a factor of three here.

Where the two objectives agree and where they do not. Six horizontal bars: the principal angles between the six-dimensional subspaces the two objectives can see. Three of them are zero to within a rounding error — which is forced, because two six-dimensional subspaces of a nine-dimensional space must share at least three directions, and their being exactly zero is the check that the angles are being computed correctly. The other three are 52.6°, 41.5°, 8.1°. The two objectives are neither the same question asked twice nor two independent questions; they overlap in half of what they can see.
Fig. 5 The angles between the two objectives’ seen subspaces, which is why any exchange exists: neither aligned, which would make one basis serve both, nor perpendicular, which would make them independent.
The error a curvature leaves halves when the step does. Three curves on logarithmic axes. Each is the error in predicting a basis's excess cost from the quadratic form alone, plotted against the fraction of the way from the optimum to that basis. All three fall along a line of slope one: halving the displacement halves the error, which is what the third term of a Taylor series does and is the evidence that the matrix being used is a second derivative rather than a fit. At an eighth of the way the prediction is out by up to 39 per cent; at a hundred and twenty-eighth it is out by 3.7 per cent at worst.
Fig. 6 The error the quadratic account leaves, against how far along each direction the step is taken. It halves when the step does, which is what says the rates above are a local statement rather than a global one — and the asymmetry survives inside the range where the statement holds.

Who found it, and when

Multi-objective optimisation has the vocabulary — a Pareto front, and the exchange rates along it are its normal vectors — and it has had it since Pareto in 1906 and its modern form since the 1950s. Nothing here is a new idea in that literature.

What colour appearance modelling has not done is draw the front. Bradford, CAT02, CAT16 and the receptor construction are reported as matrices with a residual against one criterion; the second criterion is measured, when it is measured at all, in a different paper by different people. The two objectives here have been in the same collection for only a short while, and the scatter that puts them on one picture makes them look like a symmetric pair because a scatter has no scale on it.

The asymmetry is the sort of thing a picture hides and a derivative reports. Two points at the ends of a diagonal look like two ends. The gradients at those two points say one of them is on a cliff.

Where the ladder goes next

Everything in this essay is measured against a census of illumination changes taken from a list of fourteen. The next question is what happens when the list is replaced by the family it was drawn from — and the answer moves more than any of the numbers here: the transform that is best on the average is the one that stops being defined at the edge.

What a constraint costs is how far it pushes, in the directions that are seen. A scatter of every constraint imposed here on the nine free numbers. The horizontal axis is the length of the displacement from the optimum measured only in the six directions the objective can see; the vertical, on a logarithmic scale, is the excess cost that displacement actually carries. Requiring the basis to be the inverse of three realisable display primaries sits at the bottom left, at 0.068 and 0.022 ΔE00 — it removes three degrees of freedom and moves the answer almost nowhere. Requiring it to hit the three dichromat confusion points removes six and pushes 13 times as far, for 0.68. The vertical spread at similar horizontal positions is the part a count of parameters cannot predict.
Fig. 7 Every constraint on the same nine numbers, priced by how far it pushes in the directions that are seen. The exchange of this essay is the same arithmetic run forwards instead of backwards.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyBasisCAT16Chromatic adaptationEigenvalueGradientHessianLevel setMacAdam's ellipsesTrade-off