What it takes to deliver it

The slope arrives before the bowl

The adaptation transforms colour management actually uses are not optima of anything. At every one of them the objective has a slope, and the slope is the larger term over almost the whole distance to the best matrix — so a table of curvatures taken there describes a bowl nobody meets on the way anywhere.

Assumes How long is the bowl, Only the flat directions keep their names and No matrix is right everywhere.

The matrices colour management uses are not the bottom of anything, and a curvature taken at a point with a slope describes the wrong thing.

At a published matrix the slope arrives long before the bowl. One row per basis in this collection's table. Each row is a logarithmic axis of distance in the nine coefficients, with two markers: the radius at which the objective's curvature becomes as large as its slope, and the distance from that basis to the optimum. The first is between 3.2 and 108 per cent of the second. So over almost the whole journey from a published matrix to the best one, the surface is a slope and not a bowl — and a table of eigenvalues taken there describes a neighbourhood the optimum is nowhere near. XYZ scaling is the exception, at 1.08 of the distance, because its slope is the steepest in the table.
Fig. 1 At each published basis, the radius where the objective’s curvature becomes as large as its slope, against the distance from that basis to the optimum. The first is a few per cent of the second.

The claim

Every adaptation transform in use is a point where this collection’s objective has a gradient, and within the neighbourhood a Hessian describes, that gradient is the larger term. So the eigenvalues at a published matrix describe a bowl the reader will not meet on the way to anywhere.

  • At a minimum the gradient vanishes and the Hessian is everything. At Bradford it does not: the cost rises linearly at first and quadratically only later.
  • The two terms are equal at a radius of 0.035 to 0.15 in the nine coefficients, for six of the seven bases measured.
  • The distance to the optimum is 1.0 to 2.1. So the curvature catches the slope at three to eleven per cent of the way there.
  • XYZ scaling is the exception, at 108 per cent — its slope is so steep and its curvature along the slope so slight that the quadratic term never catches up before the optimum is passed.
  • And every published matrix is at a saddle: three or four of the nine directions curve downwards, which a table of eigenvalue magnitudes cannot show because a magnitude has lost the sign.

Where the question came from

The previous round measured the shape of an optimum and got a great deal from it — a rank, a null space, an eigenbasis that explains why one constraint costs seventy per cent and another costs nothing. It left one line in the record: every published transform sits outside the neighbourhood a Hessian describes, and what the objective looks like at Bradford or CAT16 is a different set of six numbers and has not been computed.

It has now, and the first thing the numbers say is that the question was slightly wrong. At a published matrix, six numbers are not what describes the surface. Nine numbers and a direction are — the gradient first, and the curvature second, in that order of importance over the distances that matter.

That is worth saying plainly because it is easy to slip past. Everything this collection knows about the shape of this objective is knowledge about two points, and every matrix a colour management system loads is somewhere else.

The two terms

Along a unit direction from a point, a smooth function changes by the gradient’s component times the distance, plus half the curvature along that direction times the distance squared, plus higher terms.

The two terms are equal at t* = 2|gᵀd| / |dᵀHd|. Below that radius the linear term dominates and the surface is a slope; above it the quadratic term dominates and the surface is a bowl. The number to compare t* with is how far the point is from the answer.

Along the steepest direction the radii are: 0.035 at Bradford, 0.069 at the discrimination optimum, 0.095 at Hunt–Pointer–Estévez, 0.099 at the receptor construction, 0.134 at CAT16, 0.152 at CAT02, and 1.12 at XYZ scaling. The distances to the optimum, measured in the six directions the objective can see, are 1.09, 1.78, 2.06, 2.08, 2.04, 1.48 and 1.04.

So on six of the seven the surface is a slope over ninety per cent or more of the journey.

Bradford is the one worth pausing on, because it is the transform colour management actually uses and it has the smallest ratio in the table at three per cent. It also has the smallest gradient of the published four, at 4.0 — so it is the nearest to being at rest, and the reason its ratio is small is that its curvature along the steep direction is large rather than that its slope is.

A small t* is therefore not a sign of a bad basis. It is a sign of a basis in a steep-walled part of the surface, which is where a well-behaved matrix ought to be: the walls are what hold it in place.

The radius is measured along a direction the journey does not take

The essay’s central quantity is a radius along the steepest direction, and its central comparison is against the distance to the optimum. Those two are not on the same line, and the closing section says by how much: steepest descent at every published transform sets off 80 to 87 degrees from the straight line to the optimum.

So t* is a statement about a direction almost perpendicular to the one the reader is travelling in, and the ratio it produces — three per cent at Bradford — is not the fraction of the journey over which the surface is a slope. It is the fraction of a different journey.

The relevant radius can be estimated from the essay’s own numbers. Bradford’s gradient is 4.0 and its excess over the optimum is 1.14 − 0.97 = 0.17, at a distance of 1.09. Taking the departure angle at 83 degrees, the directional slope along the line to the optimum is 4.0 × cos 83° = 0.49, so a purely linear surface would give an excess of 0.53 — three times what there is.

Fitting a quadratic along that line to reproduce the actual 0.17 puts the two terms equal at t* = 1.60, which is 147 per cent of the distance. Along the path that matters, Bradford behaves like XYZ scaling: the linear term leads the whole way and the quadratic never catches it.

departure angle slope along the line linear term over the path t* as a share of the distance
80° 0.695 0.757 129%
83° 0.487 0.531 147%
87° 0.209 0.228 392%

The conclusion survives and the number that supports it does not. At a published matrix the slope is the larger term over the distances that matter is true along the path — more strongly than the essay claims, since the ratio is above one rather than at three per cent. What is not true is that the three per cent measures it.

The quadratic term is not confined to the last few per cent

The same arithmetic contradicts one sentence directly. Almost all of that excess is accounted for by the linear term over the path back and the bowl is a description of the last few per cent of the journey are both stated, and the excess says otherwise.

A linear extrapolation along the path gives 0.53 and the actual excess is 0.17, so the quadratic term removes 68 per cent of what the linear term alone would predict — at the 83-degree reading, and between 25 and 78 per cent across the plausible range of angles.

That is not a correction confined to the end of the journey. It is a term of the same order as the one it corrects, integrated over the whole path, and it is what makes the excess three times smaller than a plane would give. The surface is a slope in the sense that the linear term never loses its lead, and a bowl in the sense that the quadratic term is doing most of a factor of three. Both are true and the essay states only the first.

What sets t*, since it is not the gradient

Two sentences in the essay say the ordering is dominated by the gradient — they order the bases identically because both are dominated by the gradient. The seven bases do not support it.

The rank correlation between gradient length and t* is +0.04, and between gradient length and t* as a share of the distance it is also +0.04. Both are nothing. Hunt–Pointer–Estévez has the largest gradient in the table and the third smallest radius; CAT16 has the smallest gradient and the third largest.

What sets t* is the ratio, and the curvature is the more variable half of it. Recovering the curvature along each steepest direction from t* = 2|g| / c:

basis |g| curvature along it t*
CAT16 3.3 49 0.134
XYZ scaling 15.7 28 1.120
CAT02 6.9 91 0.152
the discrimination optimum 7.3 212 0.069
Bradford 4.0 229 0.035
the receptor construction 17.3 349 0.099
Hunt–Pointer–Estévez 23.5 495 0.095

The curvatures span a factor of 17.7 and the gradients a factor of 7.1, so the curvature is the larger source of variation. XYZ scaling’s exceptionalism is entirely in this column — its curvature along its own steepest direction is 28, the smallest in the table by a factor of 1.8, which is what the essay says in words and can now say in a number.

And the Bradford-against-CAT16 contrast the essay draws checks out exactly: the two have the smallest gradients, 4.0 and 3.3, and Bradford’s curvature is 4.6 times CAT16’s, which is the whole of why its radius is 3.8 times smaller.

What that means for a table of eigenvalues

A curvature is not wrong at a point with a slope — it is a real second derivative and it is exactly what it says. What it is not is the leading description.

Reading a table of eigenvalues at Bradford as the shape of the objective there invites two mistakes at once. The first is about size: a step of 0.1 from Bradford changes the cost by about 0.1 × |g| = 0.4, and by about 0.005 × λ from the quadratic term, so the quadratic term is a twentieth of the change. The second is about direction: the steepest direction and the stiffest direction are different directions, and a picture of a bowl points at the second.

Nine eigenvalues, six of which exist. Nine points on a logarithmic vertical axis: the eigenvalues of the Hessian of the adaptation residual at its own optimum, largest to smallest. The first six run from 6.8×10² down to 7.6×10⁻¹, a condition number of 890. Then the axis drops: the seventh is 1.9×10⁻⁴, and the last three are separated from the sixth by a factor of 4.0×10³. Those three are not small curvatures. They are the finite-difference truncation error on directions along which the objective is exactly constant, and a shaded band marks them as the numbers the objective does not have.
Fig. 2 Nine eigenvalues, six of which exist — at the optimum. Every quantity in this picture is exact and is a statement about one point in the table.

The useful pair of numbers at a published matrix is therefore the gradient’s length and the radius at which the curvature catches it, and both are cheap: eighteen evaluations for the gradient and a hundred and fifty-four for the Hessian.

The gradients are 3.3 at CAT16, 4.0 at Bradford, 6.9 at CAT02, 7.3 at the discrimination optimum, 15.7 at XYZ scaling, 17.3 at the receptor construction and 23.5 at Hunt–Pointer–Estévez. Against 5 × 10⁻⁴ at the optimum, which is a finite-difference zero.

There is a version of this that a reader may find more intuitive, in units of the objective rather than of the coefficients. The excess of a published matrix over the optimum is between 0.17 and 1.40 ΔE00 — Bradford is 1.14 against the optimum’s 0.97 — and almost all of that excess is accounted for by the linear term over the path back. The bowl is a description of the last few per cent of the journey, and the first ninety-five is a straight run downhill.

The exception, which is instructive

XYZ scaling is the one basis whose curvature does not catch its slope before the optimum is reached, at a ratio of 108 per cent, and the reason is not that its slope is the steepest.

Its gradient is 15.7, which is the second largest. What makes it exceptional is the curvature along that gradient, which is unusually small: the direction XYZ scaling falls fastest in is one the objective is nearly flat along, so the linear term keeps its lead the whole way.

That is the same basis that is exceptional in every measurement in this thread: the largest excess over the optimum, the longest trust radius, the most rotated stiff plane at 55 degrees. It is far enough from everything else to behave like a different problem, which is a reasonable thing for the oldest mistake still shipping to do.

Every published matrix is at a saddle

The signed curvature along each principal direction says something a table of magnitudes cannot, and it is the more alarming half of this essay.

At the optimum, all nine directions curve upwards or are zero — six positive and three at the numerical floor — which is what being a minimum means. At every published basis, three or four of the nine curve downwards: there are directions in which the objective falls away quadratically as well as linearly.

At a published matrix the surface curves downwards in some directions. Nine markers per series, on a logarithmic axis of magnitude: the objective's curvature along each of its own principal directions, at Bradford and at the optimum. Filled markers are directions that curve upwards and open ones curve downwards. At the optimum every one is upwards or zero, which is what being a minimum means. At Bradford 3 of the nine curve downwards, so the point is a saddle — there are directions in which the objective falls away quadratically as well as linearly, and describing that neighbourhood as a bowl is wrong in a way an eigenvalue table does not show, because a table of magnitudes has lost the sign.
Fig. 3 The signed curvature along each principal direction at a published matrix and at the optimum. Filled markers curve upwards, open ones curve downwards, and only the optimum has none open.

A singular value decomposition returns magnitudes, so a table of eigenvalues at Bradford looks exactly like a table at the optimum: nine positive numbers spanning several orders of magnitude. The sign is the whole difference between a bowl and a saddle and it is thrown away by the decomposition that produces the picture.

Counting them needs a threshold that is not the roundoff floor, because the three flat directions have an exactly zero form which a second difference computes as 10⁻⁸ with whichever sign the noise had. At a threshold of 10⁻⁵ — inside the gap between the smallest real eigenvalue at 10⁻⁴ and the numerical zeros at 10⁻⁸ — the count is zero at the optimum and three or four everywhere else.

What was computed, and how

Gradients and Hessians by central differences at the step this collection measured a window for, at eight named bases. The distance to the optimum has the three invisible directions projected out of it, because the optimum is not a point: three of the nine directions are exactly flat, so the objective’s minimiser is a three-parameter family and a search returns one arbitrary member of it.

Without that projection every distance would carry a component nobody chose. With it, the distances fall by five to thirty per cent and the ratios in this essay are statements about the part of the displacement the objective can see.

A quadratic is believed least far at the one place anybody takes one. One bar per basis: the radius, in the nine coefficients, within which the second-order model predicts the objective to within ten per cent in every one of eighteen directions. The shortest bar is the objective's own optimum, at 2.3×10⁻², and the longest is XYZ scaling at 1.1×10⁻¹ — several times further. The reason is not that the model is worse at a minimum but that it has less to do there: away from one the linear term is exact and carries most of the change, so a ten per cent error in the prediction takes longer to accumulate. It does not make a Hessian at a minimum wrong; it says the picture drawn from it describes the smallest neighbourhood in the table.
Fig. 4 The radius over which a second-order model predicts the objective to within ten per cent, at each basis. A different quantity from the one in this essay and it orders the bases the same way, because both track the gradient.

The two quantities in this thread are related and not the same. t* is where the bowl catches the slope; the trust radius is where the two of them together stop describing the surface. They order the bases identically because both are dominated by the gradient — which is itself a small finding, and the reason to report both rather than either.

Why this matters for a delivery chain

The published transforms are not a curiosity: they are what an ICC profile’s chromatic adaptation applies when a colour is carried between two white points, and Bradford in particular is in essentially every profile in circulation.

So the shape of this objective at Bradford is the shape at the point where all the arithmetic actually happens, and the practical consequence is a small one about interpretation. A sentence of the form this transform is nearly optimal, and the objective is flat around it is two sentences, and only the first is about Bradford: the flatness is a statement about the optimum, and Bradford is on a slope four ΔE00 units per unit of coefficient steep.

What that permits is a useful piece of advice for anybody tempted to tune one. A small perturbation of a published matrix moves the objective linearly, so tuning is worthwhile in the sense that it buys something proportional to the effort. It also means a mistake costs proportionally, where a mistake near an optimum would cost only quadratically — which is the argument for leaving a standard matrix alone that this collection had not previously been able to make quantitatively.

Where the model stops

This is one objective. The adaptation cost is this collection’s own construction: a mean over a census of changes of light, over a family of reflectances, with a von Kries gain applied completely. Every essay in this phase has now audited one of those choices, and the shape of the surface inherits all of them.

And the published matrices were not fitted to it. No matrix is right everywhere, and each of these was right about something different. Bradford was fitted to corresponding-colour data — human judgements under two lights — which this collection does not hold; CAT16 was designed to fix CAT02’s negative gains; Hunt–Pointer–Estévez is a set of cone fundamentals. None of them is an attempt to minimise the quantity being differentiated here, so finding that they are not at its minimum is not a criticism of any of them.

What it does establish is a reading rule. Anybody who computes a curvature of any objective at a matrix somebody else chose — which is a common thing to do, since the matrices are the interesting points — should check the gradient first and report the ratio.

Walking downhill from each published basis under the second objective is the check that the first step is the long one there too.

Downhill from every published matrix, one step at a time. Each curve is a steepest-descent walk from one of this collection's published bases, plotted as the objective against the distance walked in the nine coefficients. The horizontal line is the optimum. The first step of each walk is the long one — XYZ scaling closes 87 per cent of its whole gap in one — and every walk then flattens without reaching the line, because the valley floor is nearly flat and the steepest direction is nearly across it. Bradford starts closest and closes least: it is already in the flat part.
Fig. 5 A steepest-descent walk from each published basis, plotted as the discrimination objective against the distance walked in the nine coefficients, with the optimum as a horizontal line. The first step of each walk is the long one, which is the whole shape of the result.

The generalisation

The rule is short and generalises past colour entirely.

Before reading a Hessian, check where it is being taken. At a stationary point the curvature is the leading term and the eigenvalues are the shape. At any other point the gradient is the leading term over a radius that can be computed in one line, and if that radius is small compared with the distances being discussed, the curvature is a correction being read as a description.

The second half is about what to report instead. The gradient’s length and t* are the two numbers that describe a non-stationary point usefully, and both are cheaper than the Hessian that is usually reported alone.

Exactly flat everywhere, and eigenvectors at one point only. Two columns over the same nine places. On the left, how far from zero the objective's second derivative is along a row-scaling direction, on a logarithmic axis — it is between 10⁻⁹ and 10⁻⁶ of the largest eigenvalue at every one of them, which is a numerical zero. Scaling a row of the basis is a straight line along which the cost does not change, and that is true at every point, not only at the optimum. On the right, the angle between those three directions and the Hessian's own three smallest eigenvectors: 0.025 degrees at the optimum and up to 88 away from it. An invariance is a property of the function; being an eigenvector is a property of the function at a minimum, and the two coincide only where everybody computes.
Fig. 6 The quadratic form on an invariant direction, and the angle between those directions and the Hessian’s null space. Both are properties of the same points this essay differentiates, and only one of them survives the move.
Downhill from every published matrix, one step at a time. Each curve is a steepest-descent walk from one of this collection's published bases, plotted as the objective against the distance walked in the nine coefficients. The horizontal line is the optimum. The first step of each walk is the long one — XYZ scaling closes 41 per cent of its whole gap in one — and every walk then flattens without reaching the line, because the valley floor is nearly flat and the steepest direction is nearly across it. Bradford starts closest and closes least: it is already in the flat part.
Fig. 7 Steepest-descent walks from every published basis, plotted as the objective against the distance walked. The first step of each is the long one, which is what a dominant linear term looks like from the side.

That last picture is the same result seen as a path rather than as a radius, and it makes the practical shape of it obvious: the walks fall steeply and then flatten, and every one of them flattens before reaching the line. A dominant linear term is a promise that the first step is worth taking and the tenth is not, which is what the next essay measures directly.

And the third is about signs. A decomposition that returns magnitudes is the right tool at a minimum and hides the most important fact anywhere else. Carrying the signed form alongside costs one dot product per direction and is the difference between describing a bowl and describing a saddle.

Who found it, and when

Every part of the arithmetic is standard: a Taylor expansion, a comparison of terms, and the observation that a symmetric indefinite matrix has directions of both signs. The trust-region literature has the same comparison built into it, since a trust-region method’s whole job is to decide how far the model can be believed, and it is decided adaptively rather than reported.

What is unusual is having a reason to differentiate at somebody else’s matrix. In optimisation the interesting point is the optimum and everything else is a waypoint; here the interesting points are the ones in use, and none of them was chosen by minimising this. That is a common situation in any field that models a designed artefact and it does not have a standard practice attached.

Where the ladder goes next

If the surface at a published matrix is a slope, the obvious next question is what happens when it is walked down — how much a step is worth, how far the walk goes, and whether the direction it sets off in points at the answer.

It does not. At every published transform, steepest descent starts off between eighty and eighty-seven degrees away from the straight line to the optimum, and that turns out not to be an artefact of the three invisible directions.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

The Bradford transformCAT16Chromatic adaptationCondition numberDeclared inputEigenvalueThe ICC profileQuadratic form