The slope arrives before the bowl
Assumes How long is the bowl, Only the flat directions keep their names and No matrix is right everywhere.
The matrices colour management uses are not the bottom of anything, and a curvature taken at a point with a slope describes the wrong thing.
The claim
Every adaptation transform in use is a point where this collection’s objective has a gradient, and within the neighbourhood a Hessian describes, that gradient is the larger term. So the eigenvalues at a published matrix describe a bowl the reader will not meet on the way to anywhere.
- At a minimum the gradient vanishes and the Hessian is everything. At Bradford it does not: the cost rises linearly at first and quadratically only later.
- The two terms are equal at a radius of 0.035 to 0.15 in the nine coefficients, for six of the seven bases measured.
- The distance to the optimum is 1.0 to 2.1. So the curvature catches the slope at three to eleven per cent of the way there.
- XYZ scaling is the exception, at 108 per cent — its slope is so steep and its curvature along the slope so slight that the quadratic term never catches up before the optimum is passed.
- And every published matrix is at a saddle: three or four of the nine directions curve downwards, which a table of eigenvalue magnitudes cannot show because a magnitude has lost the sign.
Where the question came from
The previous round measured the shape of an optimum and got a great deal from it — a rank, a null space, an eigenbasis that explains why one constraint costs seventy per cent and another costs nothing. It left one line in the record: every published transform sits outside the neighbourhood a Hessian describes, and what the objective looks like at Bradford or CAT16 is a different set of six numbers and has not been computed.
It has now, and the first thing the numbers say is that the question was slightly wrong. At a published matrix, six numbers are not what describes the surface. Nine numbers and a direction are — the gradient first, and the curvature second, in that order of importance over the distances that matter.
That is worth saying plainly because it is easy to slip past. Everything this collection knows about the shape of this objective is knowledge about two points, and every matrix a colour management system loads is somewhere else.
The two terms
Along a unit direction from a point, a smooth function changes by the gradient’s component times the distance, plus half the curvature along that direction times the distance squared, plus higher terms.
The two terms are equal at t* = 2|gᵀd| / |dᵀHd|. Below that radius the linear term dominates and the surface is a slope; above it the quadratic term dominates and the surface is a bowl. The number to compare t* with is how far the point is from the answer.
Along the steepest direction the radii are: 0.035 at Bradford, 0.069 at the discrimination optimum, 0.095 at Hunt–Pointer–Estévez, 0.099 at the receptor construction, 0.134 at CAT16, 0.152 at CAT02, and 1.12 at XYZ scaling. The distances to the optimum, measured in the six directions the objective can see, are 1.09, 1.78, 2.06, 2.08, 2.04, 1.48 and 1.04.
So on six of the seven the surface is a slope over ninety per cent or more of the journey.
Bradford is the one worth pausing on, because it is the transform colour management actually uses and it has the smallest ratio in the table at three per cent. It also has the smallest gradient of the published four, at 4.0 — so it is the nearest to being at rest, and the reason its ratio is small is that its curvature along the steep direction is large rather than that its slope is.
A small t* is therefore not a sign of a bad basis. It is a sign of a basis in a steep-walled part of the surface, which is where a well-behaved matrix ought to be: the walls are what hold it in place.
The radius is measured along a direction the journey does not take
The essay’s central quantity is a radius along the steepest direction, and its central comparison is against the distance to the optimum. Those two are not on the same line, and the closing section says by how much: steepest descent at every published transform sets off 80 to 87 degrees from the straight line to the optimum.
So t* is a statement about a direction almost perpendicular to the one the reader is travelling in,
and the ratio it produces — three per cent at Bradford — is not the fraction of the journey over which
the surface is a slope. It is the fraction of a different journey.
The relevant radius can be estimated from the essay’s own numbers. Bradford’s gradient is 4.0 and its excess over the optimum is 1.14 − 0.97 = 0.17, at a distance of 1.09. Taking the departure angle at 83 degrees, the directional slope along the line to the optimum is 4.0 × cos 83° = 0.49, so a purely linear surface would give an excess of 0.53 — three times what there is.
Fitting a quadratic along that line to reproduce the actual 0.17 puts the two terms equal at t* = 1.60, which is 147 per cent of the distance. Along the path that matters, Bradford behaves like XYZ scaling: the linear term leads the whole way and the quadratic never catches it.
| departure angle | slope along the line | linear term over the path | t* as a share of the distance |
|---|---|---|---|
| 80° | 0.695 | 0.757 | 129% |
| 83° | 0.487 | 0.531 | 147% |
| 87° | 0.209 | 0.228 | 392% |
The conclusion survives and the number that supports it does not. At a published matrix the slope is the larger term over the distances that matter is true along the path — more strongly than the essay claims, since the ratio is above one rather than at three per cent. What is not true is that the three per cent measures it.
The quadratic term is not confined to the last few per cent
The same arithmetic contradicts one sentence directly. Almost all of that excess is accounted for by the linear term over the path back and the bowl is a description of the last few per cent of the journey are both stated, and the excess says otherwise.
A linear extrapolation along the path gives 0.53 and the actual excess is 0.17, so the quadratic term removes 68 per cent of what the linear term alone would predict — at the 83-degree reading, and between 25 and 78 per cent across the plausible range of angles.
That is not a correction confined to the end of the journey. It is a term of the same order as the one it corrects, integrated over the whole path, and it is what makes the excess three times smaller than a plane would give. The surface is a slope in the sense that the linear term never loses its lead, and a bowl in the sense that the quadratic term is doing most of a factor of three. Both are true and the essay states only the first.
What sets t*, since it is not the gradient
Two sentences in the essay say the ordering is dominated by the gradient — they order the bases identically because both are dominated by the gradient. The seven bases do not support it.
The rank correlation between gradient length and t* is +0.04, and between gradient length and
t* as a share of the distance it is also +0.04. Both are nothing. Hunt–Pointer–Estévez has the
largest gradient in the table and the third smallest radius; CAT16 has the smallest gradient and the
third largest.
What sets t* is the ratio, and the curvature is the more variable half of it. Recovering the
curvature along each steepest direction from t* = 2|g| / c:
| basis | |g| | curvature along it | t* |
|---|---|---|---|
| CAT16 | 3.3 | 49 | 0.134 |
| XYZ scaling | 15.7 | 28 | 1.120 |
| CAT02 | 6.9 | 91 | 0.152 |
| the discrimination optimum | 7.3 | 212 | 0.069 |
| Bradford | 4.0 | 229 | 0.035 |
| the receptor construction | 17.3 | 349 | 0.099 |
| Hunt–Pointer–Estévez | 23.5 | 495 | 0.095 |
The curvatures span a factor of 17.7 and the gradients a factor of 7.1, so the curvature is the larger source of variation. XYZ scaling’s exceptionalism is entirely in this column — its curvature along its own steepest direction is 28, the smallest in the table by a factor of 1.8, which is what the essay says in words and can now say in a number.
And the Bradford-against-CAT16 contrast the essay draws checks out exactly: the two have the smallest gradients, 4.0 and 3.3, and Bradford’s curvature is 4.6 times CAT16’s, which is the whole of why its radius is 3.8 times smaller.
What that means for a table of eigenvalues
A curvature is not wrong at a point with a slope — it is a real second derivative and it is exactly what it says. What it is not is the leading description.
Reading a table of eigenvalues at Bradford as the shape of the objective there invites two mistakes at once. The first is about size: a step of 0.1 from Bradford changes the cost by about 0.1 × |g| = 0.4, and by about 0.005 × λ from the quadratic term, so the quadratic term is a twentieth of the change. The second is about direction: the steepest direction and the stiffest direction are different directions, and a picture of a bowl points at the second.
The useful pair of numbers at a published matrix is therefore the gradient’s length and the radius at which the curvature catches it, and both are cheap: eighteen evaluations for the gradient and a hundred and fifty-four for the Hessian.
The gradients are 3.3 at CAT16, 4.0 at Bradford, 6.9 at CAT02, 7.3 at the discrimination optimum, 15.7 at XYZ scaling, 17.3 at the receptor construction and 23.5 at Hunt–Pointer–Estévez. Against 5 × 10⁻⁴ at the optimum, which is a finite-difference zero.
There is a version of this that a reader may find more intuitive, in units of the objective rather than of the coefficients. The excess of a published matrix over the optimum is between 0.17 and 1.40 ΔE00 — Bradford is 1.14 against the optimum’s 0.97 — and almost all of that excess is accounted for by the linear term over the path back. The bowl is a description of the last few per cent of the journey, and the first ninety-five is a straight run downhill.
The exception, which is instructive
XYZ scaling is the one basis whose curvature does not catch its slope before the optimum is reached, at a ratio of 108 per cent, and the reason is not that its slope is the steepest.
Its gradient is 15.7, which is the second largest. What makes it exceptional is the curvature along that gradient, which is unusually small: the direction XYZ scaling falls fastest in is one the objective is nearly flat along, so the linear term keeps its lead the whole way.
That is the same basis that is exceptional in every measurement in this thread: the largest excess over the optimum, the longest trust radius, the most rotated stiff plane at 55 degrees. It is far enough from everything else to behave like a different problem, which is a reasonable thing for the oldest mistake still shipping to do.
Every published matrix is at a saddle
The signed curvature along each principal direction says something a table of magnitudes cannot, and it is the more alarming half of this essay.
At the optimum, all nine directions curve upwards or are zero — six positive and three at the numerical floor — which is what being a minimum means. At every published basis, three or four of the nine curve downwards: there are directions in which the objective falls away quadratically as well as linearly.
A singular value decomposition returns magnitudes, so a table of eigenvalues at Bradford looks exactly like a table at the optimum: nine positive numbers spanning several orders of magnitude. The sign is the whole difference between a bowl and a saddle and it is thrown away by the decomposition that produces the picture.
Counting them needs a threshold that is not the roundoff floor, because the three flat directions have an exactly zero form which a second difference computes as 10⁻⁸ with whichever sign the noise had. At a threshold of 10⁻⁵ — inside the gap between the smallest real eigenvalue at 10⁻⁴ and the numerical zeros at 10⁻⁸ — the count is zero at the optimum and three or four everywhere else.
What was computed, and how
Gradients and Hessians by central differences at the step this collection measured a window for, at eight named bases. The distance to the optimum has the three invisible directions projected out of it, because the optimum is not a point: three of the nine directions are exactly flat, so the objective’s minimiser is a three-parameter family and a search returns one arbitrary member of it.
Without that projection every distance would carry a component nobody chose. With it, the distances fall by five to thirty per cent and the ratios in this essay are statements about the part of the displacement the objective can see.
The two quantities in this thread are related and not the same. t* is where the bowl catches the slope; the trust radius is where the two of them together stop describing the surface. They order the bases identically because both are dominated by the gradient — which is itself a small finding, and the reason to report both rather than either.
Why this matters for a delivery chain
The published transforms are not a curiosity: they are what an ICC profile’s chromatic adaptation applies when a colour is carried between two white points, and Bradford in particular is in essentially every profile in circulation.
So the shape of this objective at Bradford is the shape at the point where all the arithmetic actually happens, and the practical consequence is a small one about interpretation. A sentence of the form this transform is nearly optimal, and the objective is flat around it is two sentences, and only the first is about Bradford: the flatness is a statement about the optimum, and Bradford is on a slope four ΔE00 units per unit of coefficient steep.
What that permits is a useful piece of advice for anybody tempted to tune one. A small perturbation of a published matrix moves the objective linearly, so tuning is worthwhile in the sense that it buys something proportional to the effort. It also means a mistake costs proportionally, where a mistake near an optimum would cost only quadratically — which is the argument for leaving a standard matrix alone that this collection had not previously been able to make quantitatively.
Where the model stops
This is one objective. The adaptation cost is this collection’s own construction: a mean over a census of changes of light, over a family of reflectances, with a von Kries gain applied completely. Every essay in this phase has now audited one of those choices, and the shape of the surface inherits all of them.
And the published matrices were not fitted to it. No matrix is right everywhere, and each of these was right about something different. Bradford was fitted to corresponding-colour data — human judgements under two lights — which this collection does not hold; CAT16 was designed to fix CAT02’s negative gains; Hunt–Pointer–Estévez is a set of cone fundamentals. None of them is an attempt to minimise the quantity being differentiated here, so finding that they are not at its minimum is not a criticism of any of them.
What it does establish is a reading rule. Anybody who computes a curvature of any objective at a matrix somebody else chose — which is a common thing to do, since the matrices are the interesting points — should check the gradient first and report the ratio.
Walking downhill from each published basis under the second objective is the check that the first step is the long one there too.
The generalisation
The rule is short and generalises past colour entirely.
Before reading a Hessian, check where it is being taken. At a stationary point the curvature is the leading term and the eigenvalues are the shape. At any other point the gradient is the leading term over a radius that can be computed in one line, and if that radius is small compared with the distances being discussed, the curvature is a correction being read as a description.
The second half is about what to report instead. The gradient’s length and t* are the two numbers that describe a non-stationary point usefully, and both are cheaper than the Hessian that is usually reported alone.
That last picture is the same result seen as a path rather than as a radius, and it makes the practical shape of it obvious: the walks fall steeply and then flatten, and every one of them flattens before reaching the line. A dominant linear term is a promise that the first step is worth taking and the tenth is not, which is what the next essay measures directly.
And the third is about signs. A decomposition that returns magnitudes is the right tool at a minimum and hides the most important fact anywhere else. Carrying the signed form alongside costs one dot product per direction and is the difference between describing a bowl and describing a saddle.
Who found it, and when
Every part of the arithmetic is standard: a Taylor expansion, a comparison of terms, and the observation that a symmetric indefinite matrix has directions of both signs. The trust-region literature has the same comparison built into it, since a trust-region method’s whole job is to decide how far the model can be believed, and it is decided adaptively rather than reported.
What is unusual is having a reason to differentiate at somebody else’s matrix. In optimisation the interesting point is the optimum and everything else is a waypoint; here the interesting points are the ones in use, and none of them was chosen by minimising this. That is a common situation in any field that models a designed artefact and it does not have a standard practice attached.
Where the ladder goes next
If the surface at a published matrix is a slope, the obvious next question is what happens when it is walked down — how much a step is worth, how far the walk goes, and whether the direction it sets off in points at the answer.
It does not. At every published transform, steepest descent starts off between eighty and eighty-seven degrees away from the straight line to the optimum, and that turns out not to be an artefact of the three invisible directions.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- An extremum is not a sample chromatic adaptation · condition number · eigenvalue · quadratic form
- The census is a construction too the bradford transform · cat16 · chromatic adaptation · declared input
- The widths were free because nothing else was asked chromatic adaptation · condition number · declared input · eigenvalue
- Two tolerances do not meet in a tolerance chromatic adaptation · condition number · declared input · eigenvalue
- Best on the average, undefined at the edge cat16 · chromatic adaptation · the icc profile
- Dividing by the paper the bradford transform · chromatic adaptation · the icc profile
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
The Bradford transformCAT16Chromatic adaptationCondition numberDeclared inputEigenvalueThe ICC profileQuadratic form