Matching and measuring

Only the flat directions keep their names

Three of the nine numbers a colour match leaves free do nothing, and they do nothing everywhere — exactly, at every basis in this collection's table. They are the objective's own principal directions at one point only, and everywhere else the directions carrying the curvature have turned by tens of degrees.

Assumes The rank is the invariance, The three numbers a gain cannot see and How far a quadratic can be believed.

Three of the nine numbers do nothing, and the three that do nothing are the only part of the picture that survives being moved.

Exactly flat everywhere, and eigenvectors at one point only. Two columns over the same nine places. On the left, how far from zero the objective's second derivative is along a row-scaling direction, on a logarithmic axis — it is between 10⁻⁹ and 10⁻⁶ of the largest eigenvalue at every one of them, which is a numerical zero. Scaling a row of the basis is a straight line along which the cost does not change, and that is true at every point, not only at the optimum. On the right, the angle between those three directions and the Hessian's own three smallest eigenvectors: 0.025 degrees at the optimum and up to 88 away from it. An invariance is a property of the function; being an eigenvector is a property of the function at a minimum, and the two coincide only where everybody computes.
Fig. 1 Two columns over nine places: how far from zero the objective’s second derivative is along a row-scaling direction, and the angle between those directions and the Hessian’s own smallest eigenvectors. The first column is a numerical zero everywhere; the second is not.

The claim

An invariance is a property of a function and holds at every point. Being an eigenvector is a property of a function at a minimum. The two coincide at the one place anybody computes, which is why they are easy to confuse.

  • Scaling a row of the basis changes nothing, exactly. It is a straight line in the nine coefficients along which the cost is constant, so the gradient has no component along it and the quadratic form vanishes on it — at every point in this collection’s table, to between 10⁻⁹ and 10⁻⁶ of the largest eigenvalue.
  • Those three directions are the Hessian’s own null space at the optimum, to a principal angle of 0.025 degrees. That is the previous round’s rank result and it is exact.
  • Away from the optimum they are not eigenvectors at all, and the angle runs to 88 degrees — very nearly orthogonal to the directions the curvature actually singles out.
  • The rank is six at the optimum and nine everywhere else, which the previous round noticed and did not pursue.
  • And the stiff directions turn too: the two-dimensional subspace carrying most of the curvature is 12 to 55 degrees away from its counterpart at the optimum, so a tolerance computed at one basis does not transfer to another.

What the three flat directions are

A colour match leaves nine numbers free — the basis a von Kries gain runs along is a 3×3 and nothing fixes it — and three of those nine do nothing at all. Scaling a row of the basis by any factor scales the corresponding gain by its reciprocal, and the two cancel exactly.

That is an invariance and it was proved: the cost is identically constant along each of the three, not approximately so and not only near an optimum. Scaling row i by 1 + t traces a straight line in the nine coefficients, and the cost along that line is a constant function of t.

The previous round measured the consequence at the optimum and found something stronger than the proof: the Hessian’s rank is exactly six, its three null directions span the row scalings to a tenth of a degree, and the gap between the sixth eigenvalue and the seventh is four thousand. A rank rules out every direction there is, where a proof rules out only the directions it names.

Why the rank is nine anywhere else

The record of that measurement contains a sentence that has been sitting there unexplored: at the receptor basis, which is neither objective’s optimum, the rank is nine.

The reason is a two-line argument and it is the whole of this essay. Along a flat direction d the cost is constant, so dᵀ∇f = 0 and dᵀHd = 0 at every point. At a minimum the Hessian is positive semidefinite, and a positive semidefinite matrix with dᵀHd = 0 must have Hd = 0 — so the direction is an eigenvector with eigenvalue zero, the rank drops, and the null space is the flat subspace.

The three directions the curvature cannot see are the three row scalings. Three horizontal bars, one per principal angle between the null space of the Hessian of the adaptation residual and the three directions along which a row of the basis is rescaled. All three angles are under a tenth of a degree — 2.5e-2°, 7.9e-4°, 2.5e-4° — so the two three-dimensional subspaces are the same subspace. This is the algebraic invariance that a von Kries gain cannot see the scale of a row, measured as a geometry rather than assumed.
Fig. 2 The three directions the curvature cannot see at the optimum, compared with the three row scalings, as principal angles. At a minimum they are the same subspace.

Away from a minimum the Hessian is indefinite and the implication fails. A vector can satisfy dᵀHd = 0 without Hd = 0 — it merely has to lie on the cone where the quadratic form vanishes, which for an indefinite matrix is a whole cone rather than a subspace. So the three flat directions still make the form vanish, and H maps them somewhere non-zero, and the rank stays at nine.

That is not a subtlety about numerical precision. It is a statement about what an eigenvector is, and it means the three most robust facts about this objective are invisible to an eigen-decomposition taken anywhere but at the bottom.

The measurement

Both halves are measured at every basis in the table, at the same step, so the two columns can be read against each other.

The quadratic form on a scaling is between 1.4 × 10⁻⁹ and 1.2 × 10⁻⁶ of the largest eigenvalue, at every one of the nine places. That is a numerical zero for a quantity computed by second differences — the finite-difference floor for this objective at this step is around 10⁻⁸ relative — and it is the same at the optimum as at XYZ scaling.

The principal angle between those three directions and the Hessian’s three smallest eigenvectors is 0.025 degrees at the optimum and, elsewhere: 46 degrees at Bradford, 48 at XYZ scaling, 74 at the discrimination optimum, 78 at Hunt–Pointer–Estévez, 86 at CAT02, 88 at the receptor construction, 88 at CAT16.

Eighty-eight degrees is very nearly a right angle. At CAT16 the objective’s three smallest curvatures point in directions almost perpendicular to the three that are actually flat, which means an analysis that read the small eigenvalues there as the directions that do not matter would have named three directions the cost genuinely does depend on and missed three it does not.

There is a second way to state the same measurement that makes its shape clearer. The three flat directions are known in closed form at every point — scaling row i is a vector with three non-zero entries and six zeros, normalised — so they can be written down without computing anything. The three smallest eigenvectors have to be computed and are different at every point.

One of those two objects is a property of the parameterisation and the other is a property of the surface, and only at the optimum do they agree. That is a clean division and it suggests a habit: where an invariance is known analytically, use the analytic directions rather than the computed ones, because the analytic ones are right everywhere and the computed ones are right at one point.

The stiff directions turn as well

The same measurement applied to the other end of the spectrum gives a smaller number and a more practical consequence.

The two-dimensional subspace carrying most of the curvature — the stiff plane — is compared between each basis and the optimum by principal angle: 12.5 degrees at Bradford, 17.2 at the receptor construction, 17.8 at Hunt–Pointer–Estévez, 20.7 at CAT16, 27.9 at the discrimination optimum, 32.9 at CAT02, and 55.3 at XYZ scaling.

The stiff directions turn as the basis moves; the flat ones do not. One bar per basis: the largest principal angle between the two-dimensional subspace carrying most of the objective's curvature there and the same subspace at the optimum. The optimum's own bar is zero, which is the control. Everywhere else the angle is between 12 and 55 degrees, so a tolerance computed from the curvature at one basis does not transfer to another. Two dimensions rather than one because the two largest curvatures are close enough at several of these points that a single leading eigenvector is not a stable object to compare.
Fig. 3 The largest principal angle between each basis’s stiff two-dimensional subspace and the optimum’s. The optimum’s own bar is zero, which is the control.

Two more views say what the turning costs anybody who wants to use a basis as a starting point rather than as an answer.

A quadratic is believed least far at the one place anybody takes one. One bar per basis: the radius, in the nine coefficients, within which the second-order model predicts the objective to within ten per cent in every one of eighteen directions. The shortest bar is the objective's own optimum, at 2.3×10⁻², and the longest is XYZ scaling at 1.1×10⁻¹ — several times further. The reason is not that the model is worse at a minimum but that it has less to do there: away from one the linear term is exact and carries most of the change, so a ten per cent error in the prediction takes longer to accumulate. It does not make a Hessian at a minimum wrong; it says the picture drawn from it describes the smallest neighbourhood in the table.
Fig. 4 The radius over which a second-order model of the objective holds at each basis. Where the stiff subspace has turned most, the radius is smallest — which is the same statement twice.
Downhill from every published matrix, one step at a time. Each curve is a steepest-descent walk from one of this collection's published bases, plotted as the objective against the distance walked in the nine coefficients. The horizontal line is the optimum. The first step of each walk is the long one — XYZ scaling closes 41 per cent of its whole gap in one — and every walk then flattens without reaching the line, because the valley floor is nearly flat and the steepest direction is nearly across it. Bradford starts closest and closes least: it is already in the flat part.
Fig. 5 And steepest-descent walks from each of them. A direction that has kept its name is a direction a walk can follow; the rest have to be found again from where the walk actually is.

Two dimensions rather than one, because the two largest curvatures are close enough at several of these points that a single leading eigenvector is not a stable object to compare — comparing a subspace is the standard repair and it is the honest one.

The consequence is that a tolerance computed at one basis does not transfer to another. What a constraint costs is computed as the quadratic form in the direction the constraint points, in the eigen-coordinates at the optimum; the same constraint applied at Bradford points into a differently oriented eigenbasis and costs something else.

The rotation has a mechanism, and it is the same one that makes the whole surface awkward. The objective’s six real eigenvalues span a factor of eight hundred and ninety at the optimum, so the ordering of the middle ones is decided by small differences — and a small change in the point can swap two nearby eigenvalues and rotate the plane they span. A subspace spanned by two nearly equal eigenvalues is not a stable object, and comparing subspaces rather than vectors is the standard defence rather than a refinement.

XYZ scaling is the outlier at 55 degrees, and it is the outlier in every measurement in this thread: it has the largest excess over the optimum, the second largest gradient, the longest trust radius and the most turned stiff plane. It is the oldest mistake still shipping, and it is far enough from everything else that it behaves like a different problem.

What was computed, and how

Everything above is a Hessian and a gradient by central differences at the step this collection measured a window for, at eight named bases, plus one control at the objective’s own optimum.

The signed curvature along each principal direction is carried alongside the singular values, because a singular value has lost the sign and the sign is the whole difference between a bowl and a saddle. Counting the negative ones needs a threshold that is not the roundoff floor: the three flat directions have an exactly zero form which a second difference computes as something of order 10⁻⁸ with whichever sign the noise had, so a threshold at 10⁻⁹ counts those as negative curvature at a minimum, where there is none.

At a published matrix the surface curves downwards in some directions. Nine markers per series, on a logarithmic axis of magnitude: the objective's curvature along each of its own principal directions, at Bradford and at the optimum. Filled markers are directions that curve upwards and open ones curve downwards. At the optimum every one is upwards or zero, which is what being a minimum means. At Bradford 3 of the nine curve downwards, so the point is a saddle — there are directions in which the objective falls away quadratically as well as linearly, and describing that neighbourhood as a bowl is wrong in a way an eigenvalue table does not show, because a table of magnitudes has lost the sign.
Fig. 6 The signed curvature along each of the nine principal directions, at a published matrix and at the optimum. Filled markers curve upwards and open ones curve downwards; at the optimum none is open.

At 10⁻⁵, which is inside the gap — the smallest real eigenvalue is 10⁻⁴ of the largest and the numerical zeros are 10⁻⁸ — the counts are: zero negative directions at the optimum, and three or four at every published basis. Every published transform is at a saddle of this objective, which is unsurprising once said and is not visible in a table of magnitudes.

The two turns are independent

The essay measures two rotations — how far the flat directions are from being eigenvectors, and how far the stiff plane has swung — and reads them as two symptoms of one displacement from the optimum. Set side by side they are not.

basis null-space angle stiff-plane angle
Bradford 46° 12.5°
XYZ scaling 48° 55.3°
discrimination optimum 74° 27.9°
Hunt–Pointer–Estévez 78° 17.8°
CAT02 86° 32.9°
receptor construction 88° 17.2°
CAT16 88° 20.7°

The rank correlation between the two columns is −0.04, which over seven points is nothing at all. The two bases furthest from having the flat directions as eigenvectors — the receptor construction and CAT16, both at 88° — have among the least rotated stiff planes, at 17 and 21 degrees. XYZ scaling has the most rotated stiff plane in the table and the second smallest null angle.

So the flat directions and the stiff plane come loose independently, and neither is a proxy for distance from the optimum. That matters for the practical advice the essay gives, because the two failures have different remedies: the null-space rotation is repaired by using the analytic directions instead of the computed ones, which costs nothing and works at every point; the stiff-plane rotation has no such repair, because there is no closed form for where the curvature points.

It also means a reader cannot use one as a warning about the other. A basis whose flat directions are almost eigenvectors — Bradford, at 46° — is not thereby a basis whose stiff plane can be trusted, though in Bradford’s case it happens to be the least rotated. Nothing links them.

Two figures for the same eigenvalue

One number appears twice at different magnitudes, and the discrepancy is worth resolving because the threshold argument that closes the essay rests on it.

The stiff-plane section says the six real eigenvalues span a factor of 890 at the optimum, which puts the smallest real one at 1.12 × 10⁻³ of the largest. The threshold section says the smallest real eigenvalue is 10⁻⁴ of the largest. Those are eleven times apart.

The threshold argument survives either figure, which is worth saying before anything else. The numerical zeros run to 1.2 × 10⁻⁶ and a cut at 10⁻⁵ sits above them and below 1.12 × 10⁻³, with a factor of eight of margin on one side and a hundred on the other. Under the 10⁻⁴ figure the margins are eight and ten. Either way the cut is inside the gap, and the count of negative directions is not in doubt.

But the two figures cannot both describe the sixth eigenvalue, and the 890 is the one this collection uses elsewhere — it is quoted again in the essay comparing the observer’s nine numbers with a camera’s six. The margin at the cut is therefore about a hundredfold rather than tenfold, which makes the threshold choice less delicate than the essay presents it as.

The arithmetic around it does hold together. A gap of four thousand below the sixth eigenvalue puts the seventh at 2.8 × 10⁻⁷ of the largest, which lands inside the 1.4 × 10⁻⁹ to 1.2 × 10⁻⁶ band the flat directions’ quadratic forms occupy — exactly as it must, since the seventh eigenvalue is one of the flat directions. Two independently reported quantities landing on top of each other is the check that the decomposition and the invariance are describing the same three directions.

Eight places, or nine

A smaller thing, and it is the kind an essay about counting should not have. The measurement section says the quadratic form is a numerical zero at every one of the nine places, and the method section says the computation runs at eight named bases plus one control.

The angle list names seven: Bradford, XYZ scaling, the discrimination optimum, Hunt–Pointer–Estévez, CAT02, the receptor construction and CAT16. With the optimum that is eight, and the stiff-plane list names the same seven. A ninth place exists in the counts and appears in neither list, so a reader cannot reconstruct which bases were measured from what is printed.

It changes no conclusion — every listed basis behaves the same way and one more would not alter a correlation of −0.04 or a range of 46 to 88 degrees. It is worth correcting because the whole essay is an argument for preferring quantities that can be written down over quantities that have to be computed, and a list that does not match its own count is the smaller version of the same complaint.

The discrimination objective has the same flat directions for the same reason, and confirming that is what makes the argument about the parameterisation rather than about one objective.

Exactly flat everywhere, and eigenvectors at one point only. Two columns over the same nine places. On the left, how far from zero the objective's second derivative is along a row-scaling direction, on a logarithmic axis — it is between 10⁻⁹ and 10⁻⁶ of the largest eigenvalue at every one of them, which is a numerical zero. Scaling a row of the basis is a straight line along which the cost does not change, and that is true at every point, not only at the optimum. On the right, the angle between those three directions and the Hessian's own three smallest eigenvectors: 0.003 degrees at the optimum and up to 88 away from it. An invariance is a property of the function; being an eigenvector is a property of the function at a minimum, and the two coincide only where everybody computes.
Fig. 7 The discrimination objective’s second derivative along a row-scaling direction at each of nine places, on a logarithmic axis. It sits between 10⁻⁹ and 10⁻⁶ of the largest eigenvalue everywhere, which is numerically zero and is the same flatness the adaptation objective has.

What this rescues, and what it costs

The result reads as a warning and it also settles something, so the balance is worth drawing.

It settles the previous round’s loose end. At the receptor basis the rank is nine was recorded as an oddity to be careful about, and it is now a theorem with a measurement under it. Nothing about the rank result needs restating: it was always a statement about the Hessian at the optimum, and it was always carefully worded.

It costs the transferability of everything else. The eigen-coordinates at the optimum are the frame in which a constraint’s excess is a sum of squares and in which one per cent of one objective buys a quarter of the other. Both of those are statements at the optimum and both were checked against the true objective there. Neither transfers to Bradford, and the angle in this essay is how far it fails to.

And it hands back a cheap diagnostic. A known invariance evaluated as a quadratic form is a convergence test that costs one second difference: at a converged minimum the invariant direction is an eigenvector, and anywhere else it is not. That is a stronger test than a gradient norm, because a gradient norm is small near a minimum and also small on a plateau.

What the curvature does elsewhere is the other half, and the second objective answers it the same way.

The stiff directions turn as the basis moves; the flat ones do not. One bar per basis: the largest principal angle between the two-dimensional subspace carrying most of the objective's curvature there and the same subspace at the optimum. The optimum's own bar is zero, which is the control. Everywhere else the angle is between 24 and 79 degrees, so a tolerance computed from the curvature at one basis does not transfer to another. Two dimensions rather than one because the two largest curvatures are close enough at several of these points that a single leading eigenvector is not a stable object to compare.
Fig. 8 The largest principal angle between the subspace carrying most of the discrimination objective’s curvature at each basis and the same subspace at the optimum. The optimum’s own bar is zero — that is the control — and everywhere else the subspace has turned.

Where the model stops

The three flat directions are flat for this objective and for its neighbour, and not for everything. The invariance holds because a von Kries gain divides by the same row it multiplies by; an objective that used the basis without inverting it — a projection, a distance, a fit — would see all nine numbers.

And the angles are between subspaces, not between vectors. A principal angle of 88 degrees between two three-dimensional subspaces means the worst pair of directions is nearly perpendicular; the other two pairs may be much closer, and here they usually are. Reporting the largest is the conservative reading and is what a reader needs to know before treating one basis’s eigen-decomposition as another’s.

The optimum’s own null space is measured, not assumed. The 0.025-degree agreement is a measurement with a finite-difference floor under it, and a version of this arithmetic with a badly chosen step would produce a larger number without anything being wrong. That is why the step’s window was measured before any of this was computed.

The generalisation

The distinction is worth stating in a form that carries past this objective, because it is a genuine source of confusion in any analysis built on a curvature.

An invariance is a global statement: the function does not change along this direction, anywhere. A null eigenvector is a local statement: at this point, the second derivative annihilates this direction. At a minimum of a smooth function the first implies the second. Nowhere else does it.

The confusion is easy because the two coincide exactly where the analysis is normally done. Somebody who computes a Hessian at an optimum, finds three zero eigenvalues, and identifies them with a known symmetry has done something correct — and if they then carry the eigenvectors to another point and expect them still to be the flat directions, they have carried a local object as though it were a global one.

At a published matrix the slope arrives long before the bowl. One row per basis in this collection's table. Each row is a logarithmic axis of distance in the nine coefficients, with two markers: the radius at which the objective's curvature becomes as large as its slope, and the distance from that basis to the optimum. The first is between 3.2 and 108 per cent of the second. So over almost the whole journey from a published matrix to the best one, the surface is a slope and not a bowl — and a table of eigenvalues taken there describes a neighbourhood the optimum is nowhere near. XYZ scaling is the exception, at 1.08 of the distance, because its slope is the steepest in the table.
Fig. 9 Where the curvature catches the slope at each published basis, against the distance to the optimum. A second reason a Hessian at a non-optimum describes something the reader will not meet.

The check is one line and needs no new machinery: evaluate the quadratic form on the known invariant direction and compare it with the largest eigenvalue. If it is zero and the direction is not an eigenvector, the point is not a minimum — which is a useful thing to be able to detect, and is also the fastest available test that a search has actually converged.

Who found it, and when

The linear algebra is entirely standard. That a positive semidefinite matrix has Hd = 0 whenever dᵀHd = 0 is a one-line consequence of the Cauchy–Schwarz inequality applied to the form’s own square root, and it is in every course that touches quadratic forms. That an indefinite matrix has a null cone rather than a null space is equally standard.

What is not standard is checking it, because the situation in which it matters — a symmetry of an objective, examined at a point that is not a stationary point — arises rarely on purpose and often by accident. It arose here because the previous round asked what the objective looks like at the matrices people actually use, which are not optima of anything this collection minimises.

Where the ladder goes next

The flat directions survive being moved and the eigen-decomposition does not, which raises the obvious question about the rest of the picture: what does a curvature at a published matrix describe?

The answer is that it describes a bowl the reader will not meet on the way anywhere, because at every one of those points the slope arrives long before the bowl — the radius at which the quadratic term catches the linear one is a few per cent of the distance to the optimum, and inside that radius the surface is a plane with a tilt.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyChromatic adaptationCondition numberDeclared inputDegrees of freedomEigenvalueQuadratic formThe von Kries transform