Where the model breaks

How far a quadratic can be believed

A second-order model has a radius inside which it describes a surface and outside which it does not, and that radius can be measured. Measured at eight places on one objective, it is smallest at the optimum — the one place anybody ever takes a Hessian.

Assumes How long is the bowl, The rank is the invariance and An extremum is not a sample.

A curvature describes a surface over some distance, and how far that distance reaches is a measurable property rather than a figure of speech.

A quadratic is believed least far at the one place anybody takes one. One bar per basis: the radius, in the nine coefficients, within which the second-order model predicts the objective to within ten per cent in every one of eighteen directions. The shortest bar is the objective's own optimum, at 2.3×10⁻², and the longest is XYZ scaling at 1.1×10⁻¹ — several times further. The reason is not that the model is worse at a minimum but that it has less to do there: away from one the linear term is exact and carries most of the change, so a ten per cent error in the prediction takes longer to accumulate. It does not make a Hessian at a minimum wrong; it says the picture drawn from it describes the smallest neighbourhood in the table.
Fig. 1 The radius over which a second-order model of the adaptation objective predicts it to within ten per cent in every direction, at each of eight bases. The shortest bar is the optimum.

The claim

A quadratic model has a radius of validity, it can be measured rather than estimated, and on this objective it is smallest at the minimum — which is the only point at which anybody ever builds one.

  • The measurement is direct. At a stated radius, compare the model’s prediction with the objective along eighteen directions and take the worst; bisect on the radius until that worst error is ten per cent.
  • At the adaptation optimum the answer is 0.023 in the units of the nine coefficients. At XYZ scaling it is 0.113, five times further.
  • The reason is not that the model is worse at a minimum. It is that a model has less to do away from one: with a slope present, the linear term is exact and carries most of the change, so a relative error takes longer to accumulate.
  • It does not make a Hessian at a minimum wrong. The eigenvalues are exact and the previous round’s results built on them stand.
  • What it does say is that the picture drawn from those eigenvalues describes the smallest neighbourhood in the table, which is worth knowing before reading distances off it.

What a trust radius is

A second-order model of a function at a point is the value, plus the gradient’s contribution, plus half the quadratic form of the curvature. It is exact for a quadratic and approximate for everything else, and the approximation gets worse with distance.

How much worse is a fact about the third and fourth derivatives, which nobody has here and which are not obviously smaller near a minimum than anywhere else. That is the whole reason the answer had to be measured rather than reasoned about.

The bowl the eigenvalues describe and the bowl a sample found. Six points on a logarithmic vertical axis — the distance from the optimum of the adaptation residual to a 5 per cent rise along each of the six directions the objective can see — with a shaded band behind them showing the whole range 24 random directions reported. The eigen-radii run from 1.2e-2 to 3.6e-1, a factor of 29.8. The band runs from 2.2e-2 to 1.8e-1, a factor of 8.0, and sits entirely inside the ends of the true range: a random direction in nine dimensions carries a share of every eigenvector and so reports the middle of the bowl, never an end of it.
Fig. 2 The bowl the eigenvalues describe and the bowl a search actually walks out into, direction by direction. The two agree further in the stiff directions than in the flat ones, which is the opposite of the intuition.

The version of this quantity the previous round measured is per-direction: how far along each eigenvector the predicted radius to a stated rise matches the measured one. That is the right instrument for asking which directions the model handles. This one collapses the direction away deliberately, because the question here is about the point rather than about a direction — and a trust radius taken along one direction is a statement about that direction.

There is a second reason the quantity is worth having, beyond describing a picture’s reach. This collection uses the curvature at an optimum to compute things a reader is meant to act on — which of a device’s numbers a tolerance budget belongs on, and what a constraint on the basis costs. Both of those are statements about a neighbourhood, and both are only as good as the neighbourhood the model covers.

What was measured

At each of eight bases, the objective’s gradient and Hessian are taken by central differences at the step this collection measured a window for. Then, at a trial radius, the model is compared with the truth along the nine eigenvectors and their negatives — eighteen directions which between them span everything — and the worst relative discrepancy decides.

The radius is bisected on its logarithm until that worst error is ten per cent. Eighteen directions and fourteen halvings is two hundred and fifty-two evaluations of the objective per basis, which is cheap enough that the answer needed no cleverness.

At a published matrix the slope arrives long before the bowl. One row per basis in this collection's table. Each row is a logarithmic axis of distance in the nine coefficients, with two markers: the radius at which the objective's curvature becomes as large as its slope, and the distance from that basis to the optimum. The first is between 3.2 and 108 per cent of the second. So over almost the whole journey from a published matrix to the best one, the surface is a slope and not a bowl — and a table of eigenvalues taken there describes a neighbourhood the optimum is nowhere near. XYZ scaling is the exception, at 1.08 of the distance, because its slope is the steepest in the table.
Fig. 3 At each basis, the radius where the curvature becomes as large as the slope, against the distance to the optimum. The two quantities in this essay are different: this one is where the bowl catches the slope, and the trust radius is where the pair of them stop describing the surface.

The step deserves a word, because it is the one place a measurement like this can quietly go wrong. Too large a step and the second difference reports the curvature of something else; too small and it is the roundoff floor amplified by the square of the step. The window here was measured rather than chosen: over a decade of steps the six eigenvalues the objective has move by 0.42 per cent while the three it does not have grow by a factor of a hundred, which is the square law a truncation error obeys on a quantity that is exactly zero.

Taking the worst rather than the mean is the decision that makes the number mean something. A model that is excellent in seventeen directions and hopeless in the eighteenth is a model that will mislead whoever walks in the eighteenth, and averaging would hide exactly that. It also makes the answer conservative, which is the right side to be on for a quantity a reader will use to decide how far to trust a picture.

The answers, and the surprise

The eight radii run from 0.023 at the adaptation optimum to 0.113 at XYZ scaling, with the published transforms between: Hunt–Pointer–Estévez at 0.034, CAT16 at 0.042, Bradford at 0.052, CAT02 at 0.053, the receptor construction at 0.076, the discrimination optimum at 0.095.

The extremes track the gradient rather than the curvature. XYZ scaling has a large gradient and the longest reach; the optimum has no gradient at all and the shortest. How well that holds across the middle of the table is measured further down, and the answer is: not at all.

This was written expecting the opposite. The obvious argument is that a quadratic model is a Taylor expansion and a Taylor expansion is best where the function is best behaved, and a minimum ought to be well behaved. The obvious argument is about the wrong quantity: the criterion is a relative error on the predicted value, and at a minimum the whole predicted change is second order, so a given error in the cubic term is a large fraction of it. Where there is a slope, the first-order term is exact and dominates, so the same cubic error is a small fraction of a much larger predicted change.

That is a statement about the criterion as much as about the surface, and both halves are worth having. A trust radius defined on an absolute error would rank the eight differently. The relative one is the right criterion here because every quantity this collection reports from these surfaces is a ratio or a per-cent change.

Whether the gradient really orders them

That the radii track the gradient is a weak reading of what turns out to be a mostly binary fact.

Ranking the eight places by radius and by the size of the gradient gives a Spearman coefficient of 0.429. Removing the optimum — the one place with no gradient at all — leaves 0.143 over the remaining seven, which is nothing.

The seven, ordered by gradient with their radii beside them: CAT16 at 3.3 and 0.042, Bradford at 4.0 and 0.052, CAT02 at 6.9 and 0.053, the discrimination optimum at 7.3 and 0.095, XYZ scaling at 15.7 and 0.113, the receptor construction at 17.3 and 0.076, and the Hunt–Pointer–Estévez axes at 23.5 and 0.034.

That last pair settles it. HPE has the largest gradient in the table and the second-shortest radius, which no monotone reading of the gradient survives.

So the mechanism is binary rather than graded. Having a gradient at all is worth a factor of two to five on the radius — the optimum’s 0.023 against a median of 0.053 across the seven — and how large the gradient is decides nothing beyond that.

That is a milder claim and it keeps the whole argument intact. A relative error accumulates fastest at a minimum because the predicted change there is entirely second order, and that stops applying as soon as there is any first-order term for it to hide behind. Once there is one, whether it is three units or twenty-three is evidently swamped by something else: the third derivative’s own size in whichever directions the search probes, which nothing here measures and which has no reason to vary smoothly across a table of eight unrelated matrices.

What this does not undo

It would be easy to read this as an attack on the previous round’s results, and it is not, so the boundary is worth drawing precisely.

The eigenvalues at the optimum are exact. They are second derivatives of a smooth function, computed at a step whose window was measured, and cross-checked by the way the three that should be zero fall as the square of the step while the six that should not move by 0.42 per cent over a decade of it. Nothing about a short trust radius touches any of that.

The rank of the Hessian is a statement about the matrix and not about any neighbourhood. The excess a constraint costs is computed as the quadratic form in the direction the constraint points, and its accuracy was measured directly against the true objective at the displacements involved — which is the right check and is not this one.

What the short radius does affect is the reading of a distance. A statement of the form the bowl is thirty times longer one way than another describes a level set, and a level set at a five per cent rise sits well outside 0.023 in the flat directions. That statement was checked against the true objective when it was made; a reader tempted to extrapolate it further should know where the picture stops.

One more property of the ordering is worth naming, because it is what makes the result a result rather than an artefact. The radii do not track the condition number of the Hessian, which is the quantity a reader would nominate: XYZ scaling and the optimum have condition numbers within a factor of two of each other and radii a factor of five apart. Nor do they track the objective’s value, or the distance to the optimum. They track the gradient, and the gradient is the term the model gets exactly right.

The devices, which behave differently

The same measurement on the two devices this collection designs gives a different pattern, and the difference is informative.

A display’s three primaries and a camera’s three dyes are each six numbers, and each is optimised, so both are at minima. But their objectives are much better approximated by quadratics over the distances that matter — a primary’s tolerance region is a level set at a one per cent rise, and the model reproduces the true objective there closely enough that the region computed from the Hessian and the region found by bisection agree to a few per cent.

The reason is that a device has no invariance. The nine-coefficient objective has three exactly flat directions, so its level sets are unbounded cylinders and everything about distance in it is awkward; a display with a different green is a different display, and its objective is an ordinary bowl in all six numbers. A trust radius on a surface with a null space is a harder object than a trust radius on one without.

An interpretation worth resisting

There is a tempting misreading of all this and it should be blocked explicitly: a short trust radius at the optimum does not mean the optimum is a bad place to be, or that its location is uncertain.

The location of a minimum is fixed by the gradient vanishing, which is a first-order condition and has nothing to do with how far a second-order model reaches. A search finds the optimum by walking downhill on the true objective, not by trusting a quadratic; the optimum in this collection was found by a simplex, which has no derivatives in it at all.

What the radius bounds is the description, not the answer. A reader looking at a picture of eigenvalues at the optimum is looking at an exact statement about an infinitesimal neighbourhood and an increasingly approximate one further out, and the number in this essay says where “further out” begins to mean something. It is a caption on a figure rather than a caveat on a result.

Where the model stops

Ten per cent is a choice, and the radius depends on it — roughly as the square root, since the leading neglected term is cubic. A one per cent criterion gives radii about a third as large and the same ordering.

Eighteen directions is not every direction. The eigenvectors are the natural ones and their negatives catch the asymmetry a cubic term introduces, but a direction between two eigenvectors could in principle be worse than either. Spot checks at random directions inside the reported radii find nothing worse; that is a check rather than a proof.

At a published matrix the surface curves downwards in some directions. Nine markers per series, on a logarithmic axis of magnitude: the objective's curvature along each of its own principal directions, at Bradford and at the optimum. Filled markers are directions that curve upwards and open ones curve downwards. At the optimum every one is upwards or zero, which is what being a minimum means. At Bradford 3 of the nine curve downwards, so the point is a saddle — there are directions in which the objective falls away quadratically as well as linearly, and describing that neighbourhood as a bowl is wrong in a way an eigenvalue table does not show, because a table of magnitudes has lost the sign.
Fig. 4 The signed curvature along each principal direction at a published matrix and at the optimum. At the optimum every direction curves upwards; elsewhere several curve down, and a surface with directions of both kinds is harder for a quadratic to describe.

And the measurement is per point rather than per region. A trust radius says how far a model built here reaches; it does not say that a model built halfway between two points would reach further, which it might. Nothing here builds a model at a point nobody has a reason to be at.

What it changes about reading this collection

Four published results here are drawn from a curvature at an optimum, and it is worth saying which of them the radius touches.

The rank of the Hessian is six — untouched, because a rank is a property of a matrix rather than of a neighbourhood. The three directions a gain cannot see — untouched, because those are exactly flat lines and exactness has no radius.

What a constraint costs is a prediction over a finite displacement and was checked against the true objective at that displacement, which is the right check; the radius here is a reminder of why the check was necessary rather than optional. And what one per cent of one objective buys of the other is computed at a one per cent rise, which is well inside 0.023 in the stiff directions and outside it in the flat ones — so the flat-direction half of that result rests on an extrapolation, and it too was measured directly rather than predicted.

Three of the four were already checked against the truth at the distances involved. That is not because anybody had this number; it is the site’s habit of asserting what a figure claims, doing the work a trust radius would have prompted.

Exactly flat everywhere, and eigenvectors at one point only. Two columns over the same nine places. On the left, how far from zero the objective's second derivative is along a row-scaling direction, on a logarithmic axis — it is between 10⁻⁹ and 10⁻⁶ of the largest eigenvalue at every one of them, which is a numerical zero. Scaling a row of the basis is a straight line along which the cost does not change, and that is true at every point, not only at the optimum. On the right, the angle between those three directions and the Hessian's own three smallest eigenvectors: 0.025 degrees at the optimum and up to 88 away from it. An invariance is a property of the function; being an eigenvector is a property of the function at a minimum, and the two coincide only where everybody computes.
Fig. 5 The quadratic form on a row scaling, and the angle between those directions and the Hessian’s own null space. The first column is exact everywhere and has no radius; the second is a property of a point.

Two more views say what happens outside the radius the quadratic holds in, which is where a reader would actually want to use it.

Downhill from every published matrix, one step at a time. Each curve is a steepest-descent walk from one of this collection's published bases, plotted as the objective against the distance walked in the nine coefficients. The horizontal line is the optimum. The first step of each walk is the long one — XYZ scaling closes 41 per cent of its whole gap in one — and every walk then flattens without reaching the line, because the valley floor is nearly flat and the steepest direction is nearly across it. Bradford starts closest and closes least: it is already in the flat part.
Fig. 6 Steepest-descent walks from each published basis. A walk is the quadratic being trusted step by step, and where it stops improving is where the model it was following has run out.
Steepest descent sets off at right angles to the answer. One pair of bars per basis: the angle between the direction the objective falls fastest in and the straight line to the optimum, and the same angle with the three directions the objective cannot see projected out of the displacement. The obvious suspicion is that a near-ninety-degree angle is bookkeeping — a gradient is exactly orthogonal to those three by construction, and the optimum is a three-parameter family whose representative is arbitrary. The projection moves the angle by at most 4.9 degrees, so it is not: the surface really is a valley whose floor runs one way and whose walls are three orders of magnitude steeper, and the steepest way down is across it.
Fig. 7 And the angle between the steepest direction and the direction of the optimum at each basis. A quadratic that is believed too far sends a walk in a direction that is only locally downhill.

The generalisation

The result is worth stating in a form that outlives the objective it was measured on.

A second-order model’s accuracy is measured on the change it predicts, and at a minimum the change it predicts is entirely second order. So the relative accuracy of a quadratic is systematically worst at a stationary point, for every function whose third derivative does not also vanish there — which is nearly all of them.

The practical consequence is a small inversion of habit. A curvature is normally computed at an optimum because that is where it is meaningful — the gradient vanishes, the eigenvalues are the shape, and the interpretation is clean. This says the same point is where the resulting picture describes the least territory, so the cleanliness of the interpretation and the reach of the description pull in opposite directions.

The check is cheap and is almost never run. Two hundred and fifty evaluations of an objective is nothing next to the cost of finding the optimum in the first place, and the answer changes how far a reader should extrapolate every statement made from the Hessian afterwards.

Who found it, and when

Trust regions are an entire subfield of optimisation, and the idea of bounding the step by how far the quadratic model can be believed goes back to Levenberg and Marquardt in the 1940s and 1960s and to Powell’s work in the 1970s. In that literature the radius is adaptive: it is grown when the model predicts well and shrunk when it does not, and it is a control parameter rather than a reported quantity.

What is unusual here is reporting it. An optimiser adjusts its trust radius silently and throws it away; nothing in the usual practice writes down how far the model at the answer can be believed, because the optimiser stops caring once it has arrived.

The observation that the relative accuracy of a Taylor expansion is worst where the leading term vanishes is elementary and is taught, usually as a warning about computing relative errors near zeros. It is the same fact and it is easy not to recognise in this dress.

Where the ladder goes next

The trust radius is one of two numbers that describe what a curvature at a published matrix is worth, and it is the more forgiving one. The other is where the slope catches up with the bowl — and at every published transform in this table, the answer is that it does so at a small fraction of the distance to the optimum.

That is the next rung, and it says something sharper than a radius: over almost the whole journey from a published matrix to the best one, the surface is a slope, and the eigenvalues are a description of a bowl nobody meets on the way.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyChromatic adaptationCondition numberConvergenceDeclared inputDegrees of freedomEigenvalueQuadratic form