What it takes to deliver it

Downhill from a published matrix

Walking steepest descent from each adaptation transform in use closes between a quarter and nine tenths of its distance to the best one, and most of that in the first step. The direction it sets off in is eighty to eighty-seven degrees away from the answer, and that turns out not to be an artefact of the three directions nothing can see.

Assumes The slope arrives before the bowl, Only the flat directions keep their names and The best axes are not receptors.

If the surface under a published matrix is a slope, the obvious thing to do is walk down it and see where it goes.

Downhill from every published matrix, one step at a time. Each curve is a steepest-descent walk from one of this collection's published bases, plotted as the objective against the distance walked in the nine coefficients. The horizontal line is the optimum. The first step of each walk is the long one — XYZ scaling closes 41 per cent of its whole gap in one — and every walk then flattens without reaching the line, because the valley floor is nearly flat and the steepest direction is nearly across it. Bradford starts closest and closes least: it is already in the flat part.
Fig. 1 Steepest-descent walks from each published adaptation transform, plotted as the objective against the distance walked. The horizontal line is the optimum; every walk flattens before reaching it.

The claim

A single step downhill from a published transform buys most of what there is to buy, the walk then flattens without arriving, and the direction it sets off in is nearly at right angles to the straight line to the answer — which is a fact about the surface rather than about the parameterisation.

  • The first step is the long one. It closes 41 per cent of XYZ scaling’s gap in one move, 38 of Hunt–Pointer–Estévez’s, 19 of CAT16’s, 17 of Bradford’s.
  • Fourteen steps close between a quarter and nine tenths and never reach the optimum, because the valley floor is nearly flat and steepest descent turns into a crawl along it.
  • The angle between the first direction and the straight line to the optimum is 80 to 87 degrees.
  • Projecting out the three directions the objective cannot see moves that angle by at most 4.9 degrees, so it is not the bookkeeping artefact it looks like.
  • And Bradford starts closest and closes least, at 25 per cent, because it is already in the flat part.

Why walk at all

Nothing here is an optimiser. The optimum was found by a simplex, which has no derivatives in it, and the answer has been in this collection’s tables for two rounds.

A path is a different object from an answer, and the quantities that make it worth computing are all along it: how much of the total improvement the first step buys, how far the walk goes in the nine coefficients, and how far the initial direction points from the straight line to the bottom.

Steepest descent is the right walk for the question because it is what the gradient says to do, and the gradient is the dominant term at every one of these points. Asking what following the local slope achieves is asking what the surface’s leading term is worth.

What one step is worth

The first step is taken with a backtracking line search — a step that would move the point by a tenth of its own length, halved until the cost falls — and accepted on any decrease.

It closes, as a share of the whole gap to the optimum: 41 per cent from XYZ scaling, 38 from Hunt–Pointer–Estévez, 34 from the receptor construction, 19 from CAT16, 17 from Bradford, 13 from the discrimination optimum, and 5 from CAT02.

A third to a half of a published transform’s excess is available in one step of a first-order method, which is a large number and is what a dominant linear term implies. It is also a slightly uncomfortable one, since these are the matrices in use, and a first-order step is not a sophisticated thing to have not taken.

The honest counterweight is in the next section but one: what a step buys is measured against this collection’s own objective, which none of these matrices was fitted to.

There is a second reading of the same list, and it is the one worth carrying. The order of the first-step shares is nearly the order of the gradients, which is what a dominant linear term predicts: a steeper slope buys more in one step of fixed relative length. CAT02’s five per cent is the exception, and it is an exception about the line search rather than about the surface — its first accepted step is short because the halving stops as soon as the cost falls, and at CAT02 the objective rises again quickly along its own gradient.

That is a limitation of the method rather than a property of the basis, and it is left in rather than tuned away. A line search with curvature conditions would take a longer first step there and would make the first-step share a statement about the line search.

Where the walk stops

Fourteen steps close 88 per cent of XYZ scaling’s gap, 64 of Hunt–Pointer–Estévez’s, 52 of the receptor construction’s, 48 of CAT02’s, 27 of CAT16’s and of the discrimination optimum’s, and 25 of Bradford’s.

None reaches the optimum, and the shape of every curve is the same: a steep drop, then a knee, then a long flat crawl.

A quadratic is believed least far at the one place anybody takes one. One bar per basis: the radius, in the nine coefficients, within which the second-order model predicts the objective to within ten per cent in every one of eighteen directions. The shortest bar is the objective's own optimum, at 2.3×10⁻², and the longest is XYZ scaling at 1.1×10⁻¹ — several times further. The reason is not that the model is worse at a minimum but that it has less to do there: away from one the linear term is exact and carries most of the change, so a ten per cent error in the prediction takes longer to accumulate. It does not make a Hessian at a minimum wrong; it says the picture drawn from it describes the smallest neighbourhood in the table.
Fig. 2 The radius over which a second-order model of the objective holds at each basis. The walks in this essay leave that radius within their first step, which is why a quadratic picture does not describe them.

That is the textbook behaviour of steepest descent on an ill-conditioned surface and the condition number here is the reason: the six real curvatures at the optimum span a factor of 890, so the valley is nearly nine hundred times steeper across than along. A method that always goes down the steepest way spends its steps crossing the valley and makes progress along it only as a residue.

Bradford closing least is the informative row. It starts nearest the optimum — 1.14 against 0.97 — so it is already past the knee, in the part where the surface is flat and a first-order method has almost nothing to work with.

The angle, and why it is not bookkeeping

The direction of the first step and the straight line from the starting point to the optimum are 80 to 87 degrees apart at every published basis. That is very nearly perpendicular, and the first reaction to such a number should be suspicion.

The obvious suspect is the parameterisation. Three of the nine directions are row scalings the objective cannot see; the optimum is therefore a three-parameter family and a search returns one arbitrary member of it; and a gradient is exactly orthogonal to all three by construction. A displacement with a large component in that invisible subspace would produce a large angle for no reason at all.

Steepest descent sets off at right angles to the answer. One pair of bars per basis: the angle between the direction the objective falls fastest in and the straight line to the optimum, and the same angle with the three directions the objective cannot see projected out of the displacement. The obvious suspicion is that a near-ninety-degree angle is bookkeeping — a gradient is exactly orthogonal to those three by construction, and the optimum is a three-parameter family whose representative is arbitrary. The projection moves the angle by at most 4.9 degrees, so it is not: the surface really is a valley whose floor runs one way and whose walls are three orders of magnitude steeper, and the steepest way down is across it.
Fig. 3 The angle between the steepest direction and the direction of the optimum, at each basis, before and after projecting out the three directions the objective cannot see. The projection moves it by at most five degrees.

It is not that. Projecting the invisible component out of the displacement moves the angle from 80 to 76 degrees at XYZ scaling, and by less than a degree at the other six. The distances fall by five to thirty per cent and the angles do not move.

So the near-orthogonality is a property of the surface: a long valley whose floor runs one way and whose walls are three orders of magnitude steeper, on which the steepest way down is across. It is the same fact as the condition number, seen as a geometry rather than as a ratio.

There is one more thing the flattening says, and it is about what a different method would find. A conjugate-gradient or quasi-Newton walk would use the curvature to correct the direction and would close most of the remaining gap in a handful more steps — which is exactly what the simplex that found the optimum does by other means.

So the stalling is not evidence that the optimum is hard to reach. It is evidence that the local slope alone runs out of information, which is the quantity this essay set out to measure and is the reason the walk is deliberately unsophisticated.

What was computed, and how

Each walk is fourteen steps. At each, the gradient is taken by central differences at the step this collection measured a window for — eighteen evaluations — and the step length is halved from a tenth of the point’s own norm until the objective falls.

It is deliberately not an optimiser. Nothing here uses momentum, a conjugate direction, a quasi-Newton update or a line search with curvature conditions, all of which would close the gap far faster and would answer a different question. What is wanted is what the local slope is worth, and adding memory to the method would mix in what the previous slopes were worth.

The distance to the optimum has the three invisible directions projected out, for the reason above, and both versions are reported so that the projection’s effect can be seen rather than asserted. That reporting is the whole content of the angle result: a number that survives a correction is worth more than a number that was never checked.

Where the model stops

These matrices were not fitted to this objective. Bradford was fitted to corresponding-colour data — human judgements under two lights, which this collection does not hold. CAT16 exists because CAT02’s gains go negative in practice. Hunt–Pointer–Estévez is a set of cone fundamentals. None of them is an attempt to minimise a mean residual over a census of constructed illuminant changes, so a step downhill buys 17 per cent is a statement about the distance between two criteria as much as about either matrix.

And the objective inherits everything this phase has audited. It assumes complete adaptation, where the appearance model’s own formula says 0.94; it is a mean over a census five of whose rows are constructed; and it averages over one reflectance family. A walk downhill on it is a walk on a surface with those choices baked in.

The stiff directions turn as the basis moves; the flat ones do not. One bar per basis: the largest principal angle between the two-dimensional subspace carrying most of the objective's curvature there and the same subspace at the optimum. The optimum's own bar is zero, which is the control. Everywhere else the angle is between 12 and 55 degrees, so a tolerance computed from the curvature at one basis does not transfer to another. Two dimensions rather than one because the two largest curvatures are close enough at several of these points that a single leading eigenvector is not a stable object to compare.
Fig. 4 How far each basis’s stiff subspace has turned from the optimum’s. A walk that starts at one of these points is walking on a differently oriented valley from the one the optimum’s eigenvectors describe.

Fourteen steps is a choice, made because the curves have visibly flattened by then and because each step costs eighteen evaluations of an objective that is the slowest arithmetic in this collection. A longer walk would close more and would not change the two numbers this essay is about, which are both properties of the first step.

What a first-order step actually means here

It is worth putting the improvement into units a reader can weigh, because a percentage of a gap is a slippery quantity.

Bradford leaves 1.140 ΔE00 on the census and the optimum leaves 0.974. One step downhill from Bradford reaches 1.111. That is a gain of 0.029 ΔE00, which is far below any threshold at which two colours look different and is about a fortieth of the census’s own soft-ness under a plausible repainting of its constructed walls.

So the honest summary of a step buys seventeen per cent is: seventeen per cent of a gap that was never large. The percentages in this essay are large and the quantities are small, and both statements are worth having in the same paragraph.

The bowl the eigenvalues describe and the bowl a sample found. Six points on a logarithmic vertical axis — the distance from the optimum of the adaptation residual to a 5 per cent rise along each of the six directions the objective can see — with a shaded band behind them showing the whole range 24 random directions reported. The eigen-radii run from 1.2e-2 to 3.6e-1, a factor of 29.8. The band runs from 2.2e-2 to 1.8e-1, a factor of 8.0, and sits entirely inside the ends of the true range: a random direction in nine dimensions carries a share of every eigenvector and so reports the middle of the bowl, never an end of it.
Fig. 5 The bowl the eigenvalues describe at the optimum, direction by direction. The flat directions are where a walk ends up crawling, and they are flat by a factor of nearly nine hundred.

Which is the general shape of every result in this thread. The surface is interestingly shaped, the shapes are measurable, and the distances involved are small enough that no practical recommendation follows. That is a legitimate thing for a measurement to find and it is worth saying rather than dressing up.

What the walk is a map of

Read as a set rather than one at a time, the seven walks describe the surface better than any single measurement of it does.

They all end up in the same place and none of them arrives. The seven final values run from 1.10 to 1.57 against an optimum of 0.97, and the seven paths are between 0.11 and 0.46 long — so seven starting points scattered over a region two units across converge into a band a third of a unit wide and then stop.

That band is the valley floor, and its extent is the measurement worth having: the objective is within about fifteen per cent of its minimum over a region that a first-order method reaches from anywhere in a handful of steps and cannot leave. A collection of published transforms sitting outside that band, at 1.14 to 2.37, is a collection of matrices that are all a short walk from being nearly equivalent.

Which is a more useful summary of the table than any ranking of it, and it is one that only a set of paths can give. A ranking says which matrix is best on this objective. The walks say how much of the difference between them is a property of the objective at all, and the answer is: less than a step.

The band is half a unit wide, not a third

The closing reading of the seven walks as a set is the essay’s best summary and two of its numbers do not hold.

The final values are quoted as 1.10 to 1.57, which is a band 0.47 wide — half a unit, not the stated third. And against an optimum of 0.974 those two ends sit 13 and 61 per cent above it, so the objective is within about fifteen per cent of its minimum over a region describes the bottom of the band and not the band.

The compression is real and is worth stating in its place. The seven starting values run from 1.14 to 2.37, a spread of 1.23; the seven finishing values span 0.47. Fourteen first-order steps compress the spread between the published transforms by a factor of 2.6 — which is the finding, and it is about how much the differences between them shrink rather than about how close any of them gets.

Two of the endpoints check against the shares. XYZ scaling starts at 2.37 and closes 88 per cent of a gap of 1.396, finishing at 1.142; Bradford starts at 1.140 and closes 25 per cent of 0.166, finishing at 1.099. Those are the two ends of the quoted band, and the arithmetic reproduces them — so the band’s bottom is Bradford, which closed least, and its 1.14 neighbour is XYZ scaling, which closed most and started furthest away. The worst starting point and the best one finish four hundredths of a unit apart, which is a sharper way to put the convergence than either the band’s width or its distance from the optimum.

How closely the first step tracks the gradient

The claim that the order of the first-step shares is nearly the order of the gradients is the essay’s evidence that a dominant linear term is doing the work, and it is worth a number.

Ranking the seven bases by gradient length against their first-step shares gives a rank correlation of +0.57 — positive, and a good deal weaker than nearly the order. Against the fourteen-step shares it is +0.82, which is the stronger relation.

That ordering of the two correlations is informative and runs against the essay’s account. If the first step were the one dominated by the gradient, it should track it best; instead the cumulative result over fourteen steps tracks the gradient better than the first step does. The line search is adding noise to the first step and averaging out over the walk, which is exactly what the essay says about CAT02 — its five per cent is an artefact of where the halving stopped — generalised to the whole table.

So the honest version is that the gradient predicts where a walk ends up better than what its first move buys, and the CAT02 exception is not one exception but the visible end of a spread affecting every row. A line search accepted on any decrease makes the first step a property of the method, and the essay says so about one basis while resting a claim on all seven.

The first-step and fourteen-step shares correlate with each other at +0.68, which is the same story from a third angle: the two are measuring related but not identical things, and the difference between them is the line search.

What the two claims together are worth

Both corrections point the same way and neither touches the essay’s conclusion.

The band is wider than stated and the transforms still converge into it from a spread two and a half times larger. The first step tracks the gradient less well than stated and a first-order method still recovers a third to a half of a published transform’s excess in one move. The measurements are what the essay says they are and two of the summaries are tidier than the data.

The one place it matters is the sentence the essay itself flags as its most useful: how much of the difference between them is a property of the objective at all, and the answer is: less than a step. On the corrected numbers the spread falls from 1.23 to 0.47 rather than to a third of a unit, so about two thirds of the difference between the published transforms is less than a step — which is the same claim with a fraction attached, and is stronger for having one.

The second objective has the same structure at the same places, and its curvature spectrum is the comparison that says so.

At a published matrix the surface curves downwards in some directions. Nine markers per series, on a logarithmic axis of magnitude: the objective's curvature along each of its own principal directions, at Bradford and at the optimum. Filled markers are directions that curve upwards and open ones curve downwards. At the optimum every one is upwards or zero, which is what being a minimum means. At Bradford 3 of the nine curve downwards, so the point is a saddle — there are directions in which the objective falls away quadratically as well as linearly, and describing that neighbourhood as a bowl is wrong in a way an eigenvalue table does not show, because a table of magnitudes has lost the sign.
Fig. 6 The discrimination objective’s curvature along each of its own principal directions, at Bradford and at the optimum, on a logarithmic axis of magnitude. Filled markers curve upwards and open ones curve downwards, and at a published matrix some of them do.

The generalisation

Two things carry past this objective.

A dominant linear term means the first step is worth taking and the tenth is not. Where a gradient is the leading term over most of the distance to an answer, a single first-order step recovers a large share of the gap and the method then stalls — so how much does one step buy and does the method converge are almost unrelated questions.

And an angle between a gradient and a displacement should be checked against the parameterisation before it is believed. An objective with an invariance has directions the gradient is orthogonal to by construction, and any displacement with a component there inflates the angle for free. Here the check found the angle was real; the point is that it had to be run, and running it costs one projection.

Every published adaptation transform, and one computed from daylight, on every change. What each basis leaves an adapted observer with, row by row. Darker is worse. The last column is not a published transform: it is the basis in which a change from D65 to D50 is exactly diagonal, computed in closed form from the two spectra with nothing fitted. It is far the best on the daylight rows and it is beaten on the discharge lamps, which is the trade the published transforms are sitting in — they were fitted to data containing both kinds of light and are therefore optimal for neither. Over the census as a whole the winner is Bradford at ΔE00 1.14.
Fig. 7 Every published adaptation transform and one computed from the receptors. Each of these is a starting point for a walk in this essay, and each was chosen for a reason that is not this objective.
At a published matrix the slope arrives long before the bowl. One row per basis in this collection's table. Each row is a logarithmic axis of distance in the nine coefficients, with two markers: the radius at which the objective's curvature becomes as large as its slope, and the distance from that basis to the optimum. The first is between 3.2 and 108 per cent of the second. So over almost the whole journey from a published matrix to the best one, the surface is a slope and not a bowl — and a table of eigenvalues taken there describes a neighbourhood the optimum is nowhere near. XYZ scaling is the exception, at 1.08 of the distance, because its slope is the steepest in the table.
Fig. 8 Where the curvature catches the slope at each basis, against the distance to the optimum. The walks in this essay are what a dominant linear term does when it is followed.

The second is the more useful habit. Any statement of the form the fastest direction is not the right direction is a statement about conditioning, and conditioning is a property of a coordinate system as much as of a function — which means the first question is always whether the coordinates are doing it.

Who found it, and when

Steepest descent’s behaviour on an ill-conditioned quadratic is Cauchy’s method and its zig-zag is in every optimisation textbook: the convergence rate is ((κ−1)/(κ+1))² per step for a condition number κ, which at 890 is 0.9955 — so a step buys under half a per cent of what remains, asymptotically. The steep first steps here are the pre-asymptotic regime, where the point is still crossing the valley rather than crawling along it.

What is not textbook is the use. A descent path is normally a means to an answer and is discarded; here the answer was already known and the path is the measurement. That inversion is available whenever the interesting points of a problem are the ones somebody chose rather than the ones an optimiser finds.

Where the ladder goes next

Three essays have now taken the shape of one objective seriously at points that are not its optimum. The remaining question in this thread is about a chain rather than a surface: when a colour is delivered to a reader, it passes through a profile, a gamut mapping and a room, and each has a number somebody chose.

Putting the three costs in one column turns out to answer a question none of the three separate essays could, and the stage that dominates is the one nobody controls.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

The Bradford transformCAT16Chromatic adaptationCondition numberDegrees of freedomEigenvalueOptimisationQuadratic form