Matching and measuring

How long is the bowl

The optimum of an adaptation objective is twenty-nine times longer one way than another, and the two ends have names — the stiffest direction is almost entirely the short-wavelength row of the basis, the flattest almost entirely the long-wavelength one. Twenty-four random directions reported a factor of eight, and no number of them would have done better.

Assumes The rank is the invariance, A constraint costs what it points at and A gain needs a basis.

A long bowl is not the same shape as a round one, and everything that turns on how much a constraint costs turns on which way it points.

The bowl the eigenvalues describe and the bowl a sample found. Six points on a logarithmic vertical axis — the distance from the optimum of the adaptation residual to a 5 per cent rise along each of the six directions the objective can see — with a shaded band behind them showing the whole range 24 random directions reported. The eigen-radii run from 1.2e-2 to 3.6e-1, a factor of 29.8. The band runs from 2.2e-2 to 1.8e-1, a factor of 8.0, and sits entirely inside the ends of the true range: a random direction in nine dimensions carries a share of every eigenvector and so reports the middle of the bowl, never an end of it.
Fig. 1 The six distances from the optimum to a five per cent rise, one per direction the objective can see, with the band twenty-four random directions reported drawn behind them.

The claim

The optimum of the adaptation objective is a bowl whose six real directions differ in reach by a factor of 29.8, its stiffest direction is 97 per cent the short-wavelength row of the basis and its flattest 85 per cent the long-wavelength row, and a sample of twenty-four random directions reports 8.0 — an understatement no amount of sampling would repair.

  • The reach is √(2Δ/λ). With eigenvalues from 676.6 to 0.760, the distance to a five per cent rise runs from 0.012 to 0.358 in the nine coefficients.
  • The directions have names, and they are rows. The three the objective cannot see are each entirely one row rescaled; among the six it can, the stiffest and the flattest are dominated by single rows at opposite ends of the spectrum.
  • A random direction reports the mean. In nine dimensions every unit vector carries an expected ninth of its length on every eigenvector, so it measures an average curvature and never an extreme.
  • And the sample added nothing. The same six eigenvalues predict each of its twenty-four radii to a median of 7.6 per cent, with nothing fitted.
  • Shrinking the rise fixes the model and not the sample. At a rise a hundred times smaller the prediction’s worst error falls from 42 per cent to 6.5, and the sampled ratio moves from 8.0 to 8.8.

What a radius means here

The objective is the mean CIEDE2000 an adapted observer is left with after a von Kries gain, across the census of illumination changes this collection models. Minimised over all nine entries of the basis it reaches 0.9741, and that floor is what every constraint is priced against.

Around that point the objective is, to second order, a quadratic form. Along an eigen-direction with eigenvalue λ the cost rises as ½λt², so the distance to a rise of Δ is √(2Δ/λ) — no search, no bisection, one square root. At a rise of five per cent of the floor, the six reaches are

0.012, 0.032, 0.065, 0.091, 0.146, 0.358.

Bisecting the true objective outward along the same six directions gives 0.018, 0.031, 0.074, 0.106, 0.147, 0.518, which is the quadratic model’s accuracy at a distance where it has no particular right to be accurate. The agreement is better in the stiff directions than in the flat ones, which is the opposite of what intuition suggests and is exactly what a fourth-order term does: the flat directions are the ones the walk goes furthest along.

Nine eigenvalues, six of which exist. Nine points on a logarithmic vertical axis: the eigenvalues of the Hessian of the adaptation residual at its own optimum, largest to smallest. The first six run from 6.8×10² down to 7.6×10⁻¹, a condition number of 890. Then the axis drops: the seventh is 1.9×10⁻⁴, and the last three are separated from the sixth by a factor of 4.0×10³. Those three are not small curvatures. They are the finite-difference truncation error on directions along which the objective is exactly constant, and a shaded band marks them as the numbers the objective does not have.
Fig. 2 The nine eigenvalues this comes from. Six of them, spanning 890, and then a fall of four thousand into the directions the model cannot see.

The directions have names

An eigenvector of a Hessian over nine matrix entries is a 3×3 with no obvious meaning. It gets one from where its weight sits.

Write each direction as three row-vectors and take the length of each. The three the objective cannot see come back as (0, 0, 1), (0, 1, 0) and (1, 0, 0) to three decimals: each of them rescales one row of the basis and touches nothing else, which is precisely what the invariance says a flat direction is. That is the algebraic claim appearing in the loadings without any angle having been computed, and it is the cleanest confirmation of it in this collection.

The six the objective can see are mixtures, and the two ends are not:

  • the stiffest, at λ = 676.6, is 97 per cent the short-wavelength row;
  • the flattest, at λ = 0.760, is 85 per cent the long-wavelength row.
Which row of the basis each direction moves. A grid with one column per eigen-direction of the curvature of the adaptation residual, stiffest on the left, and one row per row of the basis. Each bar is that row's share of that direction. The last three columns — the ones the objective cannot see — each have a single full-length bar and two empty ones: a direction that rescales one row and touches nothing else, which is exactly what the invariance says they are, visible without any angle being computed. Among the six the objective does see, the stiffest is 97 per cent the short-wave row and the flattest is 85 per cent the long-wave one, so an adapted observer's model is most particular about where its blue axis points and least about where its red one does.
Fig. 3 Each direction as a share of the three rows. The shaded columns are the three the objective cannot see, and each is a single full-length bar.

So an adapted observer’s model is most particular about where its blue axis points and least about where its red one does, by a factor of 890 in curvature and 30 in reach.

That is not arbitrary and it is not a fact about the search. The gain in each channel is the ratio of the two whites read in that channel, and across a census of real illuminants the short-wavelength channel’s ratio is the one that swings — a tungsten lamp against daylight moves the blue reading by a large factor and the red one by a small one. A model whose blue axis is in the wrong place gets the largest gain wrong. A model whose red axis is in the wrong place gets the smallest one slightly wrong.

It also explains something that has sat in this collection unexplained: the published transforms differ from each other and from the receptors most in their third row, and the fits that produced them were most tightly constrained exactly there.

Why twenty-four directions found a factor of eight

The previous measurement of this bowl walked outward along twenty-four random unit directions until the cost had risen by five per cent, and reported the longest radius over the shortest: 8.03.

A random unit vector in nine dimensions has, in expectation, one ninth of its squared length on every eigenvector. The curvature it feels is Σ λₖ uₖ², which concentrates near the mean of the nine eigenvalues — here about 90 — with a spread that is narrow because it is a sum of nine terms. So the radius it reports concentrates near √(2Δ/λ̄), and neither end of the true range is anywhere near it.

More samples do not help. To find the flattest direction a sample would have to land near it, and the fraction of the sphere within ten per cent of a given axis in nine dimensions is of order 10⁻⁸. Twenty-four is not a small sample for this purpose; twenty-four million would be.

Six numbers predicting twenty-four measurements. A scatter of 24 points on logarithmic axes, one per random direction walked out from the optimum of the adaptation residual. The horizontal position is the radius measured by bisecting the true objective; the vertical is the radius the six eigenvalues predict, with nothing fitted. The points lie on the diagonal to a median of 7.6 per cent. The sample measured nothing the curvature did not already contain — which is the point, because the sample nevertheless reports a bowl 8.0 times longer one way than another where the eigenvalues say 29.8.
Fig. 4 Every one of the twenty-four measured radii against the radius the eigenvalues predict for that same direction. The sample measured nothing the curvature did not already hold.

The strongest form of that statement is the picture above. The six eigenvalues predict each individual sampled radius — not the distribution, the individual measurement — to a median of 7.6 per cent at a five per cent rise. The twenty-four bisections were an expensive way of confirming a quadratic form that was already computable in two seconds.

The control that rules out the easy objection

The obvious objection is that the bowl is simply not a bowl five per cent above its floor: at radii of a third of a unit the quadratic model is being stretched, and perhaps a sample walking the true objective is right and the model is wrong.

Shrinking the rise settles it, and it settles it in two directions at once.

At a rise of 0.05 per cent rather than five, the quadratic model’s worst per-direction error falls from 42 per cent to 6.5 — the model is converging on the objective, as a second derivative must. And the sampled ratio moves from 8.03 to 8.76, against a true 29.8. The same thing happens on the other objective: the model’s worst error falls from 58 per cent to 6.3, and the sampled ratio moves from 5.92 to 5.21 against a true 27.2.

The model gets better close in and the sample does not. What the sample is missing is the dimension of the space, which does not care how far out the measurement is made.

The other objective has the same shape and different numbers

The same arithmetic applied to the second thing this collection scores a basis on — how nearly a lightness–chroma space built on it makes the discrimination contours circles — gives a bowl of the same structure. Six real eigenvalues, 188.5, 39.63, 8.381, 4.209, 1.589, 0.254, a condition number of 742, and a reach ratio of 27.2 against the adaptation objective’s 29.8.

Two long bowls in the same nine-dimensional space raise an obvious question: are they long in the same directions? If they were, one basis would serve both and the trade-off this collection keeps finding would not exist. If one’s long directions were the other’s short ones, the two could be optimised independently and there would be no trade-off either.

Neither. The principal angles between the two six-dimensional subspaces each objective can see are 52.6°, 41.5°, 8.1°, and then three zeros. The three zeros are forced — two six-dimensional subspaces of a nine-dimensional space must share at least three directions — and their coming out at zero to a rounding error is the check that the other three are computed correctly. The stiffest directions of the two are 80.9° apart.

Where the two objectives agree and where they do not. Six horizontal bars: the principal angles between the six-dimensional subspaces the two objectives can see. Three of them are zero to within a rounding error — which is forced, because two six-dimensional subspaces of a nine-dimensional space must share at least three directions, and their being exactly zero is the check that the angles are being computed correctly. The other three are 52.6°, 41.5°, 8.1°. The two objectives are neither the same question asked twice nor two independent questions; they overlap in half of what they can see.
Fig. 5 The six principal angles between what the two objectives can see. Three zeros that had to be there, and three that did not.

So the two bowls overlap in half of what they can see and are oblique in the other half, which is exactly the geometry a genuine trade-off has: some directions buy one objective at the other’s expense, and some buy nothing for either.

What was computed, and how

The Hessian is 154 central differences at a step of 2 × 10⁻⁴, taken inside a window measured rather than assumed: over a decade of steps the six real eigenvalues move by under half a per cent while the three spurious ones grow as the square of the step. Its eigendecomposition is a one-sided Jacobi singular value decomposition of the matrix itself, which computes small singular values to high relative accuracy — the property that matters when three of the nine ought to be zero.

The measured radii are bisections of the true objective along each eigen-direction, thirty-four halvings, with a doubling search first to bracket. The predicted ones are √(2Δ/λ). Neither is fitted to the other.

The assertion in the build has three parts and the third is the one that makes it a result rather than an observation: the sampled ratio must understate the true one by a stated factor; the quadratic model must predict every sampled radius; and its error must fall as the rise falls while the understatement does not. A version with only the first part would be satisfied by any estimator having a bad day.

What a constraint costs is how far it pushes, in the directions that are seen. A scatter of every constraint imposed here on the nine free numbers. The horizontal axis is the length of the displacement from the optimum measured only in the six directions the objective can see; the vertical, on a logarithmic scale, is the excess cost that displacement actually carries. Requiring the basis to be the inverse of three realisable display primaries sits at the bottom left, at 0.068 and 0.022 ΔE00 — it removes three degrees of freedom and moves the answer almost nowhere. Requiring it to hit the three dichromat confusion points removes six and pushes 13 times as far, for 0.68. The vertical spread at similar horizontal positions is the part a count of parameters cannot predict.
Fig. 6 What the reach is for. Every constraint this collection imposes on the nine numbers, plotted against how far it pushes in the six directions the objective can see.

The sampled eight is a typical draw, simulated

The argument that twenty-four random directions must report the middle is convincing and it is also checkable, because the six eigenvalues are enough to simulate the whole experiment.

Drawing twenty-four random unit directions in the six-dimensional space the objective can see, computing each one’s effective curvature as the sum of the eigenvalues weighted by its own squared components, and taking the longest radius over the shortest, gives a median ratio of 5.9, with ninety per cent of runs between 4.0 and 9.5.

The measured 8.03 sits comfortably inside that band, a little above its median. So the previous measurement was not unlucky and was not badly implemented — it returned exactly what twenty-four random directions return, and it would have returned something between four and ten had it been repeated. Against a true 29.8, the sampling procedure’s best plausible outcome is a third of the answer.

That is a stronger statement than the essay’s argument alone supports. The argument says the sample concentrates near the mean and cannot find the ends; the simulation says the whole distribution of possible answers stops well short, so no amount of re-running would have raised an alarm. A method whose spread does not contain the truth is worse than a noisy one, because its own repeatability reads as evidence.

The chance of finding the flat end

The fraction of the sphere within ten per cent of a given axis in nine dimensions is of order 10⁻⁸ is the argument’s supporting number, and it is a good deal larger than that — which makes the point differently rather than less.

Counted directly: of two million random directions in the six visible dimensions, three reported a radius within ten per cent of the flattest, a fraction of about 1.5 × 10⁻⁶. So finding the flat end needs about seven hundred thousand directions rather than a hundred million, and the essay’s twenty-four million would find about thirty-six of them.

The conclusion survives intact and gets a sharper comparison. The eigenproblem costs 154 evaluations of the objective. Finding the same answer by sampling costs of order a million. The twenty-four bisections that were actually run cost about eight hundred evaluations on their own — five times the Hessian — and bought a third of the answer, which is the whole case in one line.

It is also worth noting where the two orders of magnitude went. Ten per cent of the radius is twenty per cent of the eigenvalue, because the radius goes as one over the square root, so the target is twice as wide as a naive reading suggests; and the six-dimensional sphere is far more generous than a nine-dimensional one, at 5.8 × 10⁻³ against 3.9 × 10⁻⁴ for the same angular window. The sample lives in six dimensions rather than nine, because the three directions the objective cannot see would have returned unbounded radii and cannot have been in it.

What the mean of the eigenvalues predicts

The concentration argument makes a prediction the essay does not cash: a random direction should report a radius near √(2Δ/λ̄).

Over nine eigenvalues, six real and three zero, the mean is 90.2 and the predicted radius is 0.0329. Over the six real ones alone the mean is 135.3 and the radius is 0.0268. The true extremes are 0.0120 and 0.3580.

So the typical sampled radius sits at about a twelfth of the way from the stiff end to the flat one on a linear scale — and near the geometric middle, which is where a concentration argument puts it: √(0.0120 × 0.3580) = 0.0655, and the sampled median is nearer the stiff end than that because the large eigenvalues dominate the sum. A random direction is pulled towards the stiffest axis, not towards the middle, which is why the understatement is one-sided and why the sampled ratio understates rather than scattering about the truth.

What a long bowl is worth

A bowl thirty times longer one way than another is the reason two constraints of the same nominal size can cost nothing and seventy per cent. That was established without the shape and can now be said with it: requiring the basis to be the inverse of three realisable display primaries moves it 0.068 in the directions the objective sees, and requiring it to hit the three dichromat confusion points moves it 0.891 — thirteen times further, into directions of comparable stiffness.

There is a second use, and it is the more interesting one. A flat direction is a direction along which the objective has no opinion, so it is free to be spent on something else. What it buys, when it is spent on the other objective this collection scores a basis against, is a great deal more than the price — and the direction to spend it in is the flattest one rather than the one the second objective most wants.

Nine eigenvalues, six of which exist. Nine points on a logarithmic vertical axis: the eigenvalues of the Hessian of the ellipse anisotropy at its own optimum, largest to smallest. The first six run from 1.9×10² down to 2.5×10⁻¹, a condition number of 742. Then the axis drops: the seventh is 2.1×10⁻⁵, and the last three are separated from the sixth by a factor of 1.2×10⁴. Those three are not small curvatures. They are the finite-difference truncation error on directions along which the objective is exactly constant, and a shaded band marks them as the numbers the objective does not have.
Fig. 7 The second objective’s own nine eigenvalues, with the same six-and-a-cliff structure and a gap of twelve thousand.

Where the model stops

A Hessian describes a neighbourhood and the published transforms are not in it. At full displacement the quadratic prediction of a basis’s excess cost is out by anything from five per cent to a factor of three. Every number here is a statement about the immediate surroundings of one point, and extending it outward is done by walking rather than by asserting.

The census decides the six eigenvalues. The rank does not — that is a property of the model — but the numbers are properties of fourteen changes of illumination somebody chose, weighted equally. A census with more discharge lamps in it would move the eigenvalues and would not move the rank.

And the row names are labels rather than mechanisms. Saying the stiffest direction is the short-wavelength row is a statement about where the weight of an eigenvector sits, in a particular normalisation, at a particular point. It is not a claim that the S cone is doing anything in particular; the eye is not applying a matrix to anything, and a von Kries diagonal is a model rather than a mechanism.

Six eigenvalues that stay and three that fall as the step squared. Nine curves on logarithmic axes: each eigenvalue of the Hessian of the adaptation residual against the finite-difference step it was measured at, from 2×10⁻⁴ to 2×10⁻³. Six of them are nearly horizontal: the largest two move by 1.15 per cent over the whole decade and the worst of the six by 26, which is the objective departing from a quadratic at the larger steps rather than anything about the method. The other three grow by a factor of 280 as the step grows by 10, which is the square law a truncation error obeys on a quantity that is exactly zero. The gap between the sixth curve and the seventh is 4.0×10³ at the smallest step.
Fig. 8 The step the whole measurement rests on, and the evidence that it is inside the window: six flat lines and three of slope two.

The generalisation

A level set is an ellipsoid and an ellipsoid has axes, so a search for its extent should be an eigenproblem and not a walk. Where the objective is smooth and a second derivative is affordable — and 154 evaluations is nearly always affordable — the walk is strictly worse: slower, less accurate, and unable to find either end.

The second half is about what a random direction is for. Random directions are an excellent estimator of a typical value and a terrible one of an extreme, and the dimension decides how terrible. In two dimensions twenty-four directions would find both ends of an ellipse comfortably. In nine they find the middle, and in ninety they would find the middle more precisely.

That is worth carrying because random-direction probing is the standard tool for exactly this job in high-dimensional settings — sensitivity analysis, robustness checks, adversarial perturbation — and in each of those the quantity of interest is usually the worst direction. It is the same defect as measuring an ellipse by walking round it, one dimension count up, and the same repair applies: find the shape, do not sample it.

The three directions the curvature cannot see are the three row scalings. Three horizontal bars, one per principal angle between the null space of the Hessian of the adaptation residual and the three directions along which a row of the basis is rescaled. All three angles are under a tenth of a degree — 2.5e-2°, 7.9e-4°, 2.5e-4° — so the two three-dimensional subspaces are the same subspace. This is the algebraic invariance that a von Kries gain cannot see the scale of a row, measured as a geometry rather than assumed.
Fig. 9 The three principal angles between the Hessian’s null space and the row scalings, which is how the six were separated from the nine in the first place.

Who found it, and when

Rayleigh’s principle dates the arithmetic to 1877: the extreme values of a quadratic form on the unit sphere are its extreme eigenvalues, attained at the corresponding eigenvectors. Nothing here is new mathematics and none of it is hard.

What is not standard, in colour appearance modelling, is asking the question at all. The published adaptation transforms are reported as matrices with an error figure, and the error figure is a single number: how well the fit did. The shape of the objective around the fit — which combinations of the entries were determined and which were floating — is not reported for Bradford, CAT02 or CAT16, and cannot be recovered from the published matrices, because a matrix is a point and a point has no curvature.

The habit belongs to the identifiability literature and to experimental design, where the Fisher information matrix is exactly this object and its eigenstructure is the standard diagnostic. Colour has borrowed the fitting and not the diagnostics.

Where the ladder goes next

Two facts sit oddly together and are the next thing to reconcile. A constraint costs what the curvature in its direction says it costs — and yet the quadratic model, evaluated at the published transforms, is wrong by up to a factor of three. Both are true, and the way to show it is to walk in rather than argue at the edge.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

BasisChromatic adaptationCondition numberCone fundamentalsDegrees of freedomEigenvalueHessianLevel setSamplingThe von Kries transform