Where the model breaks

A constraint is a direction and a distance

Four restrictions on the same nine numbers cost nothing, nothing, two per cent and seventy. How many parameters each removes predicts none of it. What does is the quadratic form evaluated along the displacement — and showing that it does means walking in towards the optimum rather than arguing at the edge, because at the edge the prediction is out by a factor of three.

Assumes A constraint costs what it points at, How long is the bowl and The rank is the invariance.

Counting the parameters a restriction removes is the natural way to price it, and it is worth almost nothing.

What a constraint costs is how far it pushes, in the directions that are seen. A scatter of every constraint imposed here on the nine free numbers. The horizontal axis is the length of the displacement from the optimum measured only in the six directions the objective can see; the vertical, on a logarithmic scale, is the excess cost that displacement actually carries. Requiring the basis to be the inverse of three realisable display primaries sits at the bottom left, at 0.068 and 0.022 ΔE00 — it removes three degrees of freedom and moves the answer almost nowhere. Requiring it to hit the three dichromat confusion points removes six and pushes 13 times as far, for 0.68. The vertical spread at similar horizontal positions is the part a count of parameters cannot predict.
Fig. 1 Every constraint this collection imposes on the observer’s nine free numbers, plotted against how far it pushes the basis in the six directions the objective can actually see.

The claim

The excess cost a constraint carries is ½ Σ λₖ dₖ² in the eigen-coordinates of the displacement it forces. Two directions out of nine carry most of it in every case, the number of parameters removed carries none of it, and demonstrating this requires walking towards the optimum rather than evaluating at the constraint’s own position.

  • Removing three parameters can cost 2 per cent and removing six can cost 70. Requiring the basis to be the inverse of three realisable display primaries takes three degrees of freedom and leaves 0.996 ΔE00 against a free floor of 0.974. Requiring it to hit three dichromat confusion points takes six and leaves 1.651.
  • The difference is the distance, not the count. In the six directions the objective can see, the display constraint moves the basis 0.068 and the confusion points move it 0.891.
  • Two of nine directions decide it. Across every constraint here, the largest two terms carry between 62 and 95 per cent of the predicted excess.
  • At full displacement the quadratic model is not to be trusted. It is out by 5 per cent on one entry and a factor of three on another.
  • So the demonstration is a walk. Scale the displacement by t and the prediction must approach the truth and do it with an error proportional to t, which is the next term of the series. Both hold: the worst entry is within 3.7 per cent at t = 1/128, and the error halves whenever t does.

Why the count is worthless

The tempting arithmetic is that nine numbers minus six constraints leaves three, that three free parameters cannot get as close to an optimum as six can, and that the cost should scale with how many were taken away.

Every step of that is wrong here, and the first step is wrong in a way worth stating on its own: three of the nine numbers do nothing to an adaptation model at all, so a constraint that removes three parameters might remove three that were already invisible and cost exactly nothing.

But the failure is not only about the null space. Take the two constraints that both remove three visible parameters and are both satisfied by real objects, and they still differ by a factor of twenty in cost. The reason is geometric: the objective’s level sets are ellipsoids — and an ellipsoid’s extent is an eigenproblem rather than a walkthirty times longer one way than another, and a constraint is a surface. What it costs is how far it pushes the answer off the optimum and how stiff the objective is along that push.

The bowl the eigenvalues describe and the bowl a sample found. Six points on a logarithmic vertical axis — the distance from the optimum of the adaptation residual to a 5 per cent rise along each of the six directions the objective can see — with a shaded band behind them showing the whole range 24 random directions reported. The eigen-radii run from 1.2e-2 to 3.6e-1, a factor of 29.8. The band runs from 2.2e-2 to 1.8e-1, a factor of 8.0, and sits entirely inside the ends of the true range: a random direction in nine dimensions carries a share of every eigenvector and so reports the middle of the bowl, never an end of it.
Fig. 2 The six reaches the constraint costs are measured against: a factor of thirty between the stiffest and the flattest direction the objective can see.

The arithmetic, and what it needs

Let B* be the optimum and B the best basis satisfying some constraint. Write d = B − B* as a vector of nine entries and decompose it on the Hessian’s eigenvectors: dₖ = d · vₖ. Then to second order the excess is

ΔE = ½ Σₖ λₖ dₖ².

Three of the λₖ are zero, so three of the nine components contribute nothing however large they are. That has a practical consequence worth naming: the ordinary Euclidean distance between two bases is not a meaningful quantity. Two matrices can differ substantially in entries and identically in behaviour. What is meaningful is the length of the displacement restricted to the six directions the objective can see, and that is the horizontal axis of every picture here.

Two more quantities fall out of the decomposition and are worth reporting beside the total. The effective curvature λ_eff = Σλₖdₖ² / Σdₖ² is the stiffness the constraint actually meets, averaged along its own direction. The concentration is what share of the predicted excess the largest two terms carry.

constraint seen distance excess, ΔE00 top two share
the inverse of three realisable primaries 0.068 0.022 69%
Bradford 0.539 0.166 62%
CAT16 0.824 0.338 77%
CAT02 0.673 0.349 64%
Hunt–Pointer–Estévez 0.856 0.610 95%
the three confusion points 0.891 0.677 74%
XYZ scaling 0.553 1.399 80%

The last row is the one that shows the count is not the story from the other side. XYZ scaling is not a constraint at all in the usual sense — it is the identity, a choice rather than a restriction — and it sits closer to the optimum in the seen directions than four of the entries above it while costing more than any of them. It points the wrong way.

The honest way to show it

Here is where the argument could have gone wrong, and nearly did.

Evaluating ½ dᵀHd at each of these bases and comparing against the true excess gives agreement ranging from very good to very bad. Hunt–Pointer–Estévez predicts 0.651 against a measured 0.610 — seven per cent, at a distance where a quadratic model has no right to anything. CAT16 predicts 1.098 against a measured 0.338, out by a factor of three.

Quoting the good rows would be selection. Quoting the average would be worse. Quoting the bad rows as evidence against the model would be wrong, because a Hessian is a statement about a limit and every one of these bases is a long way outside it.

So the displacement is scaled instead. Walk a fraction t of the way from the optimum towards each basis and compare the prediction, which scales as , against the measured excess. Two things must happen if the Hessian is what it claims to be: the ratio must go to one, and it must approach one with an error proportional to t, because the remainder of a Taylor series is the next term.

Both do. For the receptor construction the ratio runs

1.389, 1.248, 1.136, 1.071, 1.037 at t = 1/8, 1/16, 1/32, 1/64, 1/128,

which is an error of 0.389 falling to 0.037 — halving each time the displacement halves, three times over. Every one of the six bases behaves the same way; the worst at t = 1/128 is 3.7 per cent from one.

The error a curvature leaves halves when the step does. Three curves on logarithmic axes. Each is the error in predicting a basis's excess cost from the quadratic form alone, plotted against the fraction of the way from the optimum to that basis. All three fall along a line of slope one: halving the displacement halves the error, which is what the third term of a Taylor series does and is the evidence that the matrix being used is a second derivative rather than a fit. At an eighth of the way the prediction is out by up to 39 per cent; at a hundred and twenty-eighth it is out by 3.7 per cent at worst.
Fig. 3 The error the quadratic prediction leaves, against the fraction of the way to each basis, on logarithmic axes. Three curves of slope one.

The second condition is the one that makes this a demonstration rather than a coincidence. A prediction that merely converged would be satisfied by any function agreeing at small displacements. A prediction whose error is first order in the step is a second derivative, and nothing else is.

Why two directions out of nine

Concentration between 62 and 95 per cent is high enough to be worth explaining, and the explanation is the same geometry.

The excess is a sum of nine terms ½λₖdₖ², six of which can be nonzero. The λₖ span a factor of 890. So unless a displacement happens to be almost exactly orthogonal to the stiff directions, the stiffest one or two components dominate — a component of half the size on a direction ten times stiffer contributes two and a half times as much.

The exception is instructive. The display constraint’s top two are directions three and two rather than one and two, because its displacement is nearly orthogonal to the stiffest direction: the primaries it uses were chosen by minimising the same objective, so the search pushed the residual out of the direction it cost most in. A constraint optimised against the objective ends up orthogonal to the objective’s stiffest direction, which is a compact way of saying what optimisation under a constraint does.

What was computed, and how

The Hessian is 154 central differences at a step measured to be inside the window where the six real eigenvalues do not depend on it. The eigendecomposition is a one-sided Jacobi singular value decomposition, which resolves the three that ought to be zero rather than losing them in a squared condition number.

Each named basis is normalised on the adopting white before the displacement is taken, so the comparison is between directions rather than between directions and arbitrary units. The designed display’s basis is not one of the eight named matrices — it is the inverse of the primaries a search returns when asked to minimise the same objective over three realisable chromaticities — so it enters the table as a matrix rather than by name.

The assertion in the build has four parts: the walk must converge to within six per cent at the smallest step; the error must halve over the last two intervals; the top two directions must carry at least 55 per cent of every predicted excess; and the cheap constraint’s seen distance must be at least five times smaller than every other entry’s. The third and fourth are what turn a convergence check into a statement about constraints.

Only the last two intervals are used for the halving test, and that restriction is honest rather than convenient: at an eighth of the way to CAT16 the fourth-order term is still visible and the error falls by 0.68 rather than by a half. Requiring the whole walk to halve would be requiring the series to have no terms beyond the third.

Where the two objectives agree and where they do not. Six horizontal bars: the principal angles between the six-dimensional subspaces the two objectives can see. Three of them are zero to within a rounding error — which is forced, because two six-dimensional subspaces of a nine-dimensional space must share at least three directions, and their being exactly zero is the check that the angles are being computed correctly. The other three are 52.6°, 41.5°, 8.1°. The two objectives are neither the same question asked twice nor two independent questions; they overlap in half of what they can see.
Fig. 4 The two objectives’ seen subspaces, and how much of each other they contain. A constraint that is cheap for one is not therefore cheap for the other.
The cheapest direction to give ground in is the flattest one. Six bars, one per direction the adaptation objective can see, showing how much of the other objective a fixed budget of adaptation buys if it is spent along that direction. The rate is the slope of the second objective divided by the square root of the first's curvature, so it rewards a direction the second objective wants and punishes one the first is stiff in. The flattest direction wins at 10.68 against 3.29 for the next best and 0.54 for the stiffest — a factor of 20. Spending 1 per cent of the adaptation optimum there moves the anisotropy from 7.70 to 5.02.
Fig. 5 And the same nine numbers priced as an exchange rate rather than as a cost: how much of the second objective a fixed budget of the first buys, direction by direction. A constraint is cheap where the bowl is flat, and the flattest direction is not the one anybody names.

What this settles about the confusion points

The receptor construction costs seventy per cent above the unconstrained floor, and the decomposition says where that goes: 74 per cent of it in two directions, the third and the sixth, at a seen distance of 0.891.

The sixth direction is the flattest the objective has. A displacement of that size along the flattest direction would cost almost nothing; the construction pays because it also moves 0.4 along the third, which is thirty times stiffer. So the price is not that the confusion points are far away — Bradford is 0.539 away and costs a quarter as much — it is that they are far away in a direction the model minds.

That is the sentence the earlier measurement was reaching for and could not state, because a bowl measured in random directions has no directions in it to name.

The same question, asked of a device

The arithmetic is not about the observer’s nine numbers in particular, and pointing it at something manufactured makes that plain.

A display’s design is six numbers — three primary chromaticities — and there is no invariance among them: a display with a different green is a different display, so the curvature has rank six out of six and nothing is invisible. What the eigenvalues say instead is a tolerance. They span a factor of 3,792, so the design can move sixty-two times further in its cheapest combination than in its dearest for the same cost in adaptation, and the cheapest combination moves the red and green primaries together and leaves the blue one where it is.

A camera’s three dyes are the same shape of object: three centre wavelengths and three bandwidths, all in nanometres, so the six eigenvalues are directly comparable without any weighting having to be invented. There the condition number is 623, the stiffest direction is the blue dye’s centre at a weight of 0.96, and the flattest is the three bandwidths together.

Both are constraints in the sense of this essay — a manufacturer restricted to a tolerance is choosing a surface — and both show the same thing: which numbers a design can be wrong about is a fact about the objective’s shape rather than about how many numbers there are.

The table’s third column, which it does not print

The essay names the effective curvature as a quantity worth reporting beside the total and then reports the total. Dividing twice the excess by the squared distance recovers it for every row:

constraint seen distance excess effective curvature
the inverse of three realisable primaries 0.068 0.022 9.52
XYZ scaling 0.553 1.399 9.15
the three confusion points 0.891 0.677 1.71
Hunt–Pointer–Estévez 0.856 0.610 1.66
CAT02 0.673 0.349 1.54
Bradford 0.539 0.166 1.14
CAT16 0.824 0.338 1.00

The seven fall into two groups. Five published transforms meet an effective stiffness between 1.00 and 1.71 — a spread of 1.7 — and the two objects that are not published transforms meet 9.5 and 9.15, a spread of four per cent.

That is the essay’s it points the wrong way made into a number, and it says more than the phrase does. The display constraint and XYZ scaling meet the same stiffness to four per cent, and their excesses differ by 63.6 against a squared-distance ratio of 66.1 — agreeing to four per cent as well. So the two are, as far as this objective can tell, the same direction at different distances: one is the identity and the other is a design optimised to be as near it as three real primaries allow.

And the five published transforms all point somewhere else, into a region seven times slacker. Whatever the fitting to corresponding-colour data was doing, it was moving the answer out of the direction the adaptation objective minds most — which is what fitting to an objective does, and these were not fitted to this one.

The concentration is lower than chance, and one row is not

Between 62 and 95 per cent is offered as high, and the eigenvalue spread it is explained by predicts higher.

For a direction drawn at random in the six visible dimensions, the two stiffest carry (676.6 + 95.1) / 811.9 = 95.0 per cent of the expected excess. So a random constraint would concentrate at 95, and the measured mean across the seven is 74.

Every row but one concentrates less than chance, and that is the finding rather than the concentration. A displacement that fell where it liked would be almost entirely the stiffest direction; these fall systematically off it. The essay explains the display constraint that way — a constraint optimised against the objective ends up orthogonal to the objective’s stiffest direction — and the table says the same is true of nearly all of them.

The exception is Hunt–Pointer–Estévez, at 95 per cent, exactly the random figure. It is also the one transform in the list that was constructed as a set of cone fundamentals rather than fitted to anything, so nothing pushed it off the stiff direction and it landed where an arbitrary displacement would. The fitted transforms have been pulled away from the objective’s stiffest direction by their own fitting; the constructed one has not, and the concentration column is where that shows.

That is a cleaner statement of what fitting buys than the excesses give. HPE sits at 0.856 and costs 0.610; CAT16 sits at 0.824 — almost the same distance — and costs 0.338. Nearly the same displacement, nearly half the cost, and the whole of the difference is which directions the displacement went into.

The walk’s rate, read off

The convergence ratios are quoted as halving each time the displacement halves, three times over, and the four ratios are 1.57, 1.82, 1.92 and 1.92.

So the halving is reached rather than assumed: the first interval is 21 per cent short of two, the second 9 per cent, and the last two are within four per cent. The rate itself converges, which is a stronger statement than the errors falling — it says the third-order term is dying out of the ratio and not merely out of the error.

That also justifies the gate’s decision to test only the last two intervals. At an eighth of the way the ratio is 1.57 and a test demanding two would fail on a series that is behaving exactly as a Taylor series should. The restriction is not leniency; it is the correct window for the claim being made, and the two intervals inside it agree with each other to a part in five hundred.

Where the model stops

Everything here is second order. The quadratic model is demonstrably the second derivative and demonstrably wrong at the distances the published transforms sit at. Nothing in this essay predicts a cost at full displacement; what it does is explain the ordering and the concentration, and give an arithmetic that is exact in the limit and checkable on the way in.

The eigenvalues are a property of the census. Fourteen changes of illumination, equally weighted, over this collection’s family of smooth reflectances. A different census moves every λ and therefore every prediction. The structure — six directions, two dominating, a count that predicts nothing — would survive.

And a constraint here is a surface, not a penalty. Each row is the best basis satisfying its own restriction exactly. A soft constraint, of the kind a real design uses, sits somewhere between the optimum and that point, and its excess is not this excess.

The generalisation

The price of a restriction is a distance times a stiffness, and the number of parameters it removes appears nowhere in that product.

That reads as obvious and is systematically ignored, because a parameter count is available before any work is done and a displacement is not. The habit worth having is the cheap version: before accepting that a constraint is expensive because it is restrictive, check how far it actually moves the answer, in the directions the objective can see. Both halves matter, and the second half is where the invariance lives.

The second lesson is about how to demonstrate a local model at all. Evaluate at the point of interest and the result is unreadable; scale the displacement and the same model becomes testable. The convergence rate is the test: first order in the step for a second derivative, second order for a third. A model that converges at the wrong rate is the wrong model even where it agrees.

Who found it, and when

The arithmetic is the second-order condition from any optimisation text, and the identifiability literature’s version — how a restriction interacts with the curvature of a likelihood — is standard in experimental design, and is the same counting this collection does with a rank, where the whole business of choosing measurements is choosing which directions of a Fisher information matrix to make stiff.

Colour appearance modelling has the constraints and not the diagnostics. Bradford, CAT02 and CAT16 were each obtained by minimising an error over corresponding-colour data; each is reported as a matrix and a residual; and the question of how far each sits from an unconstrained optimum, in which directions, is not asked anywhere this collection can find. It cannot be recovered from a published matrix either, because the matrix is a point and a point has no curvature.

The one place the discipline does reason this way is display design, where the primaries are known to trade against each other and the trade is understood as a surface. That is the cheap constraint in the table above, and it being cheap is not a coincidence: it is the one whose designers were optimising.

Where the ladder goes next

A flat direction is a direction the objective has no opinion about, which makes it free to spend on something else. What it buys, and how asymmetric the exchange turns out to be, is the next question, and the answer is not the one the shape of the trade-off suggests.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

BasisChromatic adaptationConfusion pointConstraintDegrees of freedomEigenvalueHessianIdentifiabilityLevel setTaylor series