A constraint is a direction and a distance
Assumes A constraint costs what it points at, How long is the bowl and The rank is the invariance.
Counting the parameters a restriction removes is the natural way to price it, and it is worth almost nothing.
The claim
The excess cost a constraint carries is ½ Σ λₖ dₖ² in the eigen-coordinates of the displacement it forces. Two directions out of nine carry most of it in every case, the number of parameters removed carries none of it, and demonstrating this requires walking towards the optimum rather than evaluating at the constraint’s own position.
- Removing three parameters can cost 2 per cent and removing six can cost 70. Requiring the basis to be the inverse of three realisable display primaries takes three degrees of freedom and leaves 0.996 ΔE00 against a free floor of 0.974. Requiring it to hit three dichromat confusion points takes six and leaves 1.651.
- The difference is the distance, not the count. In the six directions the objective can see, the display constraint moves the basis 0.068 and the confusion points move it 0.891.
- Two of nine directions decide it. Across every constraint here, the largest two terms carry between 62 and 95 per cent of the predicted excess.
- At full displacement the quadratic model is not to be trusted. It is out by 5 per cent on one entry and a factor of three on another.
- So the demonstration is a walk. Scale the displacement by
tand the prediction must approach the truth and do it with an error proportional tot, which is the next term of the series. Both hold: the worst entry is within 3.7 per cent att = 1/128, and the error halves whenevertdoes.
Why the count is worthless
The tempting arithmetic is that nine numbers minus six constraints leaves three, that three free parameters cannot get as close to an optimum as six can, and that the cost should scale with how many were taken away.
Every step of that is wrong here, and the first step is wrong in a way worth stating on its own: three of the nine numbers do nothing to an adaptation model at all, so a constraint that removes three parameters might remove three that were already invisible and cost exactly nothing.
But the failure is not only about the null space. Take the two constraints that both remove three visible parameters and are both satisfied by real objects, and they still differ by a factor of twenty in cost. The reason is geometric: the objective’s level sets are ellipsoids — and an ellipsoid’s extent is an eigenproblem rather than a walk — thirty times longer one way than another, and a constraint is a surface. What it costs is how far it pushes the answer off the optimum and how stiff the objective is along that push.
The arithmetic, and what it needs
Let B* be the optimum and B the best basis satisfying some constraint. Write d = B − B* as a vector of nine entries and decompose it on the Hessian’s eigenvectors: dₖ = d · vₖ. Then to second order the excess is
ΔE = ½ Σₖ λₖ dₖ².
Three of the λₖ are zero, so three of the nine components contribute nothing however large they are. That has a practical consequence worth naming: the ordinary Euclidean distance between two bases is not a meaningful quantity. Two matrices can differ substantially in entries and identically in behaviour. What is meaningful is the length of the displacement restricted to the six directions the objective can see, and that is the horizontal axis of every picture here.
Two more quantities fall out of the decomposition and are worth reporting beside the total. The effective curvature λ_eff = Σλₖdₖ² / Σdₖ² is the stiffness the constraint actually meets, averaged along its own direction. The concentration is what share of the predicted excess the largest two terms carry.
| constraint | seen distance | excess, ΔE00 | top two share |
|---|---|---|---|
| the inverse of three realisable primaries | 0.068 | 0.022 | 69% |
| Bradford | 0.539 | 0.166 | 62% |
| CAT16 | 0.824 | 0.338 | 77% |
| CAT02 | 0.673 | 0.349 | 64% |
| Hunt–Pointer–Estévez | 0.856 | 0.610 | 95% |
| the three confusion points | 0.891 | 0.677 | 74% |
| XYZ scaling | 0.553 | 1.399 | 80% |
The last row is the one that shows the count is not the story from the other side. XYZ scaling is not a constraint at all in the usual sense — it is the identity, a choice rather than a restriction — and it sits closer to the optimum in the seen directions than four of the entries above it while costing more than any of them. It points the wrong way.
The honest way to show it
Here is where the argument could have gone wrong, and nearly did.
Evaluating ½ dᵀHd at each of these bases and comparing against the true excess gives agreement ranging from very good to very bad. Hunt–Pointer–Estévez predicts 0.651 against a measured 0.610 — seven per cent, at a distance where a quadratic model has no right to anything. CAT16 predicts 1.098 against a measured 0.338, out by a factor of three.
Quoting the good rows would be selection. Quoting the average would be worse. Quoting the bad rows as evidence against the model would be wrong, because a Hessian is a statement about a limit and every one of these bases is a long way outside it.
So the displacement is scaled instead. Walk a fraction t of the way from the optimum towards each basis and compare the prediction, which scales as t², against the measured excess. Two things must happen if the Hessian is what it claims to be: the ratio must go to one, and it must approach one with an error proportional to t, because the remainder of a Taylor series is the next term.
Both do. For the receptor construction the ratio runs
1.389, 1.248, 1.136, 1.071, 1.037 at t = 1/8, 1/16, 1/32, 1/64, 1/128,
which is an error of 0.389 falling to 0.037 — halving each time the displacement halves, three times over. Every one of the six bases behaves the same way; the worst at t = 1/128 is 3.7 per cent from one.
The second condition is the one that makes this a demonstration rather than a coincidence. A prediction that merely converged would be satisfied by any function agreeing at small displacements. A prediction whose error is first order in the step is a second derivative, and nothing else is.
Why two directions out of nine
Concentration between 62 and 95 per cent is high enough to be worth explaining, and the explanation is the same geometry.
The excess is a sum of nine terms ½λₖdₖ², six of which can be nonzero. The λₖ span a factor of 890. So unless a displacement happens to be almost exactly orthogonal to the stiff directions, the stiffest one or two components dominate — a component of half the size on a direction ten times stiffer contributes two and a half times as much.
The exception is instructive. The display constraint’s top two are directions three and two rather than one and two, because its displacement is nearly orthogonal to the stiffest direction: the primaries it uses were chosen by minimising the same objective, so the search pushed the residual out of the direction it cost most in. A constraint optimised against the objective ends up orthogonal to the objective’s stiffest direction, which is a compact way of saying what optimisation under a constraint does.
What was computed, and how
The Hessian is 154 central differences at a step measured to be inside the window where the six real eigenvalues do not depend on it. The eigendecomposition is a one-sided Jacobi singular value decomposition, which resolves the three that ought to be zero rather than losing them in a squared condition number.
Each named basis is normalised on the adopting white before the displacement is taken, so the comparison is between directions rather than between directions and arbitrary units. The designed display’s basis is not one of the eight named matrices — it is the inverse of the primaries a search returns when asked to minimise the same objective over three realisable chromaticities — so it enters the table as a matrix rather than by name.
The assertion in the build has four parts: the walk must converge to within six per cent at the smallest step; the error must halve over the last two intervals; the top two directions must carry at least 55 per cent of every predicted excess; and the cheap constraint’s seen distance must be at least five times smaller than every other entry’s. The third and fourth are what turn a convergence check into a statement about constraints.
Only the last two intervals are used for the halving test, and that restriction is honest rather than convenient: at an eighth of the way to CAT16 the fourth-order term is still visible and the error falls by 0.68 rather than by a half. Requiring the whole walk to halve would be requiring the series to have no terms beyond the third.
What this settles about the confusion points
The receptor construction costs seventy per cent above the unconstrained floor, and the decomposition says where that goes: 74 per cent of it in two directions, the third and the sixth, at a seen distance of 0.891.
The sixth direction is the flattest the objective has. A displacement of that size along the flattest direction would cost almost nothing; the construction pays because it also moves 0.4 along the third, which is thirty times stiffer. So the price is not that the confusion points are far away — Bradford is 0.539 away and costs a quarter as much — it is that they are far away in a direction the model minds.
That is the sentence the earlier measurement was reaching for and could not state, because a bowl measured in random directions has no directions in it to name.
The same question, asked of a device
The arithmetic is not about the observer’s nine numbers in particular, and pointing it at something manufactured makes that plain.
A display’s design is six numbers — three primary chromaticities — and there is no invariance among them: a display with a different green is a different display, so the curvature has rank six out of six and nothing is invisible. What the eigenvalues say instead is a tolerance. They span a factor of 3,792, so the design can move sixty-two times further in its cheapest combination than in its dearest for the same cost in adaptation, and the cheapest combination moves the red and green primaries together and leaves the blue one where it is.
A camera’s three dyes are the same shape of object: three centre wavelengths and three bandwidths, all in nanometres, so the six eigenvalues are directly comparable without any weighting having to be invented. There the condition number is 623, the stiffest direction is the blue dye’s centre at a weight of 0.96, and the flattest is the three bandwidths together.
Both are constraints in the sense of this essay — a manufacturer restricted to a tolerance is choosing a surface — and both show the same thing: which numbers a design can be wrong about is a fact about the objective’s shape rather than about how many numbers there are.
The table’s third column, which it does not print
The essay names the effective curvature as a quantity worth reporting beside the total and then reports the total. Dividing twice the excess by the squared distance recovers it for every row:
| constraint | seen distance | excess | effective curvature |
|---|---|---|---|
| the inverse of three realisable primaries | 0.068 | 0.022 | 9.52 |
| XYZ scaling | 0.553 | 1.399 | 9.15 |
| the three confusion points | 0.891 | 0.677 | 1.71 |
| Hunt–Pointer–Estévez | 0.856 | 0.610 | 1.66 |
| CAT02 | 0.673 | 0.349 | 1.54 |
| Bradford | 0.539 | 0.166 | 1.14 |
| CAT16 | 0.824 | 0.338 | 1.00 |
The seven fall into two groups. Five published transforms meet an effective stiffness between 1.00 and 1.71 — a spread of 1.7 — and the two objects that are not published transforms meet 9.5 and 9.15, a spread of four per cent.
That is the essay’s it points the wrong way made into a number, and it says more than the phrase does. The display constraint and XYZ scaling meet the same stiffness to four per cent, and their excesses differ by 63.6 against a squared-distance ratio of 66.1 — agreeing to four per cent as well. So the two are, as far as this objective can tell, the same direction at different distances: one is the identity and the other is a design optimised to be as near it as three real primaries allow.
And the five published transforms all point somewhere else, into a region seven times slacker. Whatever the fitting to corresponding-colour data was doing, it was moving the answer out of the direction the adaptation objective minds most — which is what fitting to an objective does, and these were not fitted to this one.
The concentration is lower than chance, and one row is not
Between 62 and 95 per cent is offered as high, and the eigenvalue spread it is explained by predicts higher.
For a direction drawn at random in the six visible dimensions, the two stiffest carry (676.6 + 95.1) / 811.9 = 95.0 per cent of the expected excess. So a random constraint would concentrate at 95, and the measured mean across the seven is 74.
Every row but one concentrates less than chance, and that is the finding rather than the concentration. A displacement that fell where it liked would be almost entirely the stiffest direction; these fall systematically off it. The essay explains the display constraint that way — a constraint optimised against the objective ends up orthogonal to the objective’s stiffest direction — and the table says the same is true of nearly all of them.
The exception is Hunt–Pointer–Estévez, at 95 per cent, exactly the random figure. It is also the one transform in the list that was constructed as a set of cone fundamentals rather than fitted to anything, so nothing pushed it off the stiff direction and it landed where an arbitrary displacement would. The fitted transforms have been pulled away from the objective’s stiffest direction by their own fitting; the constructed one has not, and the concentration column is where that shows.
That is a cleaner statement of what fitting buys than the excesses give. HPE sits at 0.856 and costs 0.610; CAT16 sits at 0.824 — almost the same distance — and costs 0.338. Nearly the same displacement, nearly half the cost, and the whole of the difference is which directions the displacement went into.
The walk’s rate, read off
The convergence ratios are quoted as halving each time the displacement halves, three times over, and the four ratios are 1.57, 1.82, 1.92 and 1.92.
So the halving is reached rather than assumed: the first interval is 21 per cent short of two, the second 9 per cent, and the last two are within four per cent. The rate itself converges, which is a stronger statement than the errors falling — it says the third-order term is dying out of the ratio and not merely out of the error.
That also justifies the gate’s decision to test only the last two intervals. At an eighth of the way the ratio is 1.57 and a test demanding two would fail on a series that is behaving exactly as a Taylor series should. The restriction is not leniency; it is the correct window for the claim being made, and the two intervals inside it agree with each other to a part in five hundred.
Where the model stops
Everything here is second order. The quadratic model is demonstrably the second derivative and demonstrably wrong at the distances the published transforms sit at. Nothing in this essay predicts a cost at full displacement; what it does is explain the ordering and the concentration, and give an arithmetic that is exact in the limit and checkable on the way in.
The eigenvalues are a property of the census. Fourteen changes of illumination, equally weighted, over this collection’s family of smooth reflectances. A different census moves every λ and therefore every prediction. The structure — six directions, two dominating, a count that predicts nothing — would survive.
And a constraint here is a surface, not a penalty. Each row is the best basis satisfying its own restriction exactly. A soft constraint, of the kind a real design uses, sits somewhere between the optimum and that point, and its excess is not this excess.
The generalisation
The price of a restriction is a distance times a stiffness, and the number of parameters it removes appears nowhere in that product.
That reads as obvious and is systematically ignored, because a parameter count is available before any work is done and a displacement is not. The habit worth having is the cheap version: before accepting that a constraint is expensive because it is restrictive, check how far it actually moves the answer, in the directions the objective can see. Both halves matter, and the second half is where the invariance lives.
The second lesson is about how to demonstrate a local model at all. Evaluate at the point of interest and the result is unreadable; scale the displacement and the same model becomes testable. The convergence rate is the test: first order in the step for a second derivative, second order for a third. A model that converges at the wrong rate is the wrong model even where it agrees.
Who found it, and when
The arithmetic is the second-order condition from any optimisation text, and the identifiability literature’s version — how a restriction interacts with the curvature of a likelihood — is standard in experimental design, and is the same counting this collection does with a rank, where the whole business of choosing measurements is choosing which directions of a Fisher information matrix to make stiff.
Colour appearance modelling has the constraints and not the diagnostics. Bradford, CAT02 and CAT16 were each obtained by minimising an error over corresponding-colour data; each is reported as a matrix and a residual; and the question of how far each sits from an unconstrained optimum, in which directions, is not asked anywhere this collection can find. It cannot be recovered from a published matrix either, because the matrix is a point and a point has no curvature.
The one place the discipline does reason this way is display design, where the primaries are known to trade against each other and the trade is understood as a surface. That is the cheap constraint in the table above, and it being cheap is not a coincidence: it is the one whose designers were optimising.
Where the ladder goes next
A flat direction is a direction the objective has no opinion about, which makes it free to spend on something else. What it buys, and how asymmetric the exchange turns out to be, is the next question, and the answer is not the one the shape of the trade-off suggests.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A tolerance is a region chromatic adaptation · eigenvalue · hessian · level set
- The best axes are not receptors basis · chromatic adaptation · confusion point · identifiability
- The price is also the person basis · chromatic adaptation · confusion point · identifiability
- A mean has a set under it basis · chromatic adaptation · degrees of freedom
- A template cannot place a point basis · confusion point · identifiability
- Downhill from a published matrix chromatic adaptation · degrees of freedom · eigenvalue
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
BasisChromatic adaptationConfusion pointConstraintDegrees of freedomEigenvalueHessianIdentifiabilityLevel setTaylor series