What the eye does

The rank is the invariance

A von Kries gain cannot see the scale of a row of its basis. That is an identity, proved in a line, and it can be measured instead — as the rank of a second-derivative matrix. Both objectives this collection minimises over the observer's nine free numbers have a Hessian of rank exactly six, and the three directions they cannot see are the three scalings, to a hundredth of a degree.

Assumes The three numbers a gain cannot see, The matches do not name the cones and A gain needs a basis.

An invariance proved on paper says the answer does not move. An invariance measured as a rank says something stronger: that nothing else is invisible either.

Nine eigenvalues, six of which exist. Nine points on a logarithmic vertical axis: the eigenvalues of the Hessian of the adaptation residual at its own optimum, largest to smallest. The first six run from 6.8×10² down to 7.6×10⁻¹, a condition number of 890. Then the axis drops: the seventh is 1.9×10⁻⁴, and the last three are separated from the sixth by a factor of 4.0×10³. Those three are not small curvatures. They are the finite-difference truncation error on directions along which the objective is exactly constant, and a shaded band marks them as the numbers the objective does not have.
Fig. 1 Nine eigenvalues of the second-derivative matrix of an adaptation objective at its own optimum, on a logarithmic axis. Six of them, then a fall of four thousand.

The claim

The Hessian of either objective this collection minimises over the observer’s nine free numbers has rank exactly six at its own optimum, and the three-dimensional space it annihilates is the space of row rescalings, to a principal angle of a hundredth of a degree.

  • The gap is decisive. The sixth eigenvalue is 0.760 and the seventh is 1.9 × 10⁻⁴, a factor of nearly four thousand for the adaptation objective and twelve thousand for the other.
  • The three small values are not small curvatures. They are the finite-difference truncation error on a quantity that is exactly zero, and they grow as the square of the step while the six real ones move by 0.42 per cent over a decade of it.
  • The null space is the row scalings, at principal angles of 0.025°, 0.0008° and 0.0003°.
  • And the measurement says more than the proof. The algebra says three particular directions are flat. The rank says there are no others: the remaining six are all seen, at curvatures spanning a factor of 890.
  • Away from the optimum the rank is nine, and that is not a contradiction. It is the difference between an invariance that is also an extremum and one that is not.

What the algebra already said

The observer’s three curves are determined by colour matching only up to a nonsingular 3×3, so nine numbers are loose and a match cannot name the cones that made it. An adaptation model in the von Kries form takes a basis B, reads the two whites in it, divides one by the other channel by channel, and comes back: the operator is B⁻¹ diag(d) B.

Scale row i of B by any nonzero k. Then (Bw)ᵢ is multiplied by k for both whites, so dᵢ is unchanged; and column i of B⁻¹ is divided by k, which cancels the k that row i introduced. The operator is identical. Three of the nine numbers do nothing at all, whatever the gain is trying to remove, which is where the counting argument this collection is built on comes from: the six a dichromat experiment supplies are exactly the six the model can see.

That argument is settled and this essay does not improve it. What it does is measure the same fact in a form the algebra does not reach.

Why a rank says more than an identity

The identity is a statement about three specific directions. It rules those three out and says nothing about the other six. A model could easily be nearly flat in a fourth direction for a reason nobody had noticed — a near-degeneracy rather than an exact one, invisible to any amount of algebra and fatal to any claim that a fit has six parameters to fit.

The Hessian answers that. Its eigenvalues are the curvatures along its own principal directions, so the number of them that are nonzero is the number of directions the objective can distinguish at all, and their spread is how unequally it distinguishes them. Rank six is the statement that there is no seventh flat direction hiding anywhere in the nine.

Measured at the adaptation optimum, the nine eigenvalues are

676.6, 93.59, 23.11, 11.89, 4.553, 0.760, then 1.9 × 10⁻⁴, 3.6 × 10⁻⁵, 1.8 × 10⁻⁵.

Six real ones spanning a factor of 890, and then a cliff. The rank decision is backed by a ratio of 3,973 between the sixth and the seventh, which is a fact rather than an opinion — a rank taken at a threshold with a gap of three is a guess, and this one is not.

At the other objective’s own optimum the same structure appears with different numbers: 188.5, 39.63, 8.381, 4.209, 1.589, 0.254, then values four orders of magnitude below, and a gap of 11,940.

The three directions the curvature cannot see are the three row scalings. Three horizontal bars, one per principal angle between the null space of the Hessian of the ellipse anisotropy and the three directions along which a row of the basis is rescaled. All three angles are under a tenth of a degree — 3.3e-3°, 5.3e-4°, 1.1e-4° — so the two three-dimensional subspaces are the same subspace. This is the algebraic invariance that a von Kries gain cannot see the scale of a row, measured as a geometry rather than assumed.
Fig. 2 The three principal angles between the Hessian’s null space and the three directions along which a row is rescaled. All three are under a hundredth of a degree.

How to tell a zero from a small number

The three small eigenvalues are not zero on a computer, and the difference between small and zero is the whole of the claim. A finite-difference second derivative has two errors: truncation, which falls as the square of the step, and roundoff, which rises as one over the square of it. A quantity that is genuinely zero shows only the first.

So the test is not the size of the number, it is how it behaves when the step changes. Over a decade of steps, from 2 × 10⁻⁴ to 2 × 10⁻³:

  • the six real eigenvalues move by 0.42 per cent on the discrimination objective and by a few per cent on the other, which is the objective departing from a quadratic at the larger steps rather than anything about the method;
  • the three small ones grow by a factor of a hundred, for a step ten times larger, which is exactly.

A curvature that scales as the square of the ruler is not a curvature. It is the ruler.

Six eigenvalues that stay and three that fall as the step squared. Nine curves on logarithmic axes: each eigenvalue of the Hessian of the ellipse anisotropy against the finite-difference step it was measured at, from 2×10⁻⁴ to 2×10⁻³. Six of them are nearly horizontal: the largest two move by 0.02 per cent over the whole decade and the worst of the six by 0, which is the objective departing from a quadratic at the larger steps rather than anything about the method. The other three grow by a factor of 100 as the step grows by 10, which is the square law a truncation error obeys on a quantity that is exactly zero. The gap between the sixth curve and the seventh is 1.2×10⁴ at the smallest step.
Fig. 3 Nine eigenvalues against the step they were measured at, on logarithmic axes. Six horizontal lines and three lines of slope two.

That is also the answer to the obvious objection, which is that the step was chosen to produce this. It was not chosen at all: the window is measured, and the assertion in the build requires both halves — the kept eigenvalues to be flat and the dropped ones to obey the square law — so a step outside the window fails on one or the other.

Away from the optimum the rank is nine

Here is the part that had to be got right, and that a looser statement of the claim would have got wrong.

Take the same two objectives and differentiate them twice at the receptor basis, which is neither one’s optimum. Both Hessians come back with rank nine. Nothing has broken: the invariance still holds exactly there, because scaling a row still leaves the operator identical.

What changed is what an invariance looks like. At a minimum, the objective is constant along a three-dimensional orbit of minima; the orbit is a set of minima, so the gradient vanishes on it and the Hessian annihilates its tangent space, and the rank falls. Away from a minimum the orbit is still a set of constant value, but the gradient is not zero — it is merely orthogonal to the orbit — and what survives is weaker and still exact: the quadratic form vanishes on those directions, dᵀHd = 0, while Hd does not.

Measured at the receptor basis, dᵀHd on the three scaling directions comes to between 7 × 10⁻¹¹ and 7 × 10⁻⁸ of the largest eigenvalue, for both objectives. The form vanishes; the matrix does not.

An invariance is a null space only where it is also an extremum. That is a small piece of differential geometry, it is the kind of thing that gets stated backwards in a caption, and the only way to be sure which version is true of a given case is to compute both.

What a nullity of four would have meant

The measurement is only worth making because it could have come out otherwise, and it is worth being explicit about what each outcome would have said.

Nullity three — the result — says the model’s freedom is exactly the freedom the algebra names, and the remaining six numbers are all doing something. A fit over the nine has six real parameters and three that any starting point resolves arbitrarily.

Nullity four or more would have said there is a redundancy nobody noticed: some fourth combination of the entries the model cannot see, which no amount of data would fix and which a fit would resolve by whichever way its search happened to drift. That is the outcome the identifiability literature spends its time on, it is common in models with more than a handful of parameters, and no proof would have found it — a proof answers the question it was asked.

Nullity three with a small gap would have been the least comfortable answer: a seventh direction with a curvature a hundredth of the sixth is not flat, but it is flat enough that any real data set would leave it undetermined. The gap here is 3,973 on one objective and 11,940 on the other, which puts the question beyond argument.

And a nullity of two would have meant the algebra was wrong, which it is not.

Which row of the basis each direction moves. A grid with one column per eigen-direction of the curvature of the adaptation residual, stiffest on the left, and one row per row of the basis. Each bar is that row's share of that direction. The last three columns — the ones the objective cannot see — each have a single full-length bar and two empty ones: a direction that rescales one row and touches nothing else, which is exactly what the invariance says they are, visible without any angle being computed. Among the six the objective does see, the stiffest is 97 per cent the short-wave row and the flattest is 85 per cent the long-wave one, so an adapted observer's model is most particular about where its blue axis points and least about where its red one does.
Fig. 4 Each direction as a share of the three rows of the basis. The three the objective cannot see are each one row and nothing else, which is the invariance visible without any angle being computed.

What the six say

Having six rather than nine is a structural fact. What the six are is a quantitative one, and the spread is the interesting half: 676.6 down to 0.760, a factor of 890.

The square root of that is the ratio of distances: a step of the same cost is 29.8 times longer along the flattest direction the objective can see than along the stiffest. That number is what a sample of random directions could not find — it reported 8.0 — and it is what makes the difference between two constraints of the same nominal size.

The bowl the eigenvalues describe and the bowl a sample found. Six points on a logarithmic vertical axis — the distance from the optimum of the adaptation residual to a 5 per cent rise along each of the six directions the objective can see — with a shaded band behind them showing the whole range 24 random directions reported. The eigen-radii run from 1.2e-2 to 3.6e-1, a factor of 29.8. The band runs from 2.2e-2 to 1.8e-1, a factor of 8.0, and sits entirely inside the ends of the true range: a random direction in nine dimensions carries a share of every eigenvector and so reports the middle of the bowl, never an end of it.
Fig. 5 The six radii the eigenvalues give, with the band twenty-four random directions reported drawn behind them.

It also settles something about the fit the literature performs. Bradford, CAT02 and CAT16 were obtained by minimising an error over corresponding-colour data with all nine entries free, and none of them is a set of cone responses. Three of those nine did nothing, so the fit had six parameters and three directions along which any starting point was as good as any other. That is not a defect in the published matrices — the resulting operator is the same whichever member of the flat family the search happened to stop at — but it does mean that the entries of a published adaptation matrix are not determined by the fit that produced them, and comparing two of them entry by entry is comparing three real numbers and three conventions.

Normalising each row on the adopting white, which is what this collection does before any comparison, removes exactly that freedom.

The two bowls have the same shape

Both objectives were measured at their own optima, and setting the two spectra beside each other gives something neither of them was measured for.

Adaptation: 676.6, 93.59, 23.11, 11.89, 4.553, 0.760. Discrimination: 188.5, 39.63, 8.381, 4.209, 1.589, 0.254. Divide one by the other entry by entry and the six ratios are 3.589, 2.361, 2.758, 2.826, 2.865 and 2.993 — a mean of 2.899, with a spread of only 1.52 between the extremes, across six curvatures that individually span factors of 890 and 742.

Normalised to their own largest eigenvalue the two spectra read 1, 0.138, 0.034, 0.018, 0.0067, 0.0011 and 1, 0.210, 0.044, 0.022, 0.0084, 0.0014. Two objectives measured on unrelated data — fourteen changes of illumination against twenty-five discrimination contours — in unrelated units, at two different points of the nine-dimensional space, have curvature profiles agreeing to within a fifth everywhere but the second entry. The two bowls are the same shape, to a factor of about three in overall depth.

Nothing predicted that, and nothing here explains it. Both objectives are means of a colour-difference formula over a family of stimuli, computed through the same basis, so a family resemblance is plausible after the fact; it was not derivable before it, and a factor of 1.5 of spread over a range of 890 is closer than a family resemblance has any right to be.

And they are pointed differently

Shape is not orientation, and the orientations part company where it counts.

Two six-dimensional subspaces of a nine-dimensional space must meet in at least three dimensions, so three of the principal angles between the objectives’ seen subspaces are zero by arithmetic rather than by agreement. The informative angles are the other three: 8.09°, 41.55° and 52.58°.

So the two objectives look at nearly the same six-dimensional slice of the nine — a fourth direction shared to within eight degrees — and separate decisively in two, at forty-two and fifty-three degrees.

That is the structure underneath everything this collection has said about the pair. They are not orthogonal, which would make them independent problems with independent answers; they are not aligned, which would let one basis serve both. Four of their six directions are shared closely and two are not, and those two are where the whole trade between them lives.

What was computed, and how

The Hessian is 46 second differences and therefore 154 evaluations of the objective: the diagonal from three each, every off-diagonal from four. For the adaptation objective one evaluation integrates a family of two hundred and forty constructed surfaces under fourteen changes of illumination, so the whole matrix is about two seconds; for the other it is forty milliseconds.

The eigendecomposition is a one-sided Jacobi singular value decomposition of the matrix itself rather than of its Gram matrix. The Hessian is symmetric and, at a minimum, positive semidefinite, so its singular values are its eigenvalues — and the Jacobi route computes small singular values to high relative accuracy, which matters here more than anywhere else, because the three numbers being asked about are the ones that ought to be zero. Forming HᵀH first would square the condition number and put them below the noise floor.

The principal angles between the null space and the scaling directions are the arccosines of the singular values of the matrix of inner products between two orthonormal bases. The three scaling directions are automatically orthogonal, whatever the basis, because each lives in a different three of the nine coefficients — which is worth saying rather than orthogonalising for, since a Gram–Schmidt there would run and change nothing and would leave the impression that the rows had to be orthogonal for the argument to work.

Six numbers predicting twenty-four measurements. A scatter of 24 points on logarithmic axes, one per random direction walked out from the optimum of the adaptation residual. The horizontal position is the radius measured by bisecting the true objective; the vertical is the radius the six eigenvalues predict, with nothing fitted. The points lie on the diagonal to a median of 7.6 per cent. The sample measured nothing the curvature did not already contain — which is the point, because the sample nevertheless reports a bowl 8.0 times longer one way than another where the eigenvalues say 29.8.
Fig. 6 Twenty-four radii measured by bisection against the twenty-four the six eigenvalues predict, with nothing fitted. If the rank were wrong, this picture would not be a diagonal.

Where the model stops

Rank six is a fact about a von Kries model. A model with an off-diagonal term, a nonlinearity between the two matrices, or a second stage would see the row scales, and its Hessian would have rank nine. The invariance is a property of B⁻¹ diag(d) B specifically and disappears the moment anything happens between the two matrices that is not a per-channel multiply.

A Hessian is a local object and both optima are points. Everything here describes the neighbourhood of two particular bases. The published transforms sit a long way outside that neighbourhood — at full displacement the quadratic model is out by anything from five per cent to a factor of three — and what the curvature does and does not predict about them has to be established by walking in rather than asserted at the edge.

And the objectives are constructions. The adaptation residual is a mean over a census of illumination changes somebody chose, and the anisotropy is a mean over twenty-five ellipses measured on one observer. The rank is a property of the model and would survive any census; the six eigenvalues would not.

The generalisation

An invariance is testable as a rank, and testing it that way is a stronger claim than proving it. A proof rules out the directions it names. A rank rules out every direction there is.

The habit generalises past this subject and is cheap: after establishing that a model is invariant to some parameters, differentiate it twice and count. If the nullity matches the number of parameters proved redundant, the account is complete. If the nullity is larger, there is a redundancy nobody has noticed — which is the interesting outcome and is the one a proof cannot produce, because a proof only answers the question it was asked.

The second half is the distinction between a null space and a vanishing form. Invariance plus extremum gives a null space; invariance alone gives a form that vanishes on a subspace. Confusing the two produces a claim that fails at every point except the one it was checked at, and it fails quietly, because a rank computed at the wrong place comes back as a plausible number with no gap in it.

What a constraint costs is how far it pushes, in the directions that are seen. A scatter of every constraint imposed here on the nine free numbers. The horizontal axis is the length of the displacement from the optimum measured only in the six directions the objective can see; the vertical, on a logarithmic scale, is the excess cost that displacement actually carries. Requiring the basis to be the inverse of three realisable display primaries sits at the bottom left, at 0.068 and 0.022 ΔE00 — it removes three degrees of freedom and moves the answer almost nowhere. Requiring it to hit the three dichromat confusion points removes six and pushes 13 times as far, for 0.68. The vertical spread at similar horizontal positions is the part a count of parameters cannot predict.
Fig. 7 Every constraint imposed on the nine numbers, plotted against how far it pushes in the six directions the objective can see. The three it cannot see contribute nothing, and contribute nothing exactly.

Who found it, and when

Von Kries proposed the diagonal in 1902 and the scale invariance is immediate from it, so the identity is as old as the model. What is not old is anybody’s having asked how many other directions the model is blind to, and the reason is that the question needs an objective — the model on its own has nothing to be flat in.

The rank-and-nullity habit belongs to the identifiability literature in systems biology and pharmacokinetics, where a model with fifteen parameters and eight identifiable combinations is the normal case and the profile-likelihood and rank tests for finding them are standard. Colour appearance modelling has not generally adopted them, and the reason is visible in the shape of the field: its models are fitted to corresponding-colour data sets by minimising an error, the fit converges, and a fit that converges does not announce that three of its parameters were doing nothing.

Where the ladder goes next

Six curvatures spanning a factor of 890 is a very long bowl, and the immediate question is how long, in what directions, and whether the directions have names anybody would recognise. They do, and the flattest of them turns out to be worth more than any of the others for a reason that has nothing to do with adaptation.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

BasisChromatic adaptationCone fundamentalsDegrees of freedomEigenvalueHessianIdentifiabilityInvarianceRankThe von Kries transform