Where the model breaks

A constraint costs what it points at

Three primary chromaticities remove three of the nine numbers in an adaptation basis and cost one per cent. Three dichromat confusion points remove six and cost seventy. Counting what a constraint removes predicts neither, because an optimum is a long bowl and what matters is which way the constraint points.

Assumes Primaries chosen for their inverse, The three numbers a gain cannot see and A fit can be exact and empty.

These essays have imposed four constraints on the same nine numbers and got four answers spanning two orders of magnitude. Three of them cost almost nothing and one costs seventy per cent, and the one that costs is not the one that removes the most parameters. Counting degrees of freedom predicts none of it.

The same optimum, along its narrowest direction and its widest. The adaptation objective along two straight lines through its own minimum, both of unit length in the nine coefficients. Along one of them the cost rises steeply; along the other the same step costs 8.0 times less, and a design constrained to move that way gives up almost nothing. That is why restricting the nine numbers to be the inverse of three realisable primaries — three degrees of freedom gone — costs about one per cent, while requiring them to hit the three dichromat confusion points costs seventy. Counting what a constraint removes predicts neither number; what matters is which way it points.
Fig. 1 The adaptation objective along two straight lines through its own minimum, both of unit length in the nine coefficients. The cost of a step depends on which way it is taken by a factor of eight.

The claim

An optimum has a shape, and what a constraint costs is set by how the constraint is oriented relative to that shape rather than by how many parameters it removes. The adaptation objective’s minimum is a bowl whose radius to a five per cent rise varies eightfold across a sample of directions — and by thirty across its own principal axes — and the four constraints imposed here land in four different places on it.

  • Three realisable primary chromaticities: removes three of nine, costs 2.3 per cent.
  • Three dye centres and widths: removes three of nine, costs nothing measurable.
  • A Luther residual ceiling on a sensor: costs 0.3 per cent when tightened by a third.
  • Three dichromat confusion points: removes six of nine, costs 70 per cent.
  • The bowl’s aspect ratio is at least 8.0, measured by walking outwards in twenty-four random directions. Taken from the objective’s own eigenvalues rather than from a sample it is 29.8, and why a sample of directions cannot find either end of a bowl in nine dimensions is a result of its own.

Four constraints, four answers

Set out together the numbers make the point without any interpretation.

The unconstrained optimum over all nine entries of an adaptation basis leaves 0.974 ΔE00 averaged over this collection’s illumination census.

Constrain the basis to be the inverse of a matrix built from three realisable primary chromaticities — a display, in other words — and the best available is 0.996. That is a restriction to a six-parameter subfamily and it costs 2.3 per cent.

Constrain it instead to be the basis of a camera whose three dyes are Gaussians with free centres and widths, held within a stated distance of the Luther condition, and the best available is 0.974 — the unconstrained optimum, to four decimal places. Another six-parameter subfamily, and it costs nothing at all.

Constrain it to have the three measured dichromat confusion points as its own — which fixes six of the nine and leaves the other three doing nothing — and the answer is 1.651. It costs seventy per cent.

How much every basis leaves an adapted observer. Eight bases ranked on the mean ΔE00 an adapted observer is left with after the gain, averaged over every change of illumination this collection models. The range runs from 0.97 for best for adaptation to 2.37 for XYZ scaling. The ordering is not the ordering on the other objective and is nearly its reverse.
Fig. 2 The bases compared here. The distance between the top of this ranking and the bottom is what the four constraints are being priced against.

What the two-per-cent numbers do not mean

Before the geometry, a caution, because a run of cheap constraints invites a conclusion the measurements do not support.

Cheap here means cheap on this objective, which is the mean residual an adapted observer is left with over a census of fourteen illumination changes. A display designed for its inverse gives up 2.3 per cent of that, and it may give up a great deal of something else: efficiency, manufacturability, cost, or the discrimination geometry the same nine numbers also decide, which for a display is terrible in every case.

The dye design’s cost of nothing is the same statement with the same scope. Those three narrow filters throw away most of the light that reaches them, and the noise consequence is not in this arithmetic at all.

So the four numbers are prices in one currency, and the exercise here is comparing prices in that currency rather than deciding what anybody should build. What makes the comparison worth making is that the four constraints are all on the same nine parameters and all priced by the same objective, which is a like-for-like comparison of a kind that is usually unavailable.

Why counting parameters fails

The intuition that a constraint removing six numbers should cost more than one removing three is an intuition about volume: fewer directions to move in, so less room to find a good answer.

It fails because the objective is not spherical. Near its minimum the adaptation residual rises quadratically, and the quadratic form has eigenvalues spread over a wide range — so some directions are cheap to be pushed along and others are expensive. A constraint that happens to lie along the cheap directions costs almost nothing however many parameters it removes; one that points along an expensive direction costs a great deal even if it removes only one.

Walking from the receptors to the best adaptation basis. Two curves along a straight path in the nine coefficients, from the basis built out of the dichromat confusion points at the left to the basis that minimises the adaptation residual at the right, with every row renormalised on the white so that each stop is a legitimate basis rather than a blend of two pictures. The adaptation residual falls from 1.65 to 0.97 and the ellipse anisotropy rises from 2.60 to 7.70. There is no stop where both are good and no kink where a compromise would sit.
Fig. 3 A straight path from the receptor basis to the adaptation optimum. The residual falls at every stop, which is what a constraint pointing down a narrow direction looks like from the inside.

The measurement of that spread is direct. Walk outwards from the optimum in twenty-four random unit directions in the nine coefficients, and in each direction find by bisection how far one can go before the objective rises by five per cent. The radii run from 0.022 to 0.177, with a median of 0.045 — a factor of eight between the narrowest direction sampled and the widest. That is a lower bound rather than the shape: a random direction in nine dimensions carries a share of every principal direction, so it reports the middle of the bowl and never an end, and the six eigenvalues put the true ratio at 29.8.

That is the whole explanation. The primaries and the dyes are, in effect, moving mostly along wide directions; the confusion points are pointing down a narrow one.

Three published primary sets, and a fourth chosen for how it adaptsThe spectral locus with four triangles inside it. sRGB covers 33.5% of the diagram and leaves an adapted observer 2.36 ΔE00; Display P3 covers 45.4% at 1.22; Rec. 2020 covers 63.3% at 1.09. The fourth triangle is the best adaptation basis available to a display asked to cover 63.5% of the diagram, at 1.02 — and it is a different triangle from Rec. 2020's rather than a smaller one. The largest triangle that fits at all covers 73.9%, which is where the axis of this argument ends.0.00.20.40.60.80.00.20.40.60.8xy460480500520540560580600620sRGB 2.36Display P3 1.22Rec. 2020 1.09designed 1.00D65the number beside each is its adaptation residualCIE 1931 2° observer · the device chooses the basis
Fig. 4 The display family’s own answer, drawn. The designed triangle is a different triangle from Rec. 2020’s rather than a smaller one.

Why the device constraints are cheap

There is a structural reason as well as a geometric one, and it is worth having because it explains why the result is not luck.

The space of distinguishable adaptation bases is six-dimensional, not nine, because three of the nine numbers do nothing at all to a von Kries gain. Three primary chromaticities are six numbers. Three dye centres and widths are six numbers. So neither device constraint is a restriction of dimension at all — both parametrise the whole effective space, and the only restriction is to the region of it that corresponds to realisable hardware.

The measurement then says something about that region: it contains the optimum, or comes within a few per cent of it. The camera family contains it outright. The display family comes within 2.3 per cent, and the shortfall is the requirement that three rows be duals of three points inside the spectral locus.

Nothing guaranteed that. The optimum could have been a basis with rows corresponding to chromaticities outside the locus, in which case no display could approach it. It is not, and that is a measured fact rather than a theorem.

What a display gives up in adaptation to cover more of the diagram. The best adaptation residual reachable by three realisable primaries, against the share of the CIE xy diagram they are required to cover. The curve runs from 1.00 ΔE00 at 53.8% to 1.07 at 71.9%, which is nearly flat: the trade a reader expects to be a wall costs 7.9% across the whole range. The three published sets are marked and all three sit above the curve — sRGB by a factor of 2.37. The dashed level is the best any nine numbers can do at 0.974, which no display quite reaches.
Fig. 5 The display family’s own trade curve, with the unconstrained optimum drawn as a level. The whole curve lives within a few per cent of it, which is what a constraint pointing along wide directions looks like.

Why the confusion points are not

The receptor constraint is different in kind. It does not restrict to a region of the effective space; it picks a point in it, because six numbers determine the basis and the other three are invisible. There is nothing left to optimise.

And the point it picks is 1.651 against 0.974. On the bowl’s own scale that is far outside every radius measured: the receptor basis is not near the optimum in any direction.

That is the answer to the question these essays were set. The misses the published transforms make against the confusion points are not slack in a fit; the confusion points are a long way from where the objective wants to be, in a direction the objective is steep along. Walking from one to the other improves the residual at every step, with no interior minimum and no shoulder.

The same shape, in the other objective

The bowl is not a property of the adaptation objective alone, and checking the second one is what turns an observation into a pattern.

The discrimination objective — how nearly a lightness–chroma space makes the ellipses circles — has its own optimum and its own shape, and measured the same way its radii span a factor of 5.8. So it is a long bowl too, less extreme than the adaptation objective’s eight but far from round.

That has a consequence for a claim made in a neighbouring essay. The discrimination optimum is sharply determined: nine of twelve searches from independent random starts reach the same matrix to four decimals. A sharp optimum and a long bowl sound contradictory and are not — the minimum is unique and well located, and the level sets around it are elongated. The first is about whether the search finds one answer and the second is about how much it costs to be pushed off it.

The two properties answer different questions and are routinely conflated. A well-determined optimum is not the same thing as an expensive one to leave.

Every basis against both objectives at once. A scatter with the mean adaptation residual across the illumination census on the horizontal axis and the mean axis ratio of MacAdam's ellipses in a lightness–chroma space on the vertical. Lower is better on both. The two winners sit at the two ends of an empty diagonal: the basis that adapts best leaves 7.70 on the vertical and the basis that discriminates best leaves 1.79 on the horizontal, each worse on the other objective than every published transform. The basis built from the dichromat confusion points is at (1.65, 2.60) — best at neither and within a factor of two of both floors, which no other entry in the picture manages.
Fig. 6 The bases compared here, on both objectives. Each of the two winners sits at the bottom of its own long bowl, and each is a long way up the other’s.

What a flat direction usually means

The corollary above deserves its own treatment, because it is the practically useful half and it has a name in other subjects.

If moving along a direction does not change the objective, then the data the objective is built from cannot distinguish the models at either end. That is non-identifiability, and this collection has met it twice already: nine numbers a colour match cannot see, and three numbers a von Kries gain cannot see. Both are exact flatnesses, and in both cases the response was to find a different experiment that could see the direction.

What the bowl adds is the graded version. A direction need not be exactly flat to be nearly uninformative, and the radii measured here run continuously from 0.022 to 0.177 with nothing special happening in between. So the useful picture is not identified parameters and unidentified ones but a spectrum of how well each combination is pinned, which is what a condition number is.

The practical reading of a cheap constraint follows directly. It is free because it points somewhere the data are quiet, so imposing it costs nothing and buys nothing in the way of validation: a model that fits equally well with the constraint as without has not been tested by the constraint. That is the trap a residual with no orbit beside it sets, and it is the same trap seen from the constraint’s side rather than the fit’s.

Two ways to be cheap, and they are opposites

The claim above is that orientation rather than parameter count sets the price. Decomposing each constraint’s displacement onto the bowl’s own eigen-directions sharpens that, and in one place corrects it.

The excess a constraint pays is half the effective curvature it meets, times the square of how far it is forced from the optimum. Both factors are measurable, and for the two extremes of this page they are:

constraint distance effective curvature excess
a display designed for its own inverse 0.099 5.08 0.025
the three confusion points 2.293 0.279 0.733

The product reproduces the measured excess in both rows to three decimal places, so this is a decomposition rather than a model.

The two cheap and expensive cases are cheap and expensive for opposite reasons. The designed display sits in a stiff part of the bowl — 5.08 against 0.279, eighteen times stiffer — and is cheap purely because the constraint’s feasible set passes within a tenth of the optimum. The confusion-point basis sits in one of the softest parts of the bowl and is expensive purely because it is dragged 2.3 units away. Distance ratio 23, squared 540; curvature ratio 18 the other way; net cost ratio 30.

So a constraint is cheap when it is near, or when it is soft, and neither alone predicts anything. Reading the display’s 2.3 per cent as evidence that a display’s constraint points somewhere harmless gets the geometry exactly backwards: it points at the third and second stiffest directions in the whole bowl, which carry 69 per cent of what it does pay. It simply does not have to go far.

The eigenvalue spectrum underneath explains why a distance can be almost free at all. The nine values run 677, 93.6, 23.1, 11.9, 4.55, 0.760, and then 1.91 × 10⁻⁴, 3.63 × 10⁻⁵ and 1.77 × 10⁻⁵. The gap between the sixth and the seventh is a factor of four thousand. Six directions are real curvature and three are numerically flat, which is the rank the objective has by construction — a basis is nine numbers and the objective cannot see a row’s scale. A constraint whose whole displacement lay in those three would be free, and none of the eight measured here manages it: the smallest distance in the table still meets a curvature of five.

What was computed, and how

The bowl’s shape is measured rather than derived from a Hessian, and the choice is deliberate: a Hessian is a local object and the constraints being priced are not local.

From the optimum, twenty-four pseudo-random unit directions in the nine coefficients are generated from a stated seed. In each, the objective is evaluated at increasing distances until it exceeds the minimum by five per cent, and the crossing is then bracketed by thirty bisections. The radii are sorted and the ratio of largest to smallest is reported.

The absolute radii are meaningless on their own — they depend on how the nine numbers are scaled, and the objective is invariant to scaling a row, so three of the directions are exactly flat and are excluded by the bisection running to its limit. The ratio does not depend on the parametrisation in the way the radii do, which is why it is the number quoted.

The four constraint costs are measurements made here, each from its own search: the display over six chromaticity coordinates with soft penalties, the sensor over three centres and three widths with a soft penalty on the Luther excess, and the receptor case with no search at all because there is nothing to search.

The assertion in the build requires the aspect ratio to be at least four and requires the receptor constraint to cost at least forty per cent, so a change that made the objective spherical or made the receptors cheap would stop the build.

A fifth constraint, for contrast

The four above are all constraints on the basis. A fifth was imposed on something else, and it behaves the third way, which completes the picture.

The gamut floor in the display design — the requirement that the primary triangle cover a stated share of the diagram — is a constraint on the feasible region rather than on the dimension. Tightened from nothing to 63.5 per cent of the diagram it moves the best available residual from 0.996 to 1.018, which is 2.2 per cent; tightened to 72 per cent it moves it to 1.075, which is eight.

So it costs something, it costs increasingly, and the cost is still small in absolute terms. That is what a constraint that points partly along a narrow direction looks like: a smooth, monotone, gentle penalty rather than either the nothing of the dye family or the seventy per cent of the receptors.

Three behaviours, then, from five constraints on one objective:

  • Free: the constraint’s feasible set contains the optimum.
  • Gentle: it excludes the optimum but the exclusion is along a wide direction.
  • Expensive: it picks a point, and the point is far away down a narrow one.

Only the third is what the parameter count would predict, and it is the only one of the three where the constraint is a measurement rather than a design restriction. That is not a coincidence and it is the last thing worth saying: a constraint derived from data can disagree with an objective, and a constraint derived from what can be manufactured usually cannot, because manufacturability was never in tension with anything.

How far from circles every basis leaves the ellipses. Eight bases ranked on the mean ratio of the long to the short axis of MacAdam's twenty-five discrimination ellipses, measured in a lightness–chroma space built on that basis. The range runs from 1.61 for best for discrimination to 7.70 for best for adaptation. The ordering is not the ordering on the other objective and is nearly its reverse.
Fig. 7 The same eight bases on the other objective, which has its own bowl and its own aspect ratio of 5.8.

Where the model stops

Twenty-four directions is a sample. The true extreme aspect ratio of the level set is at least eight and could be more; a sample of twenty-four unit vectors in nine dimensions rarely finds the extremes.

Five per cent is a choice. The bowl’s aspect ratio depends on how far out it is measured, because the objective is only approximately quadratic. At one per cent the ratio would be closer to the Hessian’s condition number and at twenty per cent it would be dominated by the objective’s global shape.

And the four constraints are not four samples of a population. They are the four that happened to be imposed here, chosen because each answers a question somebody had asked. A different four would give a different spread, and nothing here says how constraints are distributed in general.

The generalisation

Before pricing a constraint by how much freedom it removes, measure the shape of what it is being imposed on.

The check is cheap and is the one performed here: from the optimum, walk in random directions and record how far each goes before the objective rises by a stated amount. If the radii are all alike, the objective is round and the parameter count is a fair guide. If they span an order of magnitude, the parameter count says nothing and the orientation is everything.

The corollary is the more useful half. A cheap constraint is evidence that the objective is flat in the direction the constraint points, and flatness in a direction means the data cannot see that direction. So a constraint that costs nothing has usually been imposed on something the measurement never determined — which is the same observation this collection made about a residual that is flat exactly where the data are silent, read from the other end.

That gives the pair of readings a practitioner wants. A constraint that costs a great deal is a statement that the data disagree with it. A constraint that costs nothing is a statement that the data had no opinion, and imposing it is free because there was never anything there to lose.

Where the ladder goes next

Every result in these essays has been about a linear stage — a basis, a diagonal, a matrix in front of a nonlinearity. What the constraint arithmetic here suggests is that the interesting question about any of them is the shape of the objective rather than the value at its minimum, and the objectives this collection has built are all evaluated at a single point.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 12 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AssertionBasisChromatic adaptationConditioningDegrees of freedomIdentifiabilityMeasurement errorOptimisationSpecificationTrade-off