Where the model breaks

Which of these is a convention

Nine ordinary sentences from this collection, put through one test — do they mean the same thing after the observer's three curves are replaced by a nonsingular combination of themselves? Five survive and four do not, and none of the four is wrong, because each is a statement about a set of coordinates being read as a statement about an eye.

Assumes A fit can be exact and empty, The diagram has no area and The matches do not name the cones.

Colour matching leaves nine numbers in the observer undetermined, and this collection has been quoting a particular choice of them for its whole life. That is a fact about the site rather than about colour, and the interesting question is not whether it is embarrassing but which sentences it touches.

There is a mechanical test. Apply the freedom, and see what moves.

Which of this collection's own claims survive a change of basis, and which are about the paper. Nine sentences this site says, sorted by whether they mean the same thing after the observer's three curves are replaced by a nonsingular combination of themselves. 5 of the nine do. The four that do not are not thereby wrong — they are statements about a chosen set of coordinates, and they are true of those coordinates. What they cannot be is statements about the eye, which is how every one of them is usually read.
Fig. 1 Nine sentences this site says, sorted by whether they mean the same thing after a change of basis. Five survive; four do not.

The same question can be put to an instrument and to a scene rather than to a sentence, and in both places it has a number attached.

The part of a sensor a chart reaches, and the part it does not. Each channel's sensitivity, and the residue left after projecting it onto the span of the 12 illuminated patch spectra — the part no measurement of that chart constrains. A second camera differing by that residue produces identical raw values on every patch, to 2.9e-6 relative, and differs by 1.1% on a narrow light at 546 nm. The residue is 3.5% of each curve, which is small — patches that differ in where they are dark are a good probe, and that is the design rule.
Fig. 2 What a measurement leaves undetermined, on a camera: each channel’s sensitivity carries a residue no chart constrains, so two cameras differing by exactly that residue are the same camera to every fit ever made of them. What fills that gap is a convention.
How short of determining the light a photograph is, as the scene grows. Each cell is the number of unknowns left over after every equation the image supplies: three sensors, a three-dimensional illuminant, and reflectances confined to a linear model of the dimension on the left. At one and two dimensions more surfaces close the gap. At three the gap never closes, because each further surface adds three equations and three unknowns; at four it widens. The count is arithmetic and has no algorithm in it.
Fig. 3 And on a scene: how many unknowns are left once every equation an image supplies has been used. Where the count is positive the gap is filled by a choice, and a choice made the same way every time stops looking like one.

Two more of the same three columns say that this question has an arithmetic form wherever anything at all is fitted, and that the arithmetic does not settle it.

What a camera matrix reports on its own chart, and what it delivers off it. The same camera fitted on charts of increasing chromatic range. The left bar of each pair is the mean error on the chart the matrix was fitted to, which is the number a profile comes with; the right bar is the error on a saturated set it never saw. At the thinnest chart the fit reports 0.19 ΔE00 and delivers 1.65, a factor of 8.6. The gap closes as the chart widens, and it closes because the chart improves rather than because the camera does.
Fig. 4 A camera profile, where the count comes out determined. Nine numbers, nine equations, nothing free — and the sentence “this profile is right” still means something different on a different chart.

What a model keeps of an object it cannot fully represent is the third case, and it is the one where the convention is easiest to mistake for a measurement.

How many numbers a surface takes, on the friendliest set available. The cumulative share of the variance accounted for by the first few principal components of this collection's 240 constructed reflectances. Three components reach 99.94 per cent and two reach 97.99. The counting result allows two, and these are deliberately smooth curves built from a handful of Gaussians — the friendliest possible case, and it already needs 3.
Fig. 5 And how many dimensions a reflectance actually takes. Where a model keeps fewer than the object has, what fills the difference is not a measurement, and calling it one is the convention this essay is about.

The claim

A claim is about the observer only if it is invariant under the group the observer’s data leave free. Five of the nine sentences audited here are; four are statements about a chosen set of coordinates. None of the four is false, and every one of them is routinely read as though it were about vision.

  • Survives: these two spectra match; this colour is inside the gamut; these three lights lie on a line; this many surfaces are reachable; the ellipses cannot be made circles.
  • Does not survive: two thirds of the diagram is unreachable; this diagram is more uniform than that one; this is what a protanope sees; this ΔE00 is 2.4.
  • The test is not a matter of judgement. For four of the nine the answer is a residual of order 10⁻¹⁵; for the rest it is a measured spread — 4.5× for the diagram fraction, 3.8× for the difference formula’s uniformity.
  • And two of the four have invariant replacements, which is the useful half: counting stimuli replaces an area, and a floor over the group replaces a comparison between two diagrams.

The five that survive

“These two spectra match.” The strongest case and the reason the group is the right one: a match is an equality, and applying a matrix to both sides of an equality changes nothing. The residual over sixty-four random bases is 2.8 × 10⁻¹⁵.

“This colour is inside the gamut.” Membership is decided by three linear inequalities on tristimulus values, and a change of coordinates carries the inequalities along with the point. Every hatched cell in every diagram on this site stays hatched.

“These three lights lie on a line.” Collinearity in a chromaticity diagram means one light is an additive mixture of two others, which is what the coordinates are linear in. The worst relative departure over forty random bases is 9.8 × 10⁻¹⁵.

“This many surfaces are reachable.” A count of stimuli, and each membership test is invariant, so the count is. This is the invariant replacement for the first of the failures below, and it is the whole reason the audit is worth doing rather than merely being alarmed by.

“The ellipses cannot be made circles.” A floor over the entire group is invariant by construction: the statement quantifies over every basis, so no basis can be a counterexample. That is what makes 2.02 a fact about the eye rather than about anybody’s primaries.

The four that do not

“About two thirds of the diagram cannot be shown.” An area on a chosen plane. Across twelve published coordinate systems the sRGB triangle covers between 8.5 and 38.4 per cent of the visible region — a factor of 4.5, with the observer and the display fixed throughout. This site has said it since its first commit.

“This diagram is more uniform than that one.” A comparison of two members of the family, which is a meaningful engineering statement and not a statement about vision. It is invariant only when phrased as a floor.

“This is what a protanope sees.” The Brettel construction projects along a direction that is a property of the cone matrix, and the matrix this site uses puts the deuteranope’s confusion point 1.28 away in chromaticity from the measured one. Changing to the matrix built from the confusion points moves the simulated palette by up to 3.9 ΔE00.

“This ΔE00 is 2.4.” CIELAB divides by a white point and takes a cube root, and neither operation commutes with a change of basis. The same twenty-five ellipses have a mean axis ratio between 2.59 and 9.84 depending on which basis the formula is built on.

Two that are harder than they look

Two of the nine deserve more than a line, because their classification is not obvious and getting it wrong is the commonest way this audit goes astray.

“The ellipses cannot be made circles” is invariant and “these ellipses are round” is not. The first quantifies over the whole family and is therefore a property of the measurement; the second is a statement about one member. The same underlying data support both sentences and only one of them is about the eye. A universally quantified statement over the group is always invariant, which sounds like a trick and is the single most useful move available — it is how a floor on uniformity turns a comparison of diagrams into a fact about discrimination.

“This colour is inside the gamut” is invariant and “this colour is 0.08 outside it” is not. Membership is an incidence and survives; the distance by which something misses is a metric quantity measured in whatever space the miss was computed in. This site computes gamut misses in linear RGB and reports them in figures, which makes them a statement about a display’s own coordinates — defensible, since the question is about that display, and worth saying rather than assuming.

The pattern in both is the same: a predicate survives and a magnitude attached to the same predicate need not. That is the quickest way to sort a sentence without computing anything.

What the two piles have in common

The division is not arbitrary and it is not about importance. Read down the surviving list and every item is one of two things.

An equality or an incidence. Matching, membership, collinearity — all statements of the form this is that or this is inside that, and all of them survive because a nonsingular map preserves equalities and preserves the incidence structure of subspaces.

Or a count of physical things. A number of surfaces, a number of lights. Counting is invariant because each item’s membership is.

Read down the failing list and every item involves one of two things.

A measure on a chosen plane. Areas, distances, ratios of regions — none of them is projective.

Or a nonlinearity applied after the choice. A cube root, a projection along a direction, a division by a white point. The freedom is harmless while everything downstream is linear and becomes a decision at the first nonlinearity, which is the sharpest way to state the whole thing.

What was computed, and how

The invariant claims are checked by applying random nonsingular matrices and measuring a residual, with the number reported rather than a verdict — an exact statement is best drawn as a residual with its exponent visible.

The failing claims are measured as a spread over published coordinate systems rather than over random matrices, and the choice matters. An arbitrary matrix can be made nearly singular and any of these quantities pushed anywhere, which would be an argument about degenerate cases. Twelve systems the discipline has printed is an argument about practice.

The audit itself is curated. Nine sentences chosen because they are things this site actually says, in roughly the words it says them in. A longer list would sort the same way, since the sorting rule turns out to be structural, but it would be less legible and no more convincing.

Where each published matrix puts the confusion points, whether or not it meant to. Every matrix from tristimulus values to cone responses commits itself to three confusion points, because the point is the direction the other two rows annihilate. The first row is the construction from the measured points and returns them exactly. The rest were chosen for other reasons and land elsewhere — Hunt–Pointer–Estévez, which this collection uses everywhere, misses the deuteranope's point by 1.28 in chromaticity. The worst here is 4.09.
Fig. 6 The measurement behind the third failure. Five matrices in use, and where each one’s implied confusion points fall against the measured ones.

Where the model stops

Invariance is not correctness. Every one of the five surviving claims could be wrong for other reasons — a bad measurement, a false premise, an arithmetic slip. What invariance buys is that the claim is about the observer rather than about the representation, which is a precondition for being right about the observer and nothing more.

And non-invariance is not error. Each of the four failures is a true statement about a stated convention. The failure mode is the reading rather than the claim, and the repair is a phrase — on the CIE xy diagram, under the Hunt–Pointer–Estévez fundamentals, in CIELAB — attached where the convention enters.

Nor is the list of nine a random sample of anything. It was assembled by reading the site’s own figure captions and section headings for sentences that make a quantitative claim, and it therefore over-represents the claims this collection is proud of. A sample drawn from ordinary prose rather than from headline sentences would probably sort in the same proportions — the sorting rule is structural — and nobody has checked that, and it is the kind of thing that is easy to be wrong about.

The group is also not the only freedom. Individual observers differ, field size changes the functions, and rod intrusion changes the dimension. Those move the subspace the matches span, which is a measurement changing rather than coordinates changing, and the audit here holds all of them fixed.

Who found it, and when

The habit of quotienting by a symmetry group to find out what a theory actually says is the twentieth century’s most portable idea, and it arrived in physics rather than in colour science. Its application here is not a discovery; it is an import.

What is worth recording is the local history. The 1931 committee knew XYZ was a convention and said so. The knowledge survives in the standards, which name their observer and their coordinates carefully, and it does not survive in the derived literature — a display specification quotes a percentage of a diagram’s area, and by the third citation the diagram’s name has gone.

This site is not an exception to that and the audit is the evidence: two of the four failing sentences are sentences it has published, one of them since its first commit and one of them in a hundred and eighty figures.

The mechanism by which the caveat is lost is worth naming, because it is not carelessness. A convention that everybody shares is invisible: every textbook drew the CIE diagram, every laboratory used the same fundamentals, every difference was computed in CIELAB — so no comparison anybody made was ever affected, and a variable that never varies is indistinguishable from a constant. It is the same mechanism that kept a measurement condition out of a whiteness figure until instruments with different ultraviolet content became common, and the same one that will keep this one hidden for as long as CIELAB is the only space anybody quotes.

The generalisation

The procedure is three steps and none of them is difficult.

Name the group the data leave free. Here it is nonsingular 3×3 matrices on the matching functions; elsewhere it is a gauge, a reparametrisation, a choice of units, or the ordering of a set.

Apply it and measure. A residual of machine zero is an invariant claim; a spread is a conventional one, and the size of the spread is how much the convention is worth.

And for each conventional claim, look for an invariant replacement. Two of the four here have one — a count of stimuli replaces an area, a floor over the group replaces a comparison — and the replacements are not harder to compute. The other two do not, and the honest response is a named convention rather than a repair.

The reason this is worth doing on a body of work rather than on a single claim is that the results cluster. The failures all shared a structure — a measure on a plane, or a nonlinearity — and once that was visible, sorting the rest of the collection stopped requiring any computation at all.

What each fitted thing in these essays carries, what its data fix, and what is left. Three columns per row: how many numbers the model has, how many the stated data determine, and the difference — the dimension of the family that fits equally well. The third column is the one nobody publishes. A zero there does not mean the model is right; it means it is determined, which is a much weaker property and is compatible with being determined badly, as the camera row is.
Fig. 7 The other half of the same audit, on fitted objects rather than on sentences. The third column there and the failing pile here are the same thing seen from two directions.

What this site will do about it

The four failures need four different things and none of them needs a retraction.

The diagram fraction gets a name and a count. The hatching stays, because which cells are hatched is invariant; the sentence beside it names CIE xy and quotes the invariant version alongside — 92.1 per cent of this collection’s constructed surfaces are reachable and none of the monochromatic lights is.

The uniformity comparisons get a floor. The best plane there is leaves a mean axis ratio of 2.02, and quoting that beside any comparison converts a ranking into a statement about the eye.

The simulations name their matrix, as the site’s fifth figure rule already requires them to name their model and severity. The rule was written for exactly this shape of over-claim and did not reach far enough: Brettel–Viénot–Mollon at severity 1 is incomplete without the cone matrix the projection was taken along.

And the differences name their space, which they already do — every ΔE00 on this site says so — with the added observation that the space carries a basis and the basis was not chosen for difference.

Three of the four are a phrase. The fourth is a number that was already being computed.

The two largest failures swap once the axis ratios are corrected

The audit ranks its four failures implicitly by what the convention is worth, and the fourth one’s figure is out of date.

The spread quoted for the difference formula is between 2.59 and 9.84, a factor of 3.80, against the diagram fraction’s 4.52 — so as written the diagram is the more conventional of the two. But 2.59 and 9.84 are the mean axis ratios as a forty-eight-point sample reported them, and that sample understates a ratio by an amount that grows with the ratio. The closed form gives 2.60 and 12.45, a factor of 4.79.

Corrected, the two swap. The difference formula’s basis is worth 4.79 and the diagram is worth 4.52, and the audit’s most-conventional claim is the one every specification in the industry is written in rather than the one every textbook prints.

Neither number changes the classification, since both are far from one, and the correction moves in the direction that strengthens the essay’s case rather than weakening it. What it changes is which of the four repairs to spend effort on: the phrase that names a space is doing more work than the phrase that names a diagram.

Machine zero, meant literally

The audit’s residuals are worth pricing, because machine zero is a phrase this collection uses for two things seven orders apart and here it is the strict one.

Double precision resolves about 2.2 × 10⁻¹⁶. The matching residual over sixty-four bases is 2.8 × 10⁻¹⁵ — thirteen times that; the collinearity residual over forty is 9.8 × 10⁻¹⁵, forty-four times. Both are what a handful of arithmetic operations on a 3×3 accumulate, and neither leaves room for anything else.

That is the right standard for an invariance claim and it is worth having explicitly, because elsewhere in this collection a 10⁻⁷ result is offered as arithmetic noise, and 10⁻⁷ is four hundred and fifty million times double precision. Both are far below any threshold that means anything; only one of them is machine zero. An invariance is a theorem and should read like one, and these two do.

A count is a magnitude and survives anyway

The quick sorting rule — a predicate survives and a magnitude attached to the same predicate need not — sorts eight of the nine and misclassifies the fourth survivor.

This many surfaces are reachable is a magnitude. It is a number, not an incidence, and by the rule as stated it should be at risk. It is not, and the reason is worth writing into the rule: a count is built entirely out of predicates, so it inherits their invariance term by term. Nothing is measured; each surface contributes a one or a zero and each of those is invariant on its own.

The refined rule sorts all nine and is barely longer. A magnitude survives if it is assembled from invariant predicates and fails if it is assembled by measuring. An area measures; a distance measures; a cube root of a coordinate measures. A count does not, and neither does a supremum over the group, which is why the ellipses cannot be made circles survives while any particular ellipse’s roundness does not.

That also says why the two available repairs are the two they are. Both replacements substitute a counting or quantifying operation for a measuring one — a count of stimuli for an area, a floor over the group for a comparison of two members — and there is no third such substitution on offer for the remaining two failures, because a simulation and a difference are both irreducibly measurements in a chosen space.

And the count’s invariance is narrower than it looks. It is invariant under the group and not under the population: 92.1 per cent is a fact about this collection’s own constructed surfaces, and a different surface set gives a different number by the same kind of margin the diagram gave. So the repair converts a hidden convention into a stated one rather than removing a convention, which is an improvement in honesty and not in universality.

Why the exercise is worth repeating elsewhere

The audit took an afternoon and it produced two things a longer argument would not have.

The first is a sorting rule that needs no computation. Once the two piles were visible, every further sentence could be classified by inspection: a predicate or a count survives, and a measure on a plane or a quantity computed after a nonlinearity does not. That is a rule anybody can apply to their own field’s sentences without implementing anything.

The second is a list of repairs, three of which turned out to be a phrase and one a number that was already being computed. A criticism whose repair is expensive gets filed; a criticism whose repair is a clause in a caption gets made.

What the exercise cannot do is tell anybody which claims matter. Four sentences here fail the test and two of them are load-bearing — the diagram fraction is quoted constantly, the difference formula’s units are in every specification — while the other two are mostly rhetorical. Invariance sorts claims by what they are about and not by what they are for, and deciding which of the failures to spend effort on is a judgement the test does not make.

Where the ladder goes next

The one place the freedom becomes a decision rather than a nuisance is the first nonlinearity in any pipeline, and the appearance models make that decision in the open: the axes an appearance model adapts in are chosen, fitted, and named after receptors they do not match.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 16 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AssertionChromaticity planeCIELABColour vision deficiencyCorresponding coloursDiagram conventionsGamutIdentifiabilityInvarianceProjective transformationSpecificationStandard observer