Matching and measuring

The ambiguity is largest where the index is used

The metamerism index is defined for a pair that matches exactly under the reference light, no real pair does, and the two corrections the standard allows for the residual give different answers. Over eighteen pairs at eleven match qualities the gap is nearly a function of the mismatch alone — its middle half spans a factor of 1.4 at the mismatches a dyehouse reaches — so it could be tabulated. And it is largest exactly there: zero at a laboratory's match, and re-grading eight pairs in eighteen at a trade's.

Assumes The metamerism index has two corrections, The index is a choice too and A tolerance is a probability.

The metamerism index has two corrections established the problem. The special index is the colour difference a metameric pair shows under a test light, it is defined for a pair matching exactly under the reference light, and no pair matches exactly — so the standard allows two corrections for the residual, and they give different answers. That essay priced the difference at two match qualities: a grade apart at two units of reference mismatch and a rounding at a tenth. It said the number a standards committee would want is the distribution of the gap over a library of real pairs at the match quality the trade actually achieves.

How often the choice of correction changes a pair's grade. Eighteen metameric pairs walked to each of eleven reference mismatches, with the share whose index falls in a different band under the two corrections the standard allows. At an exact match the share is zero and must be: there is nothing for either correction to correct. It rises to 44 per cent at a reference mismatch of 2, which is the quality a dyehouse reaches rather than the quality a laboratory constructs. The banding is a five-step convention at 0.5, 1, 2 and 3, stated here rather than quoted, and how much the count depends on it is drawn separately.
Fig. 1 Eighteen metameric pairs walked to each of eleven reference mismatches, with the share whose index falls in a different band under the two corrections.

Predictable, and worst where it is used

The gap between the two corrections is nearly a function of how well the pair matched and hardly at all of which pair it is — so it could be tabulated. And it is zero at the match quality the index was defined for and largest at the quality a trade reaches.

  • At a reference mismatch of two the choice of correction re-grades 8 of the 18 pairs. At a tenth of a unit it re-grades none.
  • The gap’s middle half spans a factor of 2.2 at a mismatch of one and 1.4 at three, so it narrows as it grows: the ambiguity becomes more predictable as it becomes larger.
  • The median gap is 0.113 at half a unit of mismatch and 0.836 at two, rising roughly in proportion.
  • How many pairs are re-graded is partly the banding. Four bandings give 0, 1, 3 and 4 of eighteen at a mismatch of one and 1, 8, 13 and 14 at two.
  • One correction reads higher on every one of the eighteen pairs, so the choice is about severity rather than a coin toss.

A library rather than a bracket

The earlier essay’s two numbers bracketed the answer and did not settle it, and the reason is that two numbers from one pair cannot say whether the pairs differ.

The library here is eighteen metameric pairs built from natural reflectances by the construction two spectra, one colour sets out, each walked by bisection to each of eleven reference mismatches spanning what a laboratory constructs and what a dyehouse achieves. That makes the gap a distribution at each mismatch rather than a value, and the shape of the distribution is the first thing worth knowing.

The gap between the two corrections, across the library. The difference between the two corrections' indices, at each reference mismatch, as a band from the smallest pair to the largest with the middle half filled and the median drawn. The band is narrow and narrows further: the middle half spans a factor of 5.2 at a mismatch of 0.1 and 1.4 at 3. So the gap is nearly a function of how well the pair matched and hardly at all of which pair it is — which means it could be tabulated, and a standard could say what the choice of correction is worth without naming a pair.
Fig. 2 The difference between the two corrections’ indices at each reference mismatch, as a band from the smallest pair to the largest with the middle half filled and the median drawn.

The band is narrow and it narrows. The middle half spans a factor of 3.6 at a mismatch of half a unit, 2.2 at one, 1.5 at two and 1.4 at three. So the gap is very nearly a function of the mismatch alone — which is the useful shape, and not the one a reader would expect from a quantity built out of two spectra.

That means it could be tabulated. A standard that could not decide between the two corrections could at least publish what the choice is worth: a table of the median gap against the reference mismatch, with the quartiles beside it, and a laboratory reading its own mismatch off its own measurement would know how much of its index is convention. Nothing about that requires agreeing on which correction is right.

The table would be short. Eleven rows of three numbers cover the whole range a trade works in, and the numbers themselves are simple — 0.113 at half a unit of mismatch, 0.304 at one, 0.836 at two, 1.474 at three, with the middle half within a factor of two of each. A laboratory that measured a reference mismatch of 1.5 and reported an index of 3.2 could then say that the figure carries about half a unit of convention in it, and that the convention makes the index larger rather than smaller.

That is a different kind of statement from a tolerance and a more honest one. A tolerance says how much variation a decision will accept; this says how much of the number being compared against it was decided by a committee rather than measured. Both belong in a report, and only the first is ever in one.

Why the gap grows with the mismatch at all

The distribution is narrow and rising, and the reason is worth setting out because it is what makes the tabulation possible.

Both corrections do the same job: they move the sample’s tristimulus values under the test light so that it would have matched the standard under the reference light. The multiplicative one scales each channel by the ratio the reference mismatch shows; the additive one adds the difference. On a pair matching exactly the ratio is one and the difference is zero, so both are the identity and both give the same index.

As the reference mismatch grows the two corrections part company, and they part company in proportion to it — a ratio and a difference agree to first order and differ at second. So the gap grows roughly linearly in the mismatch at small mismatches and the growth is nearly the same for every pair, because the second-order term depends on the tristimulus values’ own sizes rather than on the pair’s spectra.

That is why the distribution narrows as it grows. At a small mismatch the gap is a small second-order quantity and the pairs’ individual differences are a large share of it; at a large mismatch the second-order term dominates and the pairs look alike. The spread is largest where the quantity is smallest, which is the usual arrangement and is the opposite of alarming.

The test light decides what the choice of correction is worthThe median gap between the two corrections at a reference mismatch of 2, under each test light the standard names. Under an incandescent lamp it is 0.836 colour differences and 8 of 18 pairs are re-graded; under an equal-energy source it is 0.023 and none are. The ambiguity scales with how far the test light sits from the reference, which is the same quantity the index itself is measuring — so the correction matters most exactly where the index is largest, and a specification that names a distant test light is naming a more ambiguous number.A8 of 18 re-graded0.8364D503 of 18 re-graded0.1472E2 of 18 re-graded0.0231median gap between the two corrections, ΔE₀₀at a mismatch of 2two corrections · CIE 1931 2°
Fig. 3 The same comparison of test lights at a reference mismatch of two rather than one. Every light’s gap is larger and the ordering between them is unchanged, which is what a term that scales with the mismatch does.

Zero where it was defined, largest where it is read

The share of pairs the choice re-grades is the form the ambiguity takes in practice, because an index is read as a grade rather than as a number.

At an exact match the share is zero, and it must be — the same null an exact match needs no correction reports: with nothing to correct, both corrections are the identity. That is the null the rest of the curve is read against and it is worth having as a computed result rather than as an argument, because it is the one place the two corrections are required to agree.

At a tenth and a fifth of a unit the share is still zero. Those are the mismatches a laboratory constructing a pair for a standard can reach, and they are where the index’s own validation would have been done.

At a reference mismatch of two the share is 8 of 18. That is the quality a dyehouse achieves on a difficult shade, and it is where the index is actually computed — so the choice of correction moves nearly half the library across a grade boundary exactly where the number is being used to accept or reject work.

That ordering is the finding. The convention is invisible where it was validated and decisive where it is applied, which is the shape of defect that survives a standard’s own checks indefinitely. Which index to buy an instrument for is the neighbouring case of a decision made where the consequences are not.

How much of that is the banding

A grade is a band, and a band has boundaries somebody chose. Before treating a count of re-graded pairs as a fact about the index, it is worth moving the boundaries.

How much of the grade count is the banding. Four bandings of the same indices, at two reference mismatches. The count of re-graded pairs runs 1, 3, 0, 4 of 18 at a mismatch of one and 8, 1, 14, 13 at two. So the count is a statement about the convention as well as about the indices — a tighter banding at the small mismatch re-grades three pairs where a wider one re-grades none. What survives every banding is the comparison between the two columns: at the mismatch a trade reaches, three of the four bandings move most of the library.
Fig. 4 Four bandings of the same indices, at two reference mismatches.

At a mismatch of one the four bandings give 0, 1, 3 and 4 of eighteen. So the count at the small mismatch is a statement about the convention as much as about the indices: a tighter banding finds three pairs re-graded where a wider one finds none.

At a mismatch of two they give 1, 8, 13 and 14. Three of the four move most of the library, and the exception is the tightest banding — which re-grades fewest because its boundaries are all below where these indices sit.

So the honest statement separates two claims. The gap between the corrections is a property of the indices and is what the distribution above measures. How many grades it changes is a property of the indices and the banding, and the number that survives every banding is the comparison between the two mismatches rather than either count alone.

The banding used throughout is a five-step convention at 0.5, 1, 2 and 3, stated here rather than quoted. A reader whose own standard bands differently can read the sensitivity off the figure directly.

And how much is the test light

The other thing the number depends on is which lamp the index is computed under, and it depends on it strongly.

The test light decides what the choice of correction is worthThe median gap between the two corrections at a reference mismatch of 1, under each test light the standard names. Under an incandescent lamp it is 0.304 colour differences and 1 of 18 pairs are re-graded; under an equal-energy source it is 0.009 and none are. The ambiguity scales with how far the test light sits from the reference, which is the same quantity the index itself is measuring — so the correction matters most exactly where the index is largest, and a specification that names a distant test light is naming a more ambiguous number.A1 of 18 re-graded0.3041D504 of 18 re-graded0.0654E0 of 18 re-graded0.0092median gap between the two corrections, ΔE₀₀at a mismatch of 1two corrections · CIE 1931 2°
Fig. 5 The median gap between the two corrections at a reference mismatch of one, under each test light the standard names.

Under an incandescent lamp the median gap is 0.304 and one pair in eighteen is re-graded. Under a daylight test light it is 0.065 and four are. Under an equal-energy source it is 0.009 and none is.

The scaling is the interesting part: the gap grows with how far the test light sits from the reference. That is the same quantity the index itself is measuring — a metameric pair comes apart in proportion to how differently the two lights weight its spectra — so the correction matters most exactly where the index is largest.

A specification naming a distant test light is therefore naming a more ambiguous number — and the lamp in the shop decides is why the distant light is the one that matters — and not by a small factor: thirty times between an equal-energy source and an incandescent one. The index is a choice too established that the test light is a decision rather than a given; this adds that the decision carries a second one with it.

The ambiguity has a direction

One more thing can be said about the gap and it is the thing a committee could act on.

One correction reads higher than the other on every pairEach of the 18 pairs at a reference mismatch of 1, with both corrections' indices. The additive correction reads higher on 18 of them and the multiplicative on 0 — means of 3.301 and 3.013. So choosing between them is not a coin toss whose errors cancel over a range: it is a choice about severity, and a committee naming one would be tightening or loosening the whole scale by a known amount. That is the useful form of the ambiguity, because a known bias can be absorbed into a limit and a spread cannot.additivemultiplicative#140.335#150.397#130.244#60.304#70.358#50.230#120.160#40.157#110.149#30.131#100.217#170.430#20.170#90.340#160.484#10.251#80.447#00.369the index, ΔE₀₀18 of 18 one waytwo corrections · CIE 1931 2°
Fig. 6 Each of the eighteen pairs at a reference mismatch of one, with both corrections’ indices.

The additive correction reads higher than the multiplicative one on every one of the eighteen pairs — means of 3.301 against 3.013. Not most, not usually: all of them.

That turns the ambiguity from a spread into a bias, and a bias is a much easier object. A known bias can be absorbed into a limit — a specification that names the multiplicative correction and sets its limit lower by the tabulated gap reaches the same decisions as one that names the additive correction — while a spread cannot be absorbed into anything. The offset is not a constant, since it grows with the mismatch, but it is a function of a number the laboratory already has, which is the next best thing.

So the practical form of the recommendation is narrower than “the standard should decide”. It is: the standard should decide, and until it does, an index quoted without naming its correction is quoted with a known offset in an unknown direction. A tolerance is a probability is the neighbouring argument about what an unstated term does to an acceptance decision, and the term here is one word.

One correction reads higher than the other on every pairEach of the 18 pairs at a reference mismatch of 2, with both corrections' indices. The additive correction reads higher on 18 of them and the multiplicative on 0 — means of 4.079 and 3.240. So choosing between them is not a coin toss whose errors cancel over a range: it is a choice about severity, and a committee naming one would be tightening or loosening the whole scale by a known amount. That is the useful form of the ambiguity, because a known bias can be absorbed into a limit and a spread cannot.additivemultiplicative#150.975#140.861#130.705#60.836#70.945#50.708#120.567#40.591#110.566#171.162#30.567#161.271#100.699#20.657#90.960#81.197#10.809#01.035the index, ΔE₀₀18 of 18 one waytwo corrections · CIE 1931 2°
Fig. 7 The same eighteen pairs at a reference mismatch of two, where every index is larger and the additive correction still reads higher on every one. The means are 4.079 and 3.240 against 3.301 and 3.013 at a mismatch of one, so the gaps have roughly tripled and the ordering has not changed on a single pair.

The consistency across the two mismatches is what makes the bias worth acting on rather than merely worth noting. A difference that held on eighteen pairs at one match quality and reversed on some of them at another would be a spread with a trend in it; a difference that holds on every pair at every quality is a systematic offset, and a systematic offset is a number that can be written into a limit once and forgotten.

How the library was built and walked

The pairs are metameric partners constructed from natural reflectances by adding a metameric black, which is the construction used throughout here; a base whose partner would leave the physical range is skipped rather than clipped. Each pair is then walked to a stated reference mismatch by bisecting the amplitude of a fixed smooth ripple, so that the mismatch is a size rather than a worst case and the direction is the same for every pair.

The index is the colour difference between the standard and the sample under the test light, with the sample corrected by each of the two rules the standard allows. The gap is the absolute difference between the two, and the grade is the band the index falls in under the stated convention.

The three test lights are the ones the standard names. Everything is computed under the CIE 1931 observer, which is the observer the index is defined for; an index under another observer is a different quantity and the index is one observer’s opinion is the essay about that.

What this leaves out

Eighteen pairs built one way are not a shade library. The construction adds a metameric black of a fixed shape, so the pairs differ in their bases and not in how they were made metameric — and a dyehouse’s pairs come apart because two recipes were used, which is a different family with its own distribution. The prediction that the gap is nearly a function of the mismatch is the claim most likely to change with a real library, and it is the claim most worth checking.

The walk is one direction in spectrum space. A pair walked away from its match along a different direction reaches the same mismatch with a different spectral difference, and whether the gap depends on the direction as well as on the size is a computation this does not make.

And the grading is a convention. The sensitivity to it is measured and reported, and a reader’s own standard may band differently enough that the counts do not transfer; the distribution of the gap does.

Still open: what a real shade library’s mismatches actually are

The whole of this essay’s practical weight rests on where a trade sits on the horizontal axis, and that is a fact about dyehouses rather than about colorimetry. The measurement is a survey rather than a computation: the reference-light colour difference of a few hundred accepted pairs, from the records a laboratory already keeps.

The prediction worth recording is that the distribution is wide and centred somewhere between half a unit and two. If it is centred below half a unit the ambiguity is a rounding and the standard’s silence costs nothing; if it is centred near two it re-grades half the pairs; and if it is wide — which the way acceptance limits are set makes likely — then the same standard is unambiguous for the easy shades and decisive for the hard ones, which is the worst arrangement, because the hard shades are the ones an argument is had about.

The same records would settle a second question for nothing. The gap’s near-dependence on the mismatch alone is the claim that would let a committee tabulate it, and a real library either confirms it or shows the spread that this constructed one does not have.

A convention is invisible where it was validated

The habit is about where to look for the cost of an unstated choice.

A standard’s optional clauses are tested where the standard was written, and a standard is written around the cases its authors could construct: exact matches, controlled samples, clean conditions. An option that changes nothing there passes, and it passes honestly — at those conditions the two branches agree.

The move is to evaluate the option across the range of conditions the thing is used in rather than the range it was checked in, and to plot the difference against whatever parameter separates the two. Here that parameter is one number a laboratory already records, and the plot runs from zero to half the library.

The failure mode is to test an option at the conditions that made it optional. The two corrections exist because a pair does not match exactly; checking whether they differ on a pair that nearly does is checking them at the one place they cannot.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

The objects this essay names

Each one links to every other essay that touches it.

AcceptabilityColour differenceConventionIlluminant metamerismMeasurement conditionMetamerismMetamerism indexQuality controlSpecificationTolerance