Matching and measuring

The metamerism index has two corrections

The index for a metameric pair is defined for a pair that matches exactly under the reference light, and no real pair does. The standard's remedy is to correct the sample first, and it names two corrections — scale the tristimulus values, or add the difference. On six pairs matched to one colour difference, the two answers differ by a tenth to four tenths of an index unit; at two, by a whole one.

Assumes A tolerance cannot cross a condition, One match names the observer and A tolerance is a shape.

Two samples with different spectra can be the same colour under one light and different colours under another. That is the site’s oldest picture, and the number the trade puts on it is the special metamerism index: the colour difference between the pair under a second, test light.

Two different spectra that are the same colourTwo reflectance curves differing by 92 per cent RMS, and the two patches they produce under D65: identical to ΔE00 = 6.2e-14, which is arithmetic noise rather than a small number. Both patches are inside the sRGB gamut, so neither has been clipped into agreement.400450500550600650700wavelength / nmthe two coloursΔE00 = 1.8e-13spectra differ by 37%reflectances, under D65CIE 1931 2° observer
Fig. 1 The object the index is about: two reflectances differing by a substantial fraction of their own height, identical in colour under D65 to the last digit the arithmetic carries.

The definition has a condition attached that is almost never met. It assumes the pair matches exactly under the reference light, and a dyed cloth against its standard, an ink against a proof or a plastic against a paint chip is matched to a fraction of a unit, never to nothing. Anything left over under the reference light is carried into the index, where it is indistinguishable from the metamerism the index is meant to report.

So the standard tells the user to correct the pair first — and names two corrections.

Two corrections, two answers

Which correction is applied is not a detail of arithmetic: at ordinary match quality it moves the index by a quarter of a unit, and at a loose match by a whole one.

  • At an exact reference match all three answers agree — uncorrected, multiplicative, additive — to floating point. That is the only case the definition covers and the check that the computation is right.
  • At a reference mismatch of one colour difference the two corrections differ by 0.13 to 0.37 index units across six pairs under an incandescent test light, with a median of 0.23.
  • At two units the median gap is 0.71 and the largest 1.03, which is the width of a whole grade on the scales the index is usually quoted in.
  • From half a unit of mismatch upwards the multiplicative correction reads lower on every pair. Below that the two cross, because they agree to first order and differ only in the second.

What the correction is for

A pair that fails to match under the reference light fails for a reason that has nothing to do with metamerism: it is simply not the same colour. Under the test light that residual is still there, mixed in with the part that is metamerism, and the index cannot tell them apart.

The two remedies both move the sample’s tristimulus values under the test light so that the pair would have matched under the reference one. The multiplicative correction scales each of the sample’s three values by the ratio the reference light demands. The additive correction adds the difference the reference light demands. Each is exact at the reference light by construction; they differ in what they imply about the test light, because a scale and an offset are the same to first order and diverge as the correction grows.

That is not a defect of either. It is the ordinary situation of a correction fitted at one condition and applied at another, which a tolerance cannot cross a condition met in its sharpest form: a least-squares matrix fitted between two measurement conditions left 23 per cent of the disagreement and made five patches of seventeen worse than doing nothing. Here the correction is exact where it is fitted. The question is what it does elsewhere.

The metamerism index, computed three ways, as the reference match loosens. The special metamerism index of 6 metameric pairs under an incandescent test light, against how well each pair matches under the reference light. Uncorrected, the index absorbs the reference mismatch and rises from 2.86 to 3.76. Corrected multiplicatively it rises to 3.31 and additively to 3.87. All three are the same number when the pair matches exactly, which is the only case the definition covers.
Fig. 2 The index of six metameric pairs under an incandescent test light, against how far each pair is from matching under the reference light. All three ways agree at the left edge, which is the only place the definition applies.

How the three diverge

At an exact match the three curves start together at a median index of 2.86. As the reference match loosens they separate, and they separate in a particular order.

The uncorrected index rises fastest — 2.86 at an exact match, 3.18 at one unit, 3.76 at two — because it is absorbing the reference mismatch wholesale. That is exactly what the correction exists to remove, and the size of the rise says how much the index would be overstating metamerism if nobody corrected anything.

The multiplicative correction rises least: 3.08 at one unit and 3.31 at two. The additive correction is close to the uncorrected number: 3.21 and 3.87. So the choice of correction is worth more than half of what the correction itself is worth — a strange position for a step whose purpose is to remove an artefact.

Six pairs matched to 1 colour difference, and the index each gets. Each pair is walked away from an exact match until it sits at ΔE₀₀ 1 under the reference light, and its index under incandescent is then computed three ways. The two corrections differ by 0.13 to 0.37 index units across the six, and the multiplicative one is the smaller on every one of them.
Fig. 3 Six pairs, each walked to a reference mismatch of exactly one colour difference, and the three indices each of them gets. The multiplicative correction is the lowest bar on every pair.

The per-pair picture is what makes the size legible. Pairs whose exact index is around two and a half get gaps of a tenth; pairs up around four get gaps approaching four tenths. The gap is roughly proportional to the index itself, which means the correction’s ambiguity is a percentage rather than a constant — about seven per cent of the index at a one-unit match.

Six pairs matched to 2 colour difference, and the index each gets. Each pair is walked away from an exact match until it sits at ΔE₀₀ 2 under the reference light, and its index under incandescent is then computed three ways. The two corrections differ by 0.57 to 1.03 index units across the six, and the multiplicative one is the smaller on every one of them.
Fig. 4 The same six pairs at a reference mismatch of two colour differences — a match a customer would query rather than accept. The gaps are now a fifth to a third of a unit apiece, and up to a whole one.

The test light decides how much it matters

A correction can only matter in proportion to how far the test light has moved the pair, and the test lights in use are not equally far away.

How far apart the two corrections are, under three test lights. The difference between the multiplicatively and additively corrected index, median over 6 metameric pairs, against how well the pair matches under the reference light. Under incandescent it reaches 0.71; Under a graphic-arts daylight it reaches 0.12; Under an equal-energy light it reaches 0.02 at a reference mismatch of two. A correction can only matter in proportion to how far the test light moves the pair, so the further the test light is from the reference, the more the choice of correction costs.
Fig. 5 The difference between the two corrected indices under three test lights, as the reference match loosens. Under an incandescent light it reaches seven tenths of a unit; under a near-daylight one, a twentieth.

Under an incandescent test light — the classic pairing with a daylight reference, and the one most metamerism specifications name — the gap at two units of mismatch is 0.71. Under a graphic-arts daylight it is 0.12, and under an equal-energy light 0.02. The correction’s ambiguity scales with the thing being measured, which is reassuring in one sense and not in another: a specification that pairs a daylight reference with an incandescent test is precisely the case where the choice of correction matters most, and it is also the most common one.

The case with a right answer

Nothing so far says which correction is better, because the pairs so far have no true index to be compared against — they are metameric, and how much of the test-light difference is metamerism and how much is the residual is exactly what is in dispute.

There is one family of pairs where the true answer is known, and it settles the question.

A sample that is its standard scaled, and the index each correction gives it. Six samples that are their standards multiplied by a constant — the same spectral shape, a few per cent lighter — scaled until each sits ΔE₀₀ 1 from its standard under the reference light. Such a pair is not metameric at any match quality: no light separates two reflectances of the same shape, so the index is zero and is known to be zero. The multiplicative correction returns that. The uncorrected index returns the reference mismatch itself, about one. The additive correction returns between 2.2 and 3.4 — metamerism that is not there.
Fig. 6 Six samples that are their own standards multiplied by a constant — the same spectral shape, a few per cent lighter — each matched to one colour difference under the reference light. Such a pair has no metamerism in it at all.

Take a sample that is its standard scaled: the same reflectance curve, uniformly a little lighter. That pair is not metameric at any match quality. Two reflectances of the same shape differ by the same factor under every light there is, so no illuminant separates them and the index is zero — known to be zero, before any measurement, from the shape alone.

Put those pairs through the three routes at a reference mismatch of one colour difference, and:

  • The multiplicative correction returns zero — between 2 × 10⁻¹⁴ and 3 × 10⁻¹³, which is the arithmetic’s own rounding.
  • The uncorrected index returns about 1.0, which is the reference mismatch itself, faithfully reported as metamerism.
  • The additive correction returns between 2.2 and 3.4 — two to three times the mismatch it was correcting, and all of it invented.

That last number is the argument. The additive correction does not merely disagree with the multiplicative one; on the one case where the answer is known it manufactures an index out of a pair that has no illuminant dependence whatsoever. The reason is algebraic: scaling a reflectance scales its tristimulus values by the same factor under every light, so a multiplicative correction removes it exactly, and an additive one removes it at the reference light and leaves a residue everywhere else that grows with how far the test light’s white is from the reference’s.

Which one to use, and what to write down

The practical answer is short and the reason for it is now not a matter of taste.

Use the multiplicative correction, because it is exact on the one family of pairs with a known answer, because it is the one the CIE’s own recommendation names first, and because it keeps the correction proportional to the signal, which is what a reflectance-based quantity should be. Then record which correction was used, because the number is not reproducible without that: two laboratories with the same spectra, the same illuminants and the same observer can publish index values a quarter of a unit apart, and both are correct.

And report the reference mismatch beside the index, which is the piece of information the correction is a response to. An index of 3.1 on a pair matched to 0.1 is a very different statement from an index of 3.1 on a pair matched to 2.0 — in the first case the correction changed nothing, and in the second it is doing more work than most of the terms in the budget. A tolerance needs a second number argued the same shape for delivery tolerances: the quantity every instrument already has and none of them prints is the one that makes the first number interpretable.

What the index is actually used for

The size of a quarter of a unit means nothing without knowing what decisions the index drives, and there are two.

The first is acceptance: a specification says a supplied colour must have an index below some figure under a stated test light, so that a garment matched in the showroom does not come apart under the shop’s lamps. There the ambiguity acts exactly like any other measurement uncertainty, and it is comparable in size to the uncertainties already priced elsewhere — a tolerance with an observer in it found six observer departures combining to about three units on an ordinary sample, and a tolerance is a probability found the same pair read from 0.8 to 3.7 across a population. The correction’s quarter of a unit is smaller than either, and unlike either it is entirely within the laboratory’s control.

The second is ranking: two candidate dye formulations are compared and the one with the lower index is chosen. Here the ambiguity is worse than its size suggests, because the gap is roughly proportional to the index — so two candidates whose true indices differ by less than seven per cent can be ordered either way by the choice of correction, and nothing in the report says which was used.

That is the same failure three constants nobody quotes found inside ΔE₀₀ itself: two settings of a parametric factor in ordinary use, twenty-nine per cent of acceptance decisions changing between them, and neither setting recorded in the number. The pattern is worth naming because it keeps recurring — a published figure with an unrecorded convention inside it, where the convention is worth about as much as the effect.

And it compounds with the choice of difference formula underneath, which is a third unrecorded decision: the weighting is the disagreement showed that whether a chroma difference is divided by the chroma it was measured at sorts the whole menu of formulae, and an index is a colour difference like any other. An index of 3.1 is a statement about a pair, a test light, a correction and a formula, and the usual report names one of the four.

What was computed, and how

Six metameric pairs are built on smooth natural reflectances by the usual construction: a metameric black, scaled as far as it will go while both members stay inside [0, 1], added to a base. Each pair matches exactly under D65 — to the arithmetic’s own rounding, which the first figure shows.

One member is then walked away from the match along a fixed smooth ripple, by bisection on its amplitude, until the pair sits at a stated ΔE₀₀ under the reference light: 0, a quarter, a half, one, one and a half, two. That is a mismatch of stated size rather than a worst case, and the direction is fixed rather than chosen per pair so that the six are comparable.

For each mismatch the index is computed under a test light three ways: with the sample’s tristimulus values left alone, scaled by the ratio the reference light demands, and offset by the difference it demands. The test lights are the standard incandescent illuminant, a graphic-arts daylight and an equal-energy light. Everything is ΔE₀₀ through the 1931 observer, each light against its own white point.

What this does not decide

Which correction is right in general is still not a question the arithmetic answers. The scaled-pair test settles one family, decisively, and a real mismatch is not usually a pure scale: it is part scale and part shape, and the shape part is where the two corrections have no known answer to be checked against. What the arithmetic settles is that they are not interchangeable, by how much, and that one of them fails a test the other passes.

The reference light is D65 throughout, and the choice is not neutral: a specification that names D50 as its reference and an incandescent light as its test has a different distance to cover, and therefore a different sensitivity to the correction. A delivery tolerance is three tolerances is the standing argument for writing down which conditions a number belongs to, and an index carries two of them rather than one.

The pairs here are constructed rather than measured. A real metameric pair — two dye formulations, an ink and its substitute — has a spectral difference with a shape of its own, and the reference mismatch is whatever the dyer achieved rather than a number set by bisection. The sizes would move; the structure, that the two corrections agree at zero and diverge in proportion to the residual, follows from their algebra and would not.

The walk is also a single direction in reflectance space. A mismatch of the same ΔE₀₀ taken in a different direction — one that moves lightness rather than chroma, say — would give a different gap, and the six pairs here happen to be walked the same way. That is a deliberate simplification, and it is the first thing to relax if anybody wants the distribution rather than the effect.

Still open: the correction’s own spread over real pairs

The number a standards committee would want is the distribution of this gap over a library of real metameric pairs at the match quality the trade actually achieves — a few hundred dyeings against their standards, with their measured spectra and their measured reference mismatches.

That would say whether the ambiguity is, in practice, a rounding difference or a grade difference. This essay says it is a grade difference at two units of mismatch and a rounding at a tenth, which brackets the answer without settling it; where the trade sits between those two is a fact about dyehouses rather than about colorimetry.

A definition with a condition nobody meets

The habit is about a quantity defined under an assumption that the world does not supply.

An index for pairs that match exactly, a gamut for a display with no room in it, a tolerance for two patches seen side by side and steadily: each is defined on a set of situations that has measure zero in practice, and each therefore needs a rule for what to do about the gap. The rule is usually not part of the definition, and it is usually not unique.

There is a reliable tell for which quantities are like this: the definition contains an if. An index for a pair that matches under the reference; a chromaticity for a stimulus that fills the field; a reflectance factor for a surface that is perfectly diffusing. The clause is what makes the quantity well defined, and the practical world almost never satisfies it, so what gets reported is the result of a repair rather than the quantity itself.

The move is to find the conditional in the definition and ask what happens when it fails. If the answer is a correction, ask who chose it, whether there is more than one in use, and how far apart they get — because that spread is part of the quantity’s uncertainty and is almost never quoted in it.

The failure mode is to treat the correction as bookkeeping. It is a second measurement decision, made after the samples are measured, and here it is worth a quarter of a unit on a number whose whole range is about five — and, on a pair with no metamerism in it at all, the difference between nothing and three.

The same move is worth making on every correction of this kind. A fit can be exact and empty asked what a fitted object’s residual conceals; this asks the narrower question of what a convention conceals, and the test is the same in shape: find the case where the answer is known independently, and see which candidate returns it.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CIEDE2000Colour differenceIlluminantIlluminant metamerismMeasurement conditionMetamerismQuality controlSpecificationStandard observerTolerance