Difference and uniformity

A tolerance is a probability

A colorimeter reports one number and a specification compares it with another, and both are computed for an observer who does not exist. Handed to two hundred people, the same pair at ΔE00 1.0 is read from 0.8 to 3.7 — and which of those two ranges applies depends on something no specification records.

Assumes A tolerance is a shape and Whose eyes.

A tolerance is written as a property of a pair of samples, and a tolerance is a shape before it is a number. Below ΔE00 1.0 — two objects, one number, a verdict. Every part of that sentence is about the things being measured and no part of it is about anybody looking.

There is a person in it, and the person is a seventeen-observer average from 1931. Replace them with two hundred people whose eyes differ in the five ways eyes are known to differ, and the verdict stops being a property of the pair.

One tolerance decision, at four spectral distances. Every column is a pair of samples ΔE00 1.0 apart for the observer a colorimeter models — solved to that value by bisection, so the instrument would report the same number for all four. What differs is how far apart the two spectra are, which is achieved by adding a metameric black the reference observer cannot see. The bands are what two hundred people report: from 1.13 at the ninety-fifth percentile when the spectra are the same shape to 2.32 when they are not, and the worst case reaches 3.7. No specification records the quantity on the horizontal axis.
Fig. 1 Four pairs of samples, every one of them ΔE00 1.0 apart for the observer a colorimeter models — solved to that value, so the instrument reports the same number four times. What differs is how far apart the two spectra are. The bands are what two hundred people report, and the rightmost worst case is 3.7.

The claim

A tolerance is met with a probability, and the probability depends on a quantity nobody records.

  • For two samples of the same colorant at slightly different strengths, the tolerance is nearly a fact about the pair. The population reads a reference difference of 1.00 as 1.01 at the median and 1.13 at the ninety-fifth percentile — a spread of about fifteen per cent, which is smaller than the instrument-to-instrument spread most quality systems already live with.
  • For two samples made of different colorants, it is a fact about the person. At the largest spectral distance this construction reaches, the same reference difference of 1.00 is read at 1.30 at the median, 2.26 at the ninety-fifth percentile and 3.68 at worst. The ninety-fifth percentile has doubled and nothing about the colorimetric difference has changed.
  • And the extreme case has no tolerance at all. A three-primary display driven to match a broad white exactly — residual 4 × 10⁻¹⁶ — is seen as a different colour by half the population at 9.3 units and by the worst-off at 20.

The three cases are the same measurement at three spectral distances. The specification cannot tell them apart, because it records only the number the instrument produced.

Two of the three numbers are a bias and only one is a spread

The claim block reads as though the whole effect were disagreement — a population spreading out around a verdict the instrument got roughly right. Half of it is not disagreement at all.

The population’s median reading of the furthest pair is 1.30 against a reference of 1.00. That is not two hundred people scattered about the instrument’s answer; it is two hundred people agreeing with each other that the difference is thirty per cent larger than the instrument says. The flat pair’s median is 1.01, so the bias is essentially zero there and grows with spectral distance exactly as the spread does.

Splitting the far end’s excursion into its two parts, in logarithms: the ninety-fifth percentile sits 2.26 times the reference, of which 32 per cent of the excursion is bias and 68 per cent is spread.

And the bias is not an accident of this population — it is forced by the construction. A metameric pair is metameric for the reference observer, by definition: that observer is the one for whom the two spectra are most nearly identical, because the metameric black was projected into their null space and nobody else’s. Every other set of sensitivities sees some of it. So the reference observer sits at a minimum of the population’s readings rather than in the middle of them, and every member can only report more.

That predicts the flat pair’s bias should be zero, because no metameric black was added and there is no null space being exploited. It is one per cent. The prediction and the measurement agree, which is what says the mechanism is right rather than the number being a coincidence of this population.

The distinction matters because the two halves have different repairs. A bias is correctable: if the standard observer systematically reads a spectrally distant pair thirty per cent low, a specification could tighten its tolerance by thirty per cent for such pairs, and the correction needs only the pair — not a population model, not a cohort, not anybody’s eyes. A spread is not correctable by anything. No tolerance can be set that is right for a member at 1.30 and a member at 3.68 at once; the only response is to accept a stated failure rate or to stop making acceptance decisions on metameric pairs by instrument.

So the missing field the essay asks for buys more than a classification. A spectral distance beside a ΔE00 would let the bias be removed and would leave only the part that genuinely cannot be. That is a better proposition than the essay makes for it, and it is available from the same measurement: the instrument already holds both curves, the bias is a function of how far apart they are, and calibrating that function is a one-off experiment rather than a per-decision one.

It also sharpens what the standard observer is being asked to be. For an ordinary pair it is an average, and an average is the right thing to compute a difference against. For a metameric pair it is not an average of anything — it is the extreme member, the one observer for whom the pair is a match, and using it as the reference guarantees an answer that is low for everybody. That is a much more specific complaint than the standard observer is nobody, and unlike that complaint it names which direction the error runs in — the same asymmetry that makes a match fail for somebody else rather than for everybody equally.

Why an ordinary pair is safe

The reassuring half of this result deserves stating first, because it is the case most quality control is actually in and it is the reason the whole system works as well as it does.

Two samples of the same ink at slightly different film thicknesses have reflectance curves of the same shape, differing by a scale. Every observer integrates both curves against their own sensitivities; the sensitivities differ between observers; but the two integrals differ from each other in nearly the same proportion for everybody, because the difference between the two spectra is spread across the band in the same way the sensitivities are.

Put another way: observer variation is largely a common-mode error in a difference measurement, and a difference measurement rejects common-mode errors. The lens absorbs blue for both samples. The macular pigment attenuates both. What survives is the part of the observer’s variation that interacts with the part of the spectrum where the two samples differ — and if they differ smoothly and slightly, that is not much.

This is why a press operator, a colorimeter and a client can agree at all, and it is why a fifteen per cent spread is the right answer for the case the standard was written around — an ordinary difference, in the coordinates the formula was fitted in.

A box tolerance and a ΔE tolerance around the same colourA slice through CIELAB at L* 50, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 2.1 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 74% — accepted by one specification and rejected by the other, on the same measurement.gold outline: ΔE2000 = 1blue box: ±1 in each componentelongation 2.13disagree on 74%of everything either acceptsΔa*at L* 50, a* 55, b* 25 — the shape changes elsewheresame measurement, two verdictsCIE 1931 2° observer
Fig. 2 The tolerance as a specification writes it: a shape in colour space around a target, with a verdict inside and outside. Everything about this picture is a property of the two samples. The observer is in it, invisibly, in the coordinate system it is drawn in.

Why a distant pair is not

Now let the two samples be made of different things. A print against a display proof. A dyed fabric against a pigmented plastic. A four-ink build against a spot ink. Two spectra that agree on three integrals and disagree everywhere in between.

The disagreement in between is exactly what an observer’s own variation multiplies. Change the sensitivities and the three integrals move, and they move by amounts that depend on where the two spectra parted company — which is now everywhere.

The measurement in the hero figure isolates that cleanly, because it changes the spectral distance while holding the colorimetric difference fixed. A metameric black is added to one sample’s reflectance. A metameric black is a spectral difference the reference observer cannot see at all, by construction: adding any amount of it leaves the tristimulus values unchanged. So the pair’s colorimetric difference is untouched, a smooth tilt is then solved by bisection to bring that difference back to exactly 1.00, and the only thing the sweep varies is how much of the spectrum the two samples disagree about.

The ninety-fifth percentile doubles across the range available. The sweep stops where it does because the reflectance clamps at zero — past that point the black is larger than the reflectance it is being added to and the pair stops being metameric — and rows whose reference difference has drifted are dropped rather than quietly reported.

Two different spectra that are the same colourTwo reflectance curves differing by 93 per cent RMS, and the two patches they produce under D65: identical to ΔE00 = 1.8e-13, which is arithmetic noise rather than a small number. Both patches are inside the sRGB gamut, so neither has been clipped into agreement.400450500550600650700wavelength / nmthe two coloursΔE00 = 1.8e-13spectra differ by 93%reflectances, under D65CIE 1931 2° observer
Fig. 3 The construction the sweep is built on. Two spectra with the same tristimulus values under a stated observer, differing by an amount that is invisible to that observer by definition — and visible to a different one. The pair here is the extreme case; the sweep uses fractions of it.

What was computed, and how

The population is two hundred members from a stated seed, each a set of three cone absorptances at that member’s own lens age, macular density, outer-segment density and three pigment peaks. The five spreads are quoted from the individual-observer literature and every one of them was already an argument to the site’s own retinal machinery with a default nobody had varied.

Each member is adapted to their own view of D65 by CAT16 before any difference is taken, which is the comparison a person makes. The cruder route — handing CIELAB the member’s own white point — moves the numbers by four per cent, and that the two agree is asserted rather than assumed.

Nothing is quoted against the standard observer. This population’s median member differs from the CIE 1931 observer by ΔE00 0.95 on average over twenty-four natural reflectances — about what two degrees or ten costs, which is itself the size of several effects measured on this site — so every result here is a spread within the population, which is invariant to which member is called the reference.

And the acceptance fraction is a count. The share of members whose reported difference exceeds the tolerance, out of four hundred. For the flat pair it is 60 per cent, for the furthest 72 — but the fractions are much less informative than the percentiles, because a reference difference sitting exactly on the tolerance puts half the population either side of it whatever the spread.

Where the disagreement comes from, one variate at a time. The same measurement with one source of variation live and the other three held at their medians. The largest is the lens — entered as an age, because that is what it is a function of — at 7.7 ΔE00 against lens density, entered as age 20–70. So the biggest single reason two people disagree about a colour is the one thing about an observer that is written on their passport, and no colour specification has a field for it. The shares do not sum to the whole and are not expected to: the variates enter a nonlinear function of the spectrum, so this is a ranking rather than a decomposition.
Fig. 4 Where the disagreement comes from, one variate at a time, on the extreme case. The largest term is the lens, which is a function of age — so the biggest single reason two people disagree about a colour is the one thing about an observer that is written on their passport, and no specification has a field for it.

The quantity nobody records

The whole essay reduces to one practical statement: whether a ΔE00 transfers from an instrument to a person depends on the spectral distance between the two samples, and a specification records the ΔE00 and not the distance.

Colour science has a name for the missing quantity — a metamerism index — and a standard way to compute one, since two spectra with one colour are not equally metameric. It is used when the question is explicitly about metamerism, in textiles and in car interiors where trim from different suppliers has to match. It is not part of an ordinary tolerance, and an ordinary tolerance is where most acceptance decisions are made.

The repair is not obviously expensive. A spectrophotometer already measures the whole reflectance curve of both samples and throws most of it away to produce three numbers. The root-mean-square difference between the two curves is already in the instrument’s memory when the ΔE00 is printed.

What that number would buy is a classification rather than a correction. It cannot say how much any particular person will disagree, because that needs the person. What it can say is which of two regimes the decision is in — the one where the instrument’s verdict transfers to almost everybody, and the one where it transfers to nobody in particular — and those two regimes want different tolerances, different sample-approval procedures and, in the fragile case, a human sign-off that a colorimeter cannot replace.

There is a second reason the field is missing that is worth naming, because it is not carelessness. A metamerism index is a property of a pair, and most tolerances are written against a single reference standard which suppliers are expected to match with whatever colorant they have. So the index depends on the supplier’s formulation and cannot be written into the specification before the supplier exists — which is exactly when specifications get written. The quantity is unavailable at the moment it is needed and available at the moment it is not, and that asymmetry has kept it out of the ordinary process for fifty years.

Where the model stops

A distribution over observers is not a distribution over judgements. Two people with identical eyes disagree about whether a difference is acceptable, because acceptability is a decision and not a measurement — and which of two is worse is the same question asked of one person twice — so the spread in that is larger than anything here. This essay computes only the part attributable to optics and photopigments — the part that would remain if everybody used the same criterion perfectly.

The variates are drawn independently and they are not independent. Lens density and macular pigment are both age-related; drawing them independently makes people who do not exist. The spread is therefore right in order of magnitude and wrong in its tails, in a direction nobody can compute without a covariance that has not been published.

And the sweep is one reflectance and one metameric black. The factor of two in the ninety-fifth percentile is a property of that pair’s construction as much as of the principle. What is not a property of the construction is the direction: adding spectral distance at fixed colorimetric difference makes the population disagree more, every time, and the flat case is the floor.

The generalisation

The sentence worth carrying is: every number this site computes is a statistic of a population, and the ones that survive the population are the ones taken as differences between similar things.

That is a general property of measurement rather than a fact about colour, and it explains a pattern across this whole site. Metameric pairs are fragile. Narrow primaries are fragile. A display matched to a print is fragile. A press sheet matched to yesterday’s press sheet is not. In every case the robust comparisons are between spectra of the same shape and the fragile ones are between spectra that agree only on their integrals.

The surprising connection is with the observer index. This site’s own gate requires each figure to name its observer, and the reason given was that a chromaticity is a picture of an integral whose kernel somebody chose. This essay is the same argument with a consequence attached: the kernel is not merely chosen, it is distributed, and the width of the distribution decides whether the picture means anything to the person looking at it.

The same white, matched at six primary widths. At every width the three primaries are solved to match D65 exactly for the reference member; the bands are what the population sees. A broad primary integrates the observer differences over a band and averages them away; a narrow one samples them at a point and passes them straight through. From 40 nm to 2 the ninety-fifth percentile rises from 11.3 to 17.9 ΔE00, monotonically, and the technology has been moving from left to right for thirty years.
Fig. 5 The technological version of the same trend. Narrower primaries mean a larger spectral distance between a display’s white and any broad white it is matched to — which is what a wider gamut costs arriving from the observer’s side, and the population’s disagreement grows monotonically with it — from 11.3 units at the ninety-fifth percentile to 17.9.

Who found it, and when

Observer metamerism was a known consequence of the standard observer from the day it was adopted. The 1931 functions are explicitly an average, the paper says so, and the variance was not published because seventeen observers is not enough to estimate one.

The CIE published a special metamerism index for observer differences in 1989, based on a hypothetical “standard deviate observer” — one artificial second observer standing in for the whole population. It is a good instrument for a rank ordering and it is not a distribution, and this essay’s numbers are not comparable with its.

The individual observer models arrived after the 2006 physiological observer, and they are what made a population computable from published numbers rather than from a cohort: five parameters, five spreads, and a construction that turns a draw into a set of colour-matching functions.

And the industrial pressure came from displays. Textile and paint metamerism has been managed for a century by simply requiring samples to be spectrally close. A laser projector cannot be spectrally close to anything, so the problem arrived in a form nobody could specify their way out of — which is when the numbers started being quoted on shop floors rather than in journals.

What the pictures cannot show

Nothing here draws what another observer sees, and nothing on this site ever can: a patch on this page is a stimulus, chosen by this site’s median observer, delivered to the reader’s own eyes. A figure claiming to show the other person’s view would be a figure about the reader’s eyes wearing somebody else’s label.

And the reader cannot be located in the distribution. Four of the five variates need an instrument nobody has. The fifth is the age, which is the largest term and which every reader knows — so the honest thing the figures can offer is that a reader over sixty and a reader under twenty-five are near opposite ends of the largest single source of spread in every band drawn here.

One exact match, handed to two hundred people. Three primaries 2 nm wide are solved so that their sum has exactly the tristimulus values of D65 for the reference member of the population — three equations, three unknowns, residual 4.3e-16 of the white's luminance. The histogram is what everybody else sees: a median of 9.3 ΔE00, a ninety-fifth percentile of 17.9, and a worst case of 20.1. Two members picked at random disagree with each other by 4.6 units at the median. Nothing about the two spectra changed between the reader of this caption and the next one.
Fig. 6 The extreme case as a distribution rather than a band: one exact match, two hundred people, a median of nine units. A tolerance is a line drawn across a picture like this one, and where it lands depends on how far apart the two spectra were.
A fourth primary, swept — every setting an exact match, none of them the same. Four primaries matching three numbers leave one degree of freedom. Along the horizontal axis it is the fourth primary's share of the white's luminance; at each value the other three powers are solved exactly, so every point on this plot is a floating-point-exact match for the reference member — worst residual 1.3e-15 — and no colorimeter can tell them apart. What the population sees runs from 13.7 ΔE00 at the ninety-fifth percentile to 17.3, a factor of 1.26. The best setting is the largest share the arithmetic admits, so what stops it is not colour but the requirement that four powers stay positive.
Fig. 7 What the same population says about a device with a fourth primary. The freedom a fourth emitter opens can be spent on making the match hold for more people, which is the one place in this phase where a measurement of observer spread comes with something to do about it.

Where the ladder goes next

The nearest unfinished piece is the missing field. A specification that carried a spectral distance beside its ΔE00 would separate the safe case from the fragile one, the number is already inside every spectrophotometer, and nobody would have to agree on a population model to use it — the distance is a property of the samples alone.

The second is the joint distribution. Observer variation and illuminant variation act on the same part of the spectral difference, so a pair fragile to one is fragile to the other, and the two effects should be strongly correlated. Nothing here has measured that correlation, and a shop that changes its lamps and its inspectors at the same time is exposed to both.

And the third is the one this essay had to declare out of scope: the spread in judgement. Acceptability has been measured, its spread between observers is large, and it interacts with the optical spread computed here in a way no model on this site can express — because everything here ends at a number and acceptability begins after one.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 20 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AcceptabilityCIEDE2000ΔEIlluminant metamerismIndividual variationMeasurement errorMetamerismObserver metamerismQuality controlSpecificationStandard observerTolerance