Difference and uniformity

A tolerance with an observer in it

A delivery tolerance is written in ΔE₀₀ against the 1931 observer, and six departures of that observer combine to about three of the same units on an ordinary saturated sample. A one-unit tolerance is being asked to contain a three-unit uncertainty that nothing in its budget mentions.

Assumes An observer is a contract, A delivery tolerance is three tolerances and A tolerance is a probability.

The observer audit produced six numbers in the same unit every colour tolerance is written in. Putting them beside a tolerance is the obvious next step and it has not been taken anywhere this collection can find.

Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.
Fig. 1 Each observer departure over forty-two surfaces. Every one of these is in ΔE₀₀ — the same unit a delivery tolerance is written in, and larger than most of them.

The claim

A colour tolerance is a budget, the observer is a line in it, and the line is missing.

  • Six departures combine to about 3.0 ΔE₀₀ at the median over ordinary surfaces and 8.2 at the ninety-fifth percentile.
  • A delivery tolerance is typically 1.0 to 2.0 ΔE₀₀, so the omitted term is larger than the whole budget.
  • The term scales with the sample, from 0.9 units on the palest members of the family to 8.2 on the most saturated, so a constant reserve is wrong in both directions.
  • And the omission is structural rather than an oversight. An uncertainty budget lists the terms that have standards attached, and the observer does not have one.

What a tolerance budget contains

A colour tolerance in industrial use is a single number — a maximum acceptable ΔE₀₀ between a delivered sample and a standard — and behind it is an uncertainty budget that decides what the number can be.

A delivery tolerance is three tolerances established the usual contents: the instrument’s repeatability, the inter-instrument agreement between the two parties’ devices, the sample’s own non-uniformity, and the process variation the specification is trying to control. Each is measurable, each has a standard method, and the tolerance has to be large enough to contain all of them plus the real variation it is there to bound.

Typical sizes: an instrument’s short-term repeatability is around 0.02 ΔE₀₀, inter-instrument agreement around 0.2 to 0.4, sample non-uniformity anything from 0.1 to 1.0 depending on the material. Those add to a few tenths, and a tolerance of one unit has most of itself left for the process.

The observer term is not in that list, and at three units it is larger than everything else combined.

The term is missing for a structural reason. The omission is not carelessness and understanding why is more useful than deploring it.

An uncertainty budget is assembled from terms that have standard methods of measurement. Instrument repeatability has one. Inter-instrument agreement has one, with published tile sets and procedures. Sample non-uniformity has one.

The observer has no standard method, because a colour specification’s observer is not measured at all — it is a contract, specified as the 1931 tables and computed rather than measured. There is nothing to put an uncertainty on, in the sense the budget’s other terms are uncertain.

So the term is missing because the budget is a budget of measurement uncertainty and the observer’s contribution is not a measurement uncertainty. It is the difference between what the contract computes and what the person will see, which is a different quantity that nobody’s budget has a slot for.

That distinction is real and it does not make the number go away. A specification exists so that a person accepts a delivery, and a term of three units between the number and the person is a term whichever budget it belongs in.

The size, stated properly

The six departures at their literature strengths, over forty-two analytic surfaces under a 6500 K radiator, at the median and the ninety-fifth percentile:

departure median 95th
the age of the lens 1.943 5.752
the macular pigment 1.610 4.081
the pigment peaks 1.392 2.401
the rods 1.028 1.748
the field size 0.740 2.477
the cone optical density 0.566 1.707
combined in quadrature 3.02 8.16

The quadrature combination assumes independence, and two of the six are correlated — the lens and the macular pigment both absorb in the blue. On these forty-two banded surfaces their adapted effects point the same way often enough that the true combination is somewhat larger than 3.02. That is a property of the surfaces rather than of the two filters: on smooth reflectances the same pair points nearly opposite ways and costs about half its quadrature, so the direction of the correction depends on the samples a tolerance is applied to.

The rod term should be excluded for a specification assessed at photopic levels, which most are, taking the median to about 2.84. That is the honest figure for a viewing booth: just under three colour differences of observer uncertainty on an ordinary saturated sample.

Why a constant reserve is wrong twice

The obvious repair is to add three units to every tolerance, and it is wrong in both directions.

Over the surface family the combined term runs from about 0.9 ΔE₀₀ on the flattest members to 8.2 at the ninety-fifth percentile — a factor of nine. A constant reserve of three is nearly four times too large for a pale sample and nearly three times too small for a saturated one.

That is not a distribution to be summarised. It is a function of the sample, and the function is known: the departure is exactly proportional to how far the sample sits from the light it is seen under, measured as a distance between relative cone excitations.

So the correct form of an observer allowance is not a constant but a coefficient times a quantity computed from the sample’s own reflectance and the illuminant. That is one integral, it uses data a specification already has, and it produces a reserve that is right across the range rather than at one point.

A departure against how far the sample sits from the light. The sample is mixed with a flat reflectance, from the flat one at the left to its own at the right, and two observers differing in the macular pigment look at each mixture. The straight line is the distance between their relative cone excitations, and it is straight to 0.0 per cent: the departure is a pairing, and scaling one factor scales the product. The curved line is the same sequence in ΔE₀₀, which is not a linear function of the excitations and cannot be — it has cube roots in it and a chroma weighting underneath. The identity is about the eye; the curvature belongs to the unit.
Fig. 2 The departure against how far the sample sits from the light. The straight line is what a sample-dependent allowance would be computed from.

What such an allowance would look like

The recipe is short enough to state.

Compute the sample’s excitation deviation from the illuminant: the Euclidean distance between its relative cone excitations and (1, 1, 1), in any consistent cone space. That is a number between zero for a neutral and about 0.6 for a saturated pigment.

Multiply by the observer coefficient for the viewing condition. From the measurements here, about 5 ΔE₀₀ per unit of excitation deviation under a smooth daylight source, rising to about 8 under tungsten and about 10 under a three-emitter LED.

And that is the reserve, to be combined with the instrument and process terms in the usual way.

Applied to the examples: a white paper reserves about 0.4 units, a pale tint about 1, an ordinary saturated pigment about 3, and a strong display primary rather more. Those are numbers a specification could live with, and they are very different from one another in a way a single figure cannot express.

The coefficients above are from this collection’s construction and would need re-deriving against the published observers before anybody used them. What is not construction-dependent is the shape: a linear allowance in a computable property of the sample.

Two conditions make the allowance vanish. An allowance that scales with something has values where it is zero, and those are worth naming because they are exactly the samples a specification is usually validated on.

A neutral reserves nothing. Every observer computes the same colour for a spectrally flat sample, exactly, so a grey card, a white tile and a black trap carry no observer uncertainty at all.

A sample under a light it nearly matches reserves little. The deviation is from the adapting white, so a warm sample under a warm light is closer to its illuminant than the same sample under daylight, and the allowance falls accordingly.

The first of those is the awkward one. A tolerance system validated on neutral references is validated on the samples where the missing term is zero, which is why the term’s absence has never shown up as an inconsistency. The samples that would reveal the omission are the ones nobody uses to check the system.

The conditions under which an observer's departure is exactly zero. A departure of the observer is the pairing of something belonging to the observer with something belonging to the stimulus, so emptying either factor empties the product. The axis is logarithmic in what is left when the condition is imposed. Six rows empty the stimulus's factor — a perfectly neutral sample is the same colour for every observer, at any age and any field size — and two empty the observer's, since a gain on each cone and a change of basis are both absorbed exactly. All eight are identities rather than small numbers. The last two are the same two conditions imposed in a published cone space rather than in the observer's own, and they are worth eight and thirteen units: the identity is about the eye, and the arithmetic everybody uses is in somebody else's coordinates.
Fig. 3 The ten conditions under which an observer departure is exactly zero. A tolerance system validated on the first six is validated where the term it omits is identically absent.

That figure is the reason the omission has survived. A tolerance system is checked for internal consistency by measuring known references, and the known references are white tiles, grey scales and neutral standards — every one of which sits on the identity. The system agrees with itself perfectly on them and would agree just as perfectly with the observer term included, since the term is zero there.

So the missing line has never produced an inconsistency anybody could find, and it never will from inside the system. The check that would reveal it is a comparison against a person on a saturated sample, which is not a check any tolerance system performs and is exactly the situation in which the complaints arise.

Where the term is already being paid

The observer term is missing from budgets and it is not missing from experience, and the industries where it bites are identifiable.

Automotive refinish, textile dyeing and plastics matched to painted trim all have the same structure: two samples made of different colorants, matched under one condition, and assessed by people rather than instruments. Those are metameric pairs by construction, and a metameric match is exactly what an observer difference breaks.

In all three the reported failure mode is a customer rejecting a delivery that the instrument passed, and the usual attributions are the illuminant, the geometry or the instrument. The observer is on the list of candidates in the literature and it is not in anybody’s budget, so it cannot be quantified when it happens.

What the audit adds is that it is quantifiable: three units at the median, eight at the ninety-fifth percentile, scaling with a property of the sample that is one integral away.

Every departure under every light. Six departures across six lights, each cell the difference between two observers in ΔE₀₀, drawn as a bar whose length is the number. The rows are not multiples of one another: the lens is worst under tungsten and the pigment peaks are worst under a three-emitter LED, because a departure is a pairing and which light is being paired with decides it. The laser projector's row is empty, and that is not a fact about lasers — on this collection's five-nanometre grid a three-line spectrum is a one-line spectrum, and a single wavelength is a stimulus every observer agrees about exactly.
Fig. 4 The six departures under six lights. The observer allowance is a function of the viewing condition as well as the sample, and warm lights roughly double it.

The viewing condition doubles it

The second argument of the allowance is the light, and it is not a small correction.

Under a tungsten lamp the six departures on a red pigment run from 0.79 to 4.47 ΔE₀₀ against 1.20 to 2.38 under daylight, and the combined term roughly doubles. Under a white LED it rises by about half.

That interacts badly with how specifications are written. A tolerance is stated once and a viewing condition is stated separately and often permissively — D65 or equivalent, a standard viewing booth. Two lamps of the same correlated colour temperature and different spectra can differ by a factor of two in the observer term, and correlated colour temperature is not a spectrum.

So an allowance that is right in a booth is wrong in a shop by a factor of two, in the direction that matters: the shop’s lamp decides whether a delivery is accepted, and it is usually warmer than the booth’s.

A specification has three options. A specification that wanted to carry the term has three options and they cost very different amounts.

State the term and leave it outside the tolerance. Add a sentence saying that observer variation contributes an additional uncertainty of roughly the sample’s excitation deviation times a stated coefficient, and that the tolerance does not cover it. That is free, it is honest, and it moves the risk to whoever reads the specification.

Widen the tolerance by the term. That is expensive: a one-unit tolerance becomes four on a saturated sample, which most processes could not meet and which would make the specification useless as a process control.

Or specify spectrally. Two samples that match band for band match for every observer, so a spectral tolerance has no observer term at all. It is much more demanding than a colorimetric one, it is what automotive and textile work has been moving towards for thirty years, and the reason is precisely this.

The third is the only one that removes the term rather than accounting for it, and it removes several others at the same time — a spectral match is also illuminant-independent and instrument-geometry-independent. What it costs is that it constrains the colorant rather than the colour, which is a much stronger requirement on a supplier.

Six departures of the observer, each at a stated strength, under a tungsten lamp at 2856 K. Each bar is two observers differing in one argument, looking at the same sample under the same light, in ΔE₀₀. The strengths are the literature's: the working-age lens, two standard deviations of the reported macular and density spreads, the long-wavelength polymorphism, the CIE's own second observer, and a rod contribution of a tenth. They are within a factor of 5.7 of one another, which is the point: there is no single term to fix. Every one of them is above the ΔE of about one that a delivery tolerance is written in.
Fig. 5 The six departures under a tungsten lamp. Every bar here is above a one-unit delivery tolerance and two of them are above four.

What was computed, and how

The six departures are this round’s, at their literature strengths, computed through this collection’s usual CAT16 route over the same forty-two analytic surfaces. Each is a difference between two observers rather than a spread across a population, so the quadrature combination is an estimate of what a population’s spread would be rather than a measurement of one.

That distinction matters and is the largest caveat here. A population statistic — what fraction of observers would reject a given sample — is a different quantity, and this collection has the machinery for it in its two-hundred-eye population. The two agree about direction and are not numerically comparable.

The coefficient of about 5 ΔE₀₀ per unit of excitation deviation is read off the linearity measurement rather than fitted, and it is a single figure standing in for six departures with different slopes.

There is a version of the same omission one level up that is worth naming, because it is the same argument applied to the round’s other two audits. A tolerance budget has no line for the wavelength grid either, and the grid is worth up to a colour difference under a fluorescent tube; it has no line for the room’s finish, and that is worth up to five.

Neither of those is an observer term and both are missing for the same reason: they are not measurement uncertainties, they have no standard method, and a budget assembled from standard methods cannot contain them. The three together are considerably larger than any tolerance in industrial use, which is a strange state of affairs for a system that works as well as this one does — and the resolution is in the conditions rather than in the sizes.

Where the model stops

Everything here is a matching uncertainty rather than an acceptability one. Whether a viewer rejects a sample at three units of observer-induced difference depends on the industry, the expectation and the surround, and acceptability is a probability rather than a threshold.

The departures are one parameter at a time. Real observers differ in all six at once with correlations the model does not carry, so the quadrature sum is an estimate with a known bias in one direction.

And the whole thing assumes the assessor is adapted to the light the sample is under. An assessor comparing a print in a booth while adapted to the room outside is in a different situation entirely, and that one is an appearance question rather than a matching one.

One further consequence for how a tolerance is agreed rather than computed. Two parties negotiating a tolerance are negotiating a number that both can measure, and both measure it with the same standard observer — so their instruments will agree to a few tenths whatever their eyes do. The negotiation therefore converges on a number that has nothing to do with either party’s ability to see the difference.

That is usually fine and it is the point of having a standard. It stops being fine when one party is a machine and the other is a customer’s eye, which is the arrangement at the end of every supply chain. A specification agreed between two instruments is an agreement about instruments, and the last step of the chain is not one.

The generalisation

The habit is about terms that are missing from a budget because they are the wrong kind of thing.

An uncertainty budget collects what can be measured with standard methods, and a term with no standard method is systematically absent — not because anybody decided it was small but because there was no line to write it on. The absence is invisible from inside the budget, which sums to a plausible total and contains everything it contains.

The test is to ask what the budget is for and whether its terms cover the gap between the measurement and the decision. Here the decision is a person accepting a delivery and the budget covers the instrument’s path to a number, which is most of the chain and not all of it.

The failure mode is to treat the budget’s completeness as established by its internal consistency. A budget that adds up is not a budget that is complete, and the missing terms are always the ones whose measurement has no standard.

Who found it, and when

Observer metamerism has been recognised since Wyszecki and Stiles and the CIE published a standard deviate observer in 1989 precisely to let it be quantified. That observer is a single hypothetical deviation rather than a population, and the index computed from it is a comparison between two samples rather than an allowance on one.

That the term does not appear in colour-tolerance uncertainty budgets is, so far as this collection can tell, simply true rather than argued anywhere. The budgets in ISO and ASTM practice for colour measurement are measurement budgets, and they are complete as measurement budgets.

Where the ladder goes next

An instrument does not have an observer uncertainty; it has an observer, exactly, and the difference between those two statements is the sharpest way to state what the audit found. An instrument is one observer exactly, and that is both its virtue and the whole of the problem.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 12 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AcceptabilityΔEIndividual variationMeasurement uncertaintyObserver metamerismQuality controlSpecificationStandard observerTest setTolerance