What it takes to deliver it

The instrument is one observer exactly

A spectrophotometer has no observer uncertainty. It computes through the 1931 tables to the last bit, reproducibly, forever — which is its whole value and is also why it cannot report the three colour differences of observer spread between itself and whoever is looking at the sample.

Assumes A tolerance with an observer in it, An instrument brings its own light and What the instrument reports.

An instrument’s relationship to the standard observer is unlike anything else in its uncertainty budget, and the difference is worth stating precisely because it explains why the observer term never appears.

The conditions under which an observer's departure is exactly zero. A departure of the observer is the pairing of something belonging to the observer with something belonging to the stimulus, so emptying either factor empties the product. The axis is logarithmic in what is left when the condition is imposed. Six rows empty the stimulus's factor — a perfectly neutral sample is the same colour for every observer, at any age and any field size — and two empty the observer's, since a gain on each cone and a change of basis are both absorbed exactly. All eight are identities rather than small numbers. The last two are the same two conditions imposed in a published cone space rather than in the observer's own, and they are worth eight and thirteen units: the identity is about the eye, and the arithmetic everybody uses is in somebody else's coordinates.
Fig. 1 The conditions under which an observer departure vanishes. An instrument satisfies none of them and has no departure anyway, because it is not a member of the population.

The claim

A spectrophotometer does not approximate the standard observer; it computes through it exactly, and that exactness is why the gap between it and a human viewer is invisible from inside the measurement.

  • The observer is not a source of instrument uncertainty. Every other term in the budget is a measurement; this one is a table lookup and a summation.
  • Two instruments agree about the observer perfectly, so inter-instrument agreement — which is measured, standardised and reported — carries none of it.
  • The gap it cannot see is about three ΔE₀₀ between its number and what a viewer’s own eye would compute on a saturated sample.
  • And an instrument could report the gap, because it measures the spectrum and the gap is computable from the spectrum. Nothing does.

What an instrument’s uncertainty is made of

A spectrophotometer’s reported colour has a chain behind it and every link but one is a measurement.

The spectral measurement carries wavelength accuracy, photometric accuracy, bandpass, stray light and noise. All of them are calibrated against standards, all have published methods, and all appear in the instrument’s specification.

The geometry decides which part of the sample’s return reaches the detector, and two standard geometries disagree by units on a glossy sample. That is a real uncertainty and it is stated as a condition rather than a tolerance.

The light is the instrument’s own, supplied rather than assumed, and its spectrum enters the calculation as a declared illuminant rather than as the lamp actually used.

And then the observer, which is not measured at all. The instrument multiplies its measured reflectance by a tabulated illuminant and three tabulated functions and sums. Two instruments doing that arrive at identical numbers from identical spectra, to the last bit.

One link in the chain has no uncertainty because it is arithmetic, and it happens to be the link where the largest gap to the viewer lives.

Perfect agreement is the point. That the observer contributes nothing to instrument disagreement is not a flaw; it is the reason a colour specification can exist at all.

Two laboratories on two continents measure the same sheet and get numbers agreeing to a few tenths. Almost all of that agreement is because they share three tables, and none of it would survive if each instrument used its own estimate of a real observer. A standard observer is a contract, and a contract’s value is that it is identical for everybody.

So the exactness is load-bearing and nothing here suggests changing it. What it costs is that the instrument’s number is a statement about the contract rather than about any viewer, and the instrument cannot tell the difference because from inside the calculation there is none.

That is the same shape as the identity a neutral sample has: a place where the arithmetic is exact is a place where nothing can be learned, and both the instrument’s exactness and the neutral’s identity produce systems that agree with themselves and are silent about the thing they were built to predict.

What an instrument already knows

The useful part of the argument is that the missing information is already in the machine.

An observer departure is computable from a spectrum — the sample’s reflectance and the illuminant — and a spectrophotometer measures the first and is told the second. The excitation deviation of the sample from the illuminant is one integral over data the instrument already has, and multiplying it by a coefficient gives the observer allowance.

So the instrument could print, beside its ΔE₀₀, a second number: this sample carries an observer spread of about 2.8 units under the stated illuminant. It would cost nothing to compute, it would use no additional measurement, and it would turn a term nobody budgets for into a number on a report.

The reason nothing does it is that there is no standard for what to print. The coefficient depends on a population model, the population models in the literature disagree, and an instrument manufacturer printing a number with no standard behind it is taking a position rather than reporting a measurement.

The obstacle is standardisation rather than capability, which is a different problem and a more tractable one.

Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.
Fig. 2 The six departures over forty-two surfaces. Every one is computable from a measured spectrum and a stated illuminant, which is what an instrument already has.
A departure against how far the sample sits from the light. The sample is mixed with a flat reflectance, from the flat one at the left to its own at the right, and two observers differing in the macular pigment look at each mixture. The straight line is the distance between their relative cone excitations, and it is straight to 0.0 per cent: the departure is a pairing, and scaling one factor scales the product. The curved line is the same sequence in ΔE₀₀, which is not a linear function of the excitations and cannot be — it has cube roots in it and a chroma weighting underneath. The identity is about the eye; the curvature belongs to the unit.
Fig. 3 The departure against how far the sample sits from the light. An instrument measures exactly the quantity on the horizontal axis and reports nothing about the vertical one.

That figure is the essay’s practical proposal drawn. The horizontal axis is computed from a reflectance and an illuminant, both of which a spectrophotometer has; the vertical axis is the number nobody prints. The line between them is a coefficient.

It is worth being clear about how little machinery the proposal needs. No population model is required to compute the horizontal axis, no new measurement, no change to the instrument’s optics. Only the coefficient is contentious, and a conservative one would still be far better than the zero currently implied.

What a standard deviate observer was for

The CIE did publish something for this in 1989 and it is worth saying what it does and does not supply.

The standard deviate observer is a single hypothetical set of colour-matching functions representing a typical deviation from the standard observer. An index computed with it — the observer metamerism index — reports how far a matched pair comes apart when the deviate observer is substituted for the standard one.

Two things about it limit its use here. It is about a pair rather than a single sample, so it answers “will this metameric match survive a different observer” rather than “how much observer spread does this colour carry”. And it is one deviation rather than a distribution, so it gives a representative number rather than a percentile.

Both are the right design for the question it was built for, which is colorant formulation, where the samples are pairs by construction. Neither answers the tolerance question, where the sample is a single delivery against a numerical standard.

That is why the term is missing from budgets despite a standard existing: the standard measures a different quantity, and using it for this one would be an error of a familiar kind.

The one place instruments do carry an observer difference

There is an exception and it is instructive because it is the case where the observer stops being arithmetic.

A tristimulus colorimeter — as distinct from a spectrophotometer — has three physical filters intended to approximate the colour-matching functions, and the approximation is imperfect. Its observer is a piece of glass rather than a table, so it has manufacturing tolerance, ageing and temperature dependence, and two such instruments genuinely disagree about the observer.

That disagreement is exactly the Luther condition failing, and this collection has an essay about when it would work: a filter set can be a colorimeter only if its sensitivities are a linear transform of the observer’s, and no real filter set is.

The interesting part is the direction of the trade. A colorimeter has an observer uncertainty and a spectrophotometer does not, and it is the spectrophotometer that is silent about the viewer — because the colorimeter’s uncertainty is about its own filters rather than about anybody’s eye. Having an uncertainty about the observer is not the same as knowing anything about observers.

What this means for inter-instrument agreement

Inter-instrument agreement is the term the industry works hardest on and the round has something specific to say about it.

Two spectrophotometers agreeing to 0.2 ΔE₀₀ on a tile set is a considerable engineering achievement and it is achievement about the spectral measurement, since the observer contributes exactly zero to their difference. Improving it further improves the spectral chain and nothing else.

Meanwhile the gap to a viewer is about three units and is not improved by any of it. So there is a point past which better inter-instrument agreement buys nothing for the decision the measurement supports, and by the numbers here that point is well past where the industry already is.

That is an uncomfortable thing to say about a large and careful body of metrological work, and it is not an argument against it — instrument agreement matters for process control, for arbitration between supplier and customer, and for the traceability the whole system rests on. It is an argument about which term to work on next, and the answer is not the one currently getting the attention.

Every departure under every light. Six departures across six lights, each cell the difference between two observers in ΔE₀₀, drawn as a bar whose length is the number. The rows are not multiples of one another: the lens is worst under tungsten and the pigment peaks are worst under a three-emitter LED, because a departure is a pairing and which light is being paired with decides it. The laser projector's row is empty, and that is not a fact about lasers — on this collection's five-nanometre grid a three-line spectrum is a one-line spectrum, and a single wavelength is a stimulus every observer agrees about exactly.
Fig. 4 The six departures under six lights. An instrument computes the same number under all of them, because its illuminant is a table rather than the lamp in the room.

The illuminant is a table too

The same argument applies once more, one column across, and it is worth following because the two compound.

An instrument reports a colour under a declared illuminant — D65, D50, A — computed from tabulated spectral power. The lamp in the room where the sample will be seen is not that spectrum, and there is no D65 lamp: the best daylight simulators have metamerism indices that put them a visible distance from the table.

So the instrument’s number is a colour under a table, computed for an observer that is a table, and both tables are contracts rather than descriptions. The gap to the viewer is the sum of the two, and the two are not independent — the observer departure roughly doubles under a warm lamp, so a specification checked under D65 and viewed under a shop’s lighting pays for both at once.

That compounding is the honest answer to why colour disputes are common in retail and rare in a laboratory. Both parties’ instruments agree; neither instrument is measuring the situation the dispute is about.

A report could carry two more columns. Sketching the output makes the proposal concrete and shows how small it is.

A colour report currently gives L*a*b* for the sample, L*a*b* for the standard, a ΔE₀₀ and a pass or fail against a tolerance. Adding two columns would give the observer allowance for this sample under this illuminant, and the ΔE₀₀ expressed as a fraction of the tolerance plus the allowance.

On a near-neutral the second number would be barely different from the first, which is correct — there is nothing to allow for on a neutral. On a saturated pigment it would be substantially different, and the difference is the information currently absent.

What that would change in practice is the conversation when a delivery is disputed. At present the instrument says 0.8 units and the customer says it does not match, and there is nothing to say. With the second column the instrument says 0.8 units with an observer allowance of 2.9, which does not settle the dispute and does explain it.

A number that explains a disagreement is worth having even when it does not resolve one, and this one costs an integral.

What was computed, and how

Nothing new is computed here. The three-unit figure is the round’s quadrature combination of six departure medians over forty-two surfaces, excluding the rod term for a photopic assessment, and it carries the caveats that essay states — chiefly that two of the six are correlated so the true combination is larger.

The claim that an instrument’s observer contributes exactly zero to inter-instrument disagreement is a statement about arithmetic rather than a measurement: two summations over the same tabulated functions with the same spectral input give bit-identical results, and any difference is in the spectral input.

One more asymmetry is worth recording because it decides who can act. The instrument manufacturer can compute the allowance and cannot standardise it; the standards body can standardise it and does not build instruments; and the user can do neither and is the one who bears the consequence.

That distribution of capability and responsibility is common wherever a missing term needs a convention before it can be reported, and it is why such terms stay missing for a long time after they are understood. A number nobody is authorised to print does not get printed, however easy it is to compute, and the authorisation is the scarce thing rather than the arithmetic.

Where the model stops

The essay treats a spectrophotometer as computing exactly through the standard tables, which is true of the calculation and glosses over one real effect: instruments interpolate and abridge, and a bandpass and a reporting interval are choices that differ between makes. That is a spectral term rather than an observer term and it belongs in the first half of the chain.

It also treats the viewer as a member of a population characterised by six parameters. A real viewer has all six at once with correlations, and may additionally have a colour vision deficiency, which is a categorical difference rather than a spread.

What choosing a space to divide the white out in is worth. Three pairs of routes to the same colour, over forty-two surfaces: dividing the white out in tristimulus values, in a published cone space, and in the observer's own cones. The first two agree to 0.59 ΔE₀₀ at the median. Either of them differs from the observer's own cones by more than fifteen. That is why the two exact conditions in this round are exact only in the eye's own coordinates: the identity belongs to the receptors, and every published arithmetic works in a basis somebody else chose.
Fig. 5 Three routes to the same colour. An instrument computes one of them exactly and the others equally exactly, and which one it computes is a setting rather than a measurement.

There is a second exact link in the same chain and it is worth naming beside the observer. An instrument’s conversion from tristimulus values to a colour difference goes through a chromatic adaptation transform, and the choice of transform is worth fifteen colour differences against the physiology and about half a unit between the two published options.

That choice is a setting on the instrument. It has no uncertainty, it is computed exactly, and two instruments set differently disagree by a real amount that no calibration reduces. So the observer is not the only definition in the chain, and the general form of the essay’s point is that a measurement chain accumulates definitions as well as errors, and the definitions are invisible in the budget by construction.

The generalisation

The habit is about what an exact link in a chain hides.

A chain of estimates has an uncertainty at every link, and the total is dominated by the worst. A chain with one exact link has no uncertainty there at all, which is an unambiguous improvement — and it also means that no amount of examining the chain’s uncertainties will reveal anything about that link’s correspondence to the world.

An exact link is a definition, and a definition cannot be wrong. What it can be is a definition of something other than what the chain is for, and detecting that requires stepping outside the chain and comparing against the thing it was meant to predict.

The failure mode is to read a small total uncertainty as evidence that a system predicts well. The uncertainty budget of a system with a definition in it is a budget about everything except the definition, and the definition is often where the interesting error is.

A last observation about what an instrument is for. A spectrophotometer’s real product is not a colour; it is a spectrum, and everything colorimetric it prints is a projection of that spectrum through two tables. A user who keeps the spectrum keeps everything, and can recompute under any observer, any illuminant and any adaptation route later; a user who keeps the ΔE₀₀ has kept one projection and thrown the rest away.

That is an argument for archiving spectra rather than colours, and it is stronger than the usual one about future illuminants. Every departure in this round is recoverable from a stored spectrum and none of them is recoverable from a stored colour, so a laboratory that stores spectra has already paid for whatever the next audit finds.

Who found it, and when

The CIE standardised the colorimetric calculation in 1931 and the summation has been the definition of a colour ever since. The standard deviate observer arrived in 1989 with a technical report specifically about assessing observer metamerism in colorant formulation, and a further report on observer metamerism for displays followed in 2016 as narrowband primaries reached production.

The absence of an observer term in colour tolerance uncertainty budgets is not argued against anywhere this collection can find; the budgets are measurement budgets and the term is not a measurement uncertainty. Whether it belongs somewhere else is a question the standards do not take a position on.

Where the ladder goes next

Two places carry the term more than anywhere else and both are hardware decisions. The first is the display, where narrow primaries buy a disagreement that no calibration removes.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationIndividual variationInter-instrument agreementMeasurement conditionMeasurement uncertaintyObserver metamerismQuality controlSpecificationSpectrophotometryStandard observer