An observer is a contract
Assumes The peaks move the flanks, Seventeen observers in 1931 and Which of these is a convention.
The round has now measured six ways in which no reader of this page is the observer every figure on it names. That is either a serious finding or a category error, and which one it is depends entirely on what a standard observer is for.
The claim
A standard observer is an agreement about what a number means, not a claim about anybody’s eyes, and the round’s six departures measure the cost of the agreement rather than an error in it.
- A contract cannot be wrong in the way a measurement can. It can be inconvenient, badly chosen or expensive, and it cannot fail to match a person it never claimed to describe.
- What the departures measure is the cost of the agreement — how far a real viewer will be from the number the contract produces, which is exactly what a specification needs and is not what any standard publishes.
- The 1931 observer’s known errors are a different matter. Its short-wavelength functions are wrong as a description of the mean of its own seventeen subjects, and that is a defect in the contract’s drafting rather than in its nature.
- And the reason it has never been replaced is the reason contracts are hard to change. Every specification, every instrument and every profile in the world refers to it, and the value of a shared reference is in the sharing.
What a standard observer does
A colour specification is a promise between two parties that a thing will be a certain colour. To be checkable, the promise has to reduce to a number that both parties can compute from a measurable quantity, and the reduction is what a standard observer supplies: a stated set of three weighting functions applied to a measured spectrum.
That is the whole of its job. It is an agreed weighting, and the agreement is what makes two instruments in two countries produce the same number for the same sheet of paper.
Read that way, several things that look like problems stop being problems. It does not matter that the seventeen subjects were young British men, because the contract does not claim they were representative. It does not matter that nobody’s eye matches the functions, because the functions are not a claim about eyes. And it does not matter that the two standard observers disagree, because they are two contracts and a specification names one.
This collection has an essay about which of its quantities are conventions and the standard observer belongs firmly on that list — closer to the metre than to the speed of light, an agreement that could have been made differently and is useful because it was made once.
What the departures are then a measurement of
If the contract cannot be wrong, the six numbers in this round need a different interpretation, and it is a sharper one.
They measure the gap between the contract and its purpose. The purpose of a colour specification is that a person will find the delivered thing acceptable; the mechanism is a number computed through an agreed observer; and the departures say how far the person can be from the number.
That gap is a real engineering quantity with a name — it is the observer’s contribution to the uncertainty budget of a colour specification — and it is the thing a tolerance has to cover alongside instrument repeatability, sample non-uniformity and illuminant variation.
What the round adds is its size and its structure. Six terms, between 0.57 and 1.94 ΔE₀₀ at the median over ordinary surfaces, combining to about three units, no one of them dominant, all of them scaling with the sample’s departure from the light. That is a budget line item, and it is larger than the instrument terms most specifications do budget for.
The one thing that is a defect
There is one part of the standard observer that is not merely a convention, and it is worth separating carefully from everything above.
The 1931 functions are known to misrepresent the short-wavelength behaviour of their own subjects. ȳ was constrained to equal the 1924 luminous efficiency function, and that function was too low in the blue by a factor approaching ten below about 460 nanometres — a fact established by the 1940s. Judd published a correction in 1951 and Vos refined it in 1978, and neither was adopted.
That is a defect in the contract’s drafting: the functions do not describe the data they were derived from, in a region where the discrepancy is large. A contract can be badly drafted, and this one is, in one identifiable place.
The reason it survives is instructive. Changing it would invalidate every colour specification written against it, every instrument calibrated to it and every profile computed with it, for the benefit of a correction that matters only for stimuli with substantial short-wavelength content. Until recently there were few such stimuli in commerce. Blue-pumped LEDs and narrowband displays have changed that, and the cost of the defect is rising.
The difference between those two readings is not rhetorical, and it changes what a reader should do with the chart. A demolition asks whether to keep using the standard, which has no available answer — there is nothing else in general use and the alternatives are worse-supported rather than better.
A price list asks how much to reserve, and that has an answer. On an ordinary saturated sample under an ordinary light, about three ΔE₀₀, rising to eight on the most saturated members of the surface family and falling below one on the palest. That is a number a specification can carry, and carrying it is a change of one line rather than a change of standard.
Why replacing a contract is harder than fixing a measurement
A better measurement supersedes a worse one as soon as it is published. A better contract does not, and the difference is what makes standardisation a different activity from science.
The value of a shared reference is superlinear in how many people share it. Two laboratories agreeing is worth something; the whole world agreeing is worth enormously more, because every pair of parties can transact. Replacing the reference means every pair has to switch at once or maintain both, and maintaining both is worse than either.
The CIE has published better observers. The 2006 physiological observer is parameterised by age and field size, it is built on the Stockman–Sharpe fundamentals, and it repairs the short-wavelength defect. It is not what industry uses, and the reason is not ignorance.
A standard’s inertia is a feature until it is a liability, and the transition point is when the cost of the defect exceeds the cost of switching. Nobody can compute either number, which is why such transitions happen late and abruptly.
What a specification should do instead
Given that the contract will not change, the practical question is what a specification can do about the gap, and there are three answers of increasing ambition.
Budget for it. Add an observer term to the uncertainty budget, sized from the round’s numbers and scaled by the sample’s departure from the illuminant. That is a small change to an existing calculation and no specification this collection has read makes it.
Choose conditions that shrink it. The departures are smallest under a smooth light and on samples near the illuminant, so a daylight-simulating booth and a specification written in terms a customer will actually view under are not the same choice — and the second matters more. The lamp in the shop decides, and it decides the observer disagreement too.
Or specify spectrally. A specification that states a reflectance rather than a colour has no observer in it at all, and two samples matching spectrally match for everybody. That is what a spectral tolerance would be, it is much more demanding than a colorimetric one, and it is used in exactly the industries where observer metamerism is expensive.
The third is the only one that removes the problem rather than budgeting for it, and it is the reason automotive and textile work has moved towards spectral matching over the last thirty years.
What the identities say about the contract
The conditions figure has a reading in this essay’s terms that is worth having, because it says what a perfect specification would look like.
Every departure vanishes on a sample that does not modulate its light. So a specification about neutrals is observer-free — and a specification about neutrals is a specification about almost nothing, since neutrals are the one colour everybody can already make.
Every departure vanishes when two observers differ by a gain. So a contract that only ever compared a sample against a white would be observer-free, which is what a densitometer does and is why densitometry survived so long as a press control: it measures a ratio in a way that is insensitive to a great deal, at the cost of not being about colour.
The identities therefore describe the boundary of what a colour specification can be free of, and the boundary is tight. Anything worth specifying is outside it, which is not a failure of the contract but a statement about the subject.
The round does not license two conclusions it invites. An audit of this length invites two conclusions it does not support, and both are worth closing explicitly.
It does not say that colorimetry does not work. It works extremely well within its stated conditions, and the conditions are met by most industrial colour measurement: photopic levels, a two-degree or ten-degree field, smooth reflective samples, and comparisons between samples of similar spectral shape. The departures measured here are what happens outside those conditions, and the round’s own identities say the arithmetic is exact inside them.
And it does not say that personal observers are the answer. A specification computed for the individual doing the looking would remove the departures and would remove the point of a specification, which is that two parties who have never met can agree about a number. A personalised observer is a measurement of one person and a contract has to serve everybody.
What the round licenses is narrower and more useful: knowing the size of the gap, knowing which conditions widen it, and knowing that four of the six terms respond to design choices somebody is already making for other reasons.
Two price lists on two samples make the point a specification needs. On a red pigment under daylight the six departures run from 1.20 to 2.38 ΔE₀₀; on a notch filter under the same light they run from 0.50 to 2.61 and in a different order.
So the contract’s price is not a constant even at fixed viewing conditions, and a specification carrying one number is carrying the number for whichever sample was measured.
Under a warm lamp the same budget roughly doubles for the two blue-absorbing terms and barely moves for the others, so a specification that names its viewing condition can price itself and one that does not cannot.
What was computed, and how
Nothing new is computed in this essay; it is the round’s numbers read differently. The six departures are the ones measured throughout, at their literature strengths, over the same forty-two surfaces and the same six lights.
The combination into about three units at the median is a quadrature sum over the six medians, which assumes independence and is known to be an underestimate because the lens and the macular departures are correlated across samples. It is quoted as an order of magnitude rather than as a result.
The claim about the 1924 luminous efficiency function’s blue error is from the literature rather than from this collection’s machinery, and it is one of the few numbers in this round that is quoted rather than computed.
One final thing follows from the contract reading and it concerns this collection rather than the field. Every figure here names its observer, and the naming has been treated for nineteen rounds as an epistemic virtue — a caption that says which kernel produced a number. Read as a contract, it is something slightly different and more useful: it is a statement of which agreement the figure is made under, which is what lets a reader compare it with a specification of their own.
That is why the naming survives the audit unchanged. Nothing in this round suggests the captions are wrong; what the round adds is that the named object is a choice with six arguments and a price, and a reader who has both the name and the price can do something a reader with only the name cannot.
Where the model stops
The contract framing is a way of reading a standard and not a claim anybody at the CIE has made. The standards’ own language is descriptive — the functions are presented as representing the colour-matching properties of an average observer — and the reading here is an interpretation of what they are used for rather than of what they say.
The essay also assumes the purpose of a specification is acceptability to a viewer, which is true of consumer goods and not of everything. A specification for a photometric standard, a signal light or a scientific reference has different purposes and different tolerances for the same departures.
And nothing here measures whether a viewer would actually reject a sample at three ΔE₀₀ of observer-induced difference. That is an acceptability question, it depends on the industry and the observer’s expectations, and this collection has addressed it only for instrument-reported differences.
There is a final asymmetry worth recording between the two halves of this round. The tabulation audit found defects that are repairable in software by anybody, at no cost beyond computation — a rule change, a wider range, a finer grid. The observer audit found six terms none of which is repairable at all, because they are properties of the people looking rather than of the arithmetic.
That is why the two halves end differently. The first produced a decision procedure and a list of repairs; this one produces a budget and a set of conditions that shrink it. An audit that finds an irreparable term has done its job by pricing it, and a reader expecting every audit to end in a fix will find the second half unsatisfying for the right reason.
The generalisation
The habit is about knowing which kind of object is being criticised.
A model can be wrong. A measurement can be wrong. A convention cannot be wrong; it can only be costly, and the cost is measurable in a way that rightness is not. Confusing the three produces two symmetrical errors: defending a convention as though it were true, and attacking it as though it were false.
The test is to ask what would change if it were replaced. If a better version would supersede it immediately, it is a measurement. If a better version would have to be adopted, negotiated and transitioned to, it is a convention, and the relevant question is what the transition costs against what the defect costs.
The failure mode this essay guards against is the one an audit invites. Six measured departures, all above a tolerance, none dominant, read as a demolition of the standard observer. They are a price list, and a price list is what a specification needs and what nobody has published.
Who found it, and when
The CIE adopted the 1931 observer at its eighth session in Cambridge, from Wright’s and Guild’s data reconciled and transformed, with ȳ constrained to the 1924 photopic luminous efficiency function. The constraint was a deliberate act of standardisation rather than a measurement: it welded photometry into colorimetry so that one system would serve both.
The short-wavelength defect was documented within two decades and corrected on paper by Judd and by Vos. Neither correction was adopted, and the CIE’s own 2006 observer — which is a much larger change — has not displaced the 1931 tables in industrial use either. Ninety-four years of a contract that everybody knows is imperfect is a fact about contracts rather than about colour.
Where the ladder goes next
The observer has been taken apart and the round’s second half is done. What remains is the third thing the previous round left: the scene solver has no slot for a glossy wall, and radiosity does not approximate one badly — it cannot express one at all.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A brand colour for a population acceptability · colour management · individual variation · observer metamerism · quality control · specification
- Fitted to an eye nobody has colour management · individual variation · observer metamerism · quality control · specification · standard observer
- The index is one observer's opinion acceptability · individual variation · observer metamerism · quality control · specification · standard observer
- The observer has no age acceptability · individual variation · observer metamerism · quality control · specification · standard observer
- A neutral is everyone's colour calibration · individual variation · observer metamerism · standard observer · structural choice
- A soft proof is exact for one reader colour management · individual variation · observer metamerism · specification · standard observer
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AcceptabilityCalibrationColour managementIndividual variationModelling assumptionObserver metamerismQuality controlSpecificationStandard observerStructural choice