What it takes to deliver it

The booth and the eye disagree

A viewing booth is specified to simulate daylight, and daylight is the light under which observers disagree least. So a booth is the condition in which a specification is least likely to be tested against the thing that breaks it, and the shop it will be judged in roughly doubles the term.

Assumes The lens is worst under tungsten, The booth is a luminaire and The lamp in the shop decides.

A viewing booth exists so that two people can agree about a colour. The audit says it also happens to be the condition in which they agree most easily, which makes it a poor test of whether they would agree anywhere else.

Every departure under every light. Six departures across six lights, each cell the difference between two observers in ΔE₀₀, drawn as a bar whose length is the number. The rows are not multiples of one another: the lens is worst under tungsten and the pigment peaks are worst under a three-emitter LED, because a departure is a pairing and which light is being paired with decides it. The laser projector's row is empty, and that is not a fact about lasers — on this collection's five-nanometre grid a three-line spectrum is a one-line spectrum, and a single wavelength is a stimulus every observer agrees about exactly.
Fig. 1 Every observer departure under every light. The daylight row is the lowest overall, and the tungsten row is the highest.

The claim

A daylight booth minimises observer disagreement, and the lights a delivery is actually seen under roughly double it.

  • The mean departure across the six terms is 1.77 ΔE₀₀ under a 6500 K radiator and 2.37 under tungsten — the lowest and the highest of the six lights measured.
  • The reason is smoothness. A source with no narrow feature excites the localised departures least, and daylight has none.
  • So a booth is where a specification is least likely to fail an observer test, and the shop, the office and the home are all worse.
  • And nothing in a booth’s specification mentions this. Booths are specified for daylight simulation on the grounds of what colours they render, not of how much two people will disagree about them.

What a booth is specified for

The booth is a luminaire established what one actually is: a box with a lamp whose spectrum is designed to approximate a CIE daylight illuminant, assessed by a metamerism index that says how well pairs matching under real daylight still match under it.

The specification is entirely about rendering: does this lamp make samples look as they would under D65, and does it preserve matches made under D65. Both are questions about the light’s spectrum against a standard observer.

Neither is a question about the spread between observers. That quantity does not appear in the standards for viewing conditions, in the metamerism indices used to grade simulators, or in any booth manufacturer’s datasheet.

It turns out to be strongly correlated with the property booths are already specified for, which is fortunate and is worth understanding rather than relying on.

Why smooth light minimises the spread

The mechanism is the pairing structure this round has found everywhere.

An observer departure is a pairing of something belonging to the observer with something belonging to the stimulus’s deviation from the adapting white. The observer’s factor is localised — the macular pigment is a band at 460 nanometres, the lens is a tail below 500, the peaks are derivatives on the cone flanks — so what excites each departure is how much of the stimulus’s structure lands where that departure lives.

A smooth light spreads the stimulus’s structure broadly and lands a little of it on each departure. A light with a narrow feature concentrates it, and if the feature sits where a departure lives, that departure is excited strongly.

Daylight has no narrow feature at all. A tungsten lamp has none either and is nonetheless the worst, for the different reason the lens essay gives: it is so poor in the blue that the short-wavelength cone’s relative excitation is a ratio of two small numbers, which any blue-absorbing filter disturbs.

So smoothness protects against three of the six departures and blue content protects against two, and daylight is the only one of the six lights measured that has both.

Six departures of the observer, each at a stated strength, under a 6500 K thermal radiator. Each bar is two observers differing in one argument, looking at the same sample under the same light, in ΔE₀₀. The strengths are the literature's: the working-age lens, two standard deviations of the reported macular and density spreads, the long-wavelength polymorphism, the CIE's own second observer, and a rod contribution of a tenth. They are within a factor of 2.0 of one another, which is the point: there is no single term to fix. Every one of them is above the ΔE of about one that a delivery tolerance is written in.
Fig. 2 The six departures under a 6500 K radiator, which is what a booth is designed to simulate. This is the flattest and lowest ladder of the six lights.

The numbers, per light

Averaged over the six departures on a red pigment:

light mean departure
a 6500 K radiator 1.77 ΔE₀₀
a fluorescent tube 1.64
a white LED 1.83
a three-emitter LED 2.08
tungsten at 2856 K 2.37

The fluorescent tube’s low mean deserves a caution rather than a recommendation: its lines are a nanometre wide, so on a finer grid its numbers move, and the figure quoted is on this collection’s own five-nanometre tabulation. It is not a light anybody should choose for observer robustness on the strength of that number.

The useful comparison is between the daylight radiator at 1.77 and the tungsten lamp at 2.37, a factor of 1.34 in the mean and a factor of two in the worst single term. Both are smooth, both are ordinary, and one is a booth and the other is a domestic interior.

Adding the three-emitter LED at 2.08 gives the modern retail case, which sits between them and closer to the tungsten end.

Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.
Fig. 3 Each departure over forty-two surfaces under a tungsten lamp. Every distribution is higher than its daylight counterpart, and the two blue-absorbing ones are higher by most.

Comparing the two distributions is more informative than comparing two means, because the shift is not uniform. Under daylight the six medians run from 0.57 to 1.94 ΔE₀₀; under tungsten the lens and macular medians rise substantially while the density and field medians barely move.

So the warm-light penalty is not a scaling of the daylight answer. It is concentrated in two terms, both of which are absorptions in front of the receptors, and both of which are the terms an individual has no control over and a specification cannot measure.

That concentration is useful for anybody deciding how much to worry. A warm viewing condition roughly doubles the two largest terms and leaves the other four alone, which means the penalty is carried by the population’s spread in lens and macular density rather than by anything about the sample.

What this does to testing a specification

The consequence for practice is uncomfortable and it is specific.

A specification is validated by assessors in a booth. If two assessors agree, the specification is judged sound; if they disagree, the tolerance is widened or the sample is re-made. Either way the assessment is made in the condition where they are most likely to agree.

Then the delivery is seen in a shop, an office or a house, under a warmer and often narrower source, where the same two assessors would disagree by up to twice as much. The validation and the use are in different regimes, and the difference is in the direction that makes the validation optimistic.

That is the same structural complaint the shop’s lamp has always attracted, with a second mechanism added. The known one is that a metameric pair matching under the booth’s light comes apart under the shop’s. The new one is that the two people arguing about it disagree more under the shop’s light too.

What could be done about it

Three responses, in increasing order of cost and effectiveness.

Assess under the light of use. That is standard advice for the metamerism reason and it addresses this one at the same time. It is also often impractical, because the light of use is not one light.

Assess under two lights and require agreement under both. A booth with a daylight and a tungsten setting is ordinary hardware and the second setting is the more demanding one for both mechanisms. Requiring a pass under both is a strictly stronger test at almost no cost.

Or use a panel spanning ages. The lens departure is the only one of the six that is a trajectory rather than a spread, and a panel of assessors all in their twenties agrees with itself and not with the customers. A panel spanning twenty to sixty tests the specification against the population it will meet.

The third is the one nobody does, and it costs nothing but a hiring decision.

Six departures of the observer, each at a stated strength, under a tungsten lamp at 2856 K. Each bar is two observers differing in one argument, looking at the same sample under the same light, in ΔE₀₀. The strengths are the literature's: the working-age lens, two standard deviations of the reported macular and density spreads, the long-wavelength polymorphism, the CIE's own second observer, and a rod contribution of a tenth. They are within a factor of 5.7 of one another, which is the point: there is no single term to fix. Every one of them is above the ΔE of about one that a delivery tolerance is written in.
Fig. 4 The six departures under a tungsten lamp. Two terms have separated from the rest and both are blue absorptions, which is the warm-light signature.

A two-light booth costs a rule change. The middle recommendation deserves a costing because it is the only one that is both effective and immediately available.

Booths with two or three illuminant settings are ordinary hardware and are already used, for the metamerism reason: a pair that matches under D65 and fails under tungsten is a metameric pair, and finding that out in the booth is much cheaper than finding it out from a customer. The equipment exists and the procedure exists.

What changes is the acceptance rule. At present a sample is usually assessed under the primary illuminant and checked under the secondary; making both a pass condition is stricter, will reject more samples, and is the correct test if the delivery will be seen under either.

The audit adds a second reason for the same change and it is independent of the first. Even for a sample that is not metameric with its standard — matched in spectrum, not merely in colour — the two assessors will disagree more under the warm setting, so the warm setting is the harder test of whether the tolerance is adequate for a population.

Two mechanisms recommending the same procedure is worth more than either, and the procedure is a rule change rather than a purchase.

The booth’s other silent decision

There is a second thing a booth fixes that the audit has something to say about, and it is the field size.

A booth presents a sample at a distance and a size that decides how many degrees it subtends, and the difference between the two standard observers is a field size worth 1.33 ΔE₀₀. A sample viewed at ten degrees is a different observer’s stimulus from the same sample at two.

Booth standards specify a geometry, a distance and a surround. They do not specify a subtended angle for the sample, because the sample’s size is whatever the sample is, and an assessor comparing a small swatch and a large panel is comparing them through two different observers.

That is a smaller term than the others and it is a term nobody has a way to control, since the sample’s size is a property of the product. It belongs in the list because it is the one place a booth’s own specification could carry an observer term and does not.

Six departures of the observer, each at a stated strength, under a white LED, blue pump and one phosphor. Each bar is two observers differing in one argument, looking at the same sample under the same light, in ΔE₀₀. The strengths are the literature's: the working-age lens, two standard deviations of the reported macular and density spreads, the long-wavelength polymorphism, the CIE's own second observer, and a rod contribution of a tenth. They are within a factor of 3.9 of one another, which is the point: there is no single term to fix. Every one of them is above the ΔE of about one that a delivery tolerance is written in.
Fig. 5 The six departures under a white LED, which is what a great many shops and offices now use. The macular pigment has moved to the top, because the lamp’s blue pump lands on its absorption band.

The modern retail case deserves its own paragraph because it is neither of the two lights the argument has been made with. A blue-pumped white LED is smooth over most of its range and has one narrow feature, and the feature is at 452 nanometres — inside the macular absorption band.

So it produces a ladder unlike either the daylight or the tungsten one: the macular departure leads at 3.37 ΔE₀₀, the lens is second at 2.35, and the mean is 1.83, only a little above daylight’s. A summary by mean would call it a benign light; a reader of the ladder would notice that its largest single term is the largest macular figure in the round.

A mean over the six departures hides which of them a lamp excites, and for a lamp with one narrow feature that is the whole of what there is to know. This collection has made the same point about a rendering index summarising a spectrum, and the two summaries fail in the same way for the same reason.

What was computed, and how

Each mean is the average of the six departures at their literature strengths, on a red pigment, computed through this collection’s usual CAT16 route. Averaging six departures is a summary rather than a prediction of what a person would see, and the quadrature combination is the better estimate for that.

The six lights are the audit’s constructions and the daylight one is a 6500 K thermal radiator rather than a CIE D65 reconstruction. Those differ in the ultraviolet and in the detail of their shape, and neither has a narrow feature, so the conclusion about smoothness is unaffected.

The fluorescent tube’s numbers are on the coarse grid and are noted as such above, because a light with lines narrower than the tabulation is not being computed correctly.

The conditions under which an observer's departure is exactly zero. A departure of the observer is the pairing of something belonging to the observer with something belonging to the stimulus, so emptying either factor empties the product. The axis is logarithmic in what is left when the condition is imposed. Six rows empty the stimulus's factor — a perfectly neutral sample is the same colour for every observer, at any age and any field size — and two empty the observer's, since a gain on each cone and a change of basis are both absorbed exactly. All eight are identities rather than small numbers. The last two are the same two conditions imposed in a published cone space rather than in the observer's own, and they are worth eight and thirteen units: the identity is about the eye, and the arithmetic everybody uses is in somebody else's coordinates.
Fig. 6 The conditions under which a departure vanishes, imposed under a tungsten lamp. The eight identities hold under every light, so the warm-light penalty is entirely about coloured samples.

That figure bounds the whole essay’s claim and is worth having at the end of it. A warm booth does not make two assessors disagree about a grey card — the identity holds at the floating-point floor under every light in the round. What it does is roughly double their disagreement about a coloured one.

So the practical form of the recommendation is narrower than it might sound. A neutral-balancing step, a grey-scale check, a white-point verification: all of those are as reliable in a warm booth as in a daylight one. It is the saturated samples that need the second setting, which is convenient, because they are also the ones that need the metamerism check the second setting was installed for.

Where the model stops

The comparison is between five lights on one sample. A blue sample would reorder them, since the blue-absorbing departures would rise on every light and rise most on the warm ones.

The booth’s own light is modelled as a thermal radiator rather than as a real daylight simulator, which is a filtered tungsten source or a phosphor-based fluorescent lamp with a considerably less smooth spectrum. A real booth would sit somewhere between the daylight and the fluorescent rows, and probably closer to the daylight one for the better simulators.

And nothing here is about appearance. The assessors’ adaptation state, the surround’s lightness and the sample’s size all matter for whether a difference is noticed, and this round measures whether it is there.

One more consequence belongs to whoever writes the standard rather than uses it. Every quantity a viewing-condition standard controls is a property of the light: its spectrum, its illuminance, its uniformity, the surround’s reflectance. Every quantity this round measures is a property of the pairing between a light and a person, and no standard has a place to put one.

Adding one would not be a small change. It would mean a viewing-condition standard carrying a statement about the assessors — their age distribution at minimum — and assessor requirements in current standards stop at normal colour vision and a visual acuity check. A standard that specifies the room and not the people is a standard that assumes the people are interchangeable, and this round’s whole content is that they are not, by about three colour differences.

The generalisation

The habit is about testing in the condition that flatters.

A test performed under favourable conditions is a weaker test than the same test under ordinary ones, and the tell is that the conditions were chosen for a reason other than testing — for repeatability, for convenience, for standardisation. Those are all good reasons and none of them is about the test’s power.

The move is to ask what the test’s conditions were chosen to control, and whether the thing being tested is sensitive to something those conditions also happen to suppress. Here the booth controls the illuminant for repeatability and happens to suppress observer disagreement, which is a different variable and was never the point.

The failure mode is to conclude from a pass under standard conditions that a thing works. A standard condition is standard because it is reproducible, not because it is representative, and the two are frequently in tension.

There is a last, slightly deflating observation about how much any of this matters relative to the other things a booth controls. The illuminance, the surround’s lightness and the sample’s immediate background all change what an assessor reports by amounts that are hard to compare with a colour difference and are certainly not small. A booth’s specification controls all three tightly, for good reasons found by experiment over decades.

Against that, a factor of 1.34 in the observer term is a modest addition to a well-controlled situation rather than a discovery that the situation is uncontrolled. The reason it is worth stating is that it runs in the opposite direction from the others: the booth’s other controls make an assessment more representative of ordinary viewing, and this one makes it less.

Who found it, and when

Viewing conditions for colour appraisal have been standardised since the 1960s and the current standards specify the illuminant, its metamerism index, the surround’s reflectance and the illuminance. The observer’s variation is not among the quantities they control, and the observers themselves are usually required only to have normal colour vision.

That the daylight condition minimises observer spread does not appear to be stated anywhere as a reason for it, and it is very likely never to have been a consideration — the standards are about reproducing daylight because daylight is what things are seen in, and the rest is a happy accident.

Where the ladder goes next

The audit has now reached the room a colour is judged in. What is left is the colour itself, and one class of colour carries the term more than any other: a brand colour is chosen once and seen by everybody, which is the largest population any specification faces.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AcceptabilityIlluminantIndividual variationLuminaireMeasurement conditionObserver metamerismQuality controlSpecificationViewing conditionWhite LED