The booth and the eye disagree
Assumes The lens is worst under tungsten, The booth is a luminaire and The lamp in the shop decides.
A viewing booth exists so that two people can agree about a colour. The audit says it also happens to be the condition in which they agree most easily, which makes it a poor test of whether they would agree anywhere else.
The claim
A daylight booth minimises observer disagreement, and the lights a delivery is actually seen under roughly double it.
- The mean departure across the six terms is 1.77 ΔE₀₀ under a 6500 K radiator and 2.37 under tungsten — the lowest and the highest of the six lights measured.
- The reason is smoothness. A source with no narrow feature excites the localised departures least, and daylight has none.
- So a booth is where a specification is least likely to fail an observer test, and the shop, the office and the home are all worse.
- And nothing in a booth’s specification mentions this. Booths are specified for daylight simulation on the grounds of what colours they render, not of how much two people will disagree about them.
What a booth is specified for
The booth is a luminaire established what one actually is: a box with a lamp whose spectrum is designed to approximate a CIE daylight illuminant, assessed by a metamerism index that says how well pairs matching under real daylight still match under it.
The specification is entirely about rendering: does this lamp make samples look as they would under D65, and does it preserve matches made under D65. Both are questions about the light’s spectrum against a standard observer.
Neither is a question about the spread between observers. That quantity does not appear in the standards for viewing conditions, in the metamerism indices used to grade simulators, or in any booth manufacturer’s datasheet.
It turns out to be strongly correlated with the property booths are already specified for, which is fortunate and is worth understanding rather than relying on.
Why smooth light minimises the spread
The mechanism is the pairing structure this round has found everywhere.
An observer departure is a pairing of something belonging to the observer with something belonging to the stimulus’s deviation from the adapting white. The observer’s factor is localised — the macular pigment is a band at 460 nanometres, the lens is a tail below 500, the peaks are derivatives on the cone flanks — so what excites each departure is how much of the stimulus’s structure lands where that departure lives.
A smooth light spreads the stimulus’s structure broadly and lands a little of it on each departure. A light with a narrow feature concentrates it, and if the feature sits where a departure lives, that departure is excited strongly.
Daylight has no narrow feature at all. A tungsten lamp has none either and is nonetheless the worst, for the different reason the lens essay gives: it is so poor in the blue that the short-wavelength cone’s relative excitation is a ratio of two small numbers, which any blue-absorbing filter disturbs.
So smoothness protects against three of the six departures and blue content protects against two, and daylight is the only one of the six lights measured that has both.
The numbers, per light
Averaged over the six departures on a red pigment:
| light | mean departure |
|---|---|
| a 6500 K radiator | 1.77 ΔE₀₀ |
| a fluorescent tube | 1.64 |
| a white LED | 1.83 |
| a three-emitter LED | 2.08 |
| tungsten at 2856 K | 2.37 |
The fluorescent tube’s low mean deserves a caution rather than a recommendation: its lines are a nanometre wide, so on a finer grid its numbers move, and the figure quoted is on this collection’s own five-nanometre tabulation. It is not a light anybody should choose for observer robustness on the strength of that number.
The useful comparison is between the daylight radiator at 1.77 and the tungsten lamp at 2.37, a factor of 1.34 in the mean and a factor of two in the worst single term. Both are smooth, both are ordinary, and one is a booth and the other is a domestic interior.
Adding the three-emitter LED at 2.08 gives the modern retail case, which sits between them and closer to the tungsten end.
Comparing the two distributions is more informative than comparing two means, because the shift is not uniform. Under daylight the six medians run from 0.57 to 1.94 ΔE₀₀; under tungsten the lens and macular medians rise substantially while the density and field medians barely move.
So the warm-light penalty is not a scaling of the daylight answer. It is concentrated in two terms, both of which are absorptions in front of the receptors, and both of which are the terms an individual has no control over and a specification cannot measure.
That concentration is useful for anybody deciding how much to worry. A warm viewing condition roughly doubles the two largest terms and leaves the other four alone, which means the penalty is carried by the population’s spread in lens and macular density rather than by anything about the sample.
What this does to testing a specification
The consequence for practice is uncomfortable and it is specific.
A specification is validated by assessors in a booth. If two assessors agree, the specification is judged sound; if they disagree, the tolerance is widened or the sample is re-made. Either way the assessment is made in the condition where they are most likely to agree.
Then the delivery is seen in a shop, an office or a house, under a warmer and often narrower source, where the same two assessors would disagree by up to twice as much. The validation and the use are in different regimes, and the difference is in the direction that makes the validation optimistic.
That is the same structural complaint the shop’s lamp has always attracted, with a second mechanism added. The known one is that a metameric pair matching under the booth’s light comes apart under the shop’s. The new one is that the two people arguing about it disagree more under the shop’s light too.
What could be done about it
Three responses, in increasing order of cost and effectiveness.
Assess under the light of use. That is standard advice for the metamerism reason and it addresses this one at the same time. It is also often impractical, because the light of use is not one light.
Assess under two lights and require agreement under both. A booth with a daylight and a tungsten setting is ordinary hardware and the second setting is the more demanding one for both mechanisms. Requiring a pass under both is a strictly stronger test at almost no cost.
Or use a panel spanning ages. The lens departure is the only one of the six that is a trajectory rather than a spread, and a panel of assessors all in their twenties agrees with itself and not with the customers. A panel spanning twenty to sixty tests the specification against the population it will meet.
The third is the one nobody does, and it costs nothing but a hiring decision.
A two-light booth costs a rule change. The middle recommendation deserves a costing because it is the only one that is both effective and immediately available.
Booths with two or three illuminant settings are ordinary hardware and are already used, for the metamerism reason: a pair that matches under D65 and fails under tungsten is a metameric pair, and finding that out in the booth is much cheaper than finding it out from a customer. The equipment exists and the procedure exists.
What changes is the acceptance rule. At present a sample is usually assessed under the primary illuminant and checked under the secondary; making both a pass condition is stricter, will reject more samples, and is the correct test if the delivery will be seen under either.
The audit adds a second reason for the same change and it is independent of the first. Even for a sample that is not metameric with its standard — matched in spectrum, not merely in colour — the two assessors will disagree more under the warm setting, so the warm setting is the harder test of whether the tolerance is adequate for a population.
Two mechanisms recommending the same procedure is worth more than either, and the procedure is a rule change rather than a purchase.
The booth’s other silent decision
There is a second thing a booth fixes that the audit has something to say about, and it is the field size.
A booth presents a sample at a distance and a size that decides how many degrees it subtends, and the difference between the two standard observers is a field size worth 1.33 ΔE₀₀. A sample viewed at ten degrees is a different observer’s stimulus from the same sample at two.
Booth standards specify a geometry, a distance and a surround. They do not specify a subtended angle for the sample, because the sample’s size is whatever the sample is, and an assessor comparing a small swatch and a large panel is comparing them through two different observers.
That is a smaller term than the others and it is a term nobody has a way to control, since the sample’s size is a property of the product. It belongs in the list because it is the one place a booth’s own specification could carry an observer term and does not.
The modern retail case deserves its own paragraph because it is neither of the two lights the argument has been made with. A blue-pumped white LED is smooth over most of its range and has one narrow feature, and the feature is at 452 nanometres — inside the macular absorption band.
So it produces a ladder unlike either the daylight or the tungsten one: the macular departure leads at 3.37 ΔE₀₀, the lens is second at 2.35, and the mean is 1.83, only a little above daylight’s. A summary by mean would call it a benign light; a reader of the ladder would notice that its largest single term is the largest macular figure in the round.
A mean over the six departures hides which of them a lamp excites, and for a lamp with one narrow feature that is the whole of what there is to know. This collection has made the same point about a rendering index summarising a spectrum, and the two summaries fail in the same way for the same reason.
What was computed, and how
Each mean is the average of the six departures at their literature strengths, on a red pigment, computed through this collection’s usual CAT16 route. Averaging six departures is a summary rather than a prediction of what a person would see, and the quadrature combination is the better estimate for that.
The six lights are the audit’s constructions and the daylight one is a 6500 K thermal radiator rather than a CIE D65 reconstruction. Those differ in the ultraviolet and in the detail of their shape, and neither has a narrow feature, so the conclusion about smoothness is unaffected.
The fluorescent tube’s numbers are on the coarse grid and are noted as such above, because a light with lines narrower than the tabulation is not being computed correctly.
That figure bounds the whole essay’s claim and is worth having at the end of it. A warm booth does not make two assessors disagree about a grey card — the identity holds at the floating-point floor under every light in the round. What it does is roughly double their disagreement about a coloured one.
So the practical form of the recommendation is narrower than it might sound. A neutral-balancing step, a grey-scale check, a white-point verification: all of those are as reliable in a warm booth as in a daylight one. It is the saturated samples that need the second setting, which is convenient, because they are also the ones that need the metamerism check the second setting was installed for.
Where the model stops
The comparison is between five lights on one sample. A blue sample would reorder them, since the blue-absorbing departures would rise on every light and rise most on the warm ones.
The booth’s own light is modelled as a thermal radiator rather than as a real daylight simulator, which is a filtered tungsten source or a phosphor-based fluorescent lamp with a considerably less smooth spectrum. A real booth would sit somewhere between the daylight and the fluorescent rows, and probably closer to the daylight one for the better simulators.
And nothing here is about appearance. The assessors’ adaptation state, the surround’s lightness and the sample’s size all matter for whether a difference is noticed, and this round measures whether it is there.
One more consequence belongs to whoever writes the standard rather than uses it. Every quantity a viewing-condition standard controls is a property of the light: its spectrum, its illuminance, its uniformity, the surround’s reflectance. Every quantity this round measures is a property of the pairing between a light and a person, and no standard has a place to put one.
Adding one would not be a small change. It would mean a viewing-condition standard carrying a statement about the assessors — their age distribution at minimum — and assessor requirements in current standards stop at normal colour vision and a visual acuity check. A standard that specifies the room and not the people is a standard that assumes the people are interchangeable, and this round’s whole content is that they are not, by about three colour differences.
The generalisation
The habit is about testing in the condition that flatters.
A test performed under favourable conditions is a weaker test than the same test under ordinary ones, and the tell is that the conditions were chosen for a reason other than testing — for repeatability, for convenience, for standardisation. Those are all good reasons and none of them is about the test’s power.
The move is to ask what the test’s conditions were chosen to control, and whether the thing being tested is sensitive to something those conditions also happen to suppress. Here the booth controls the illuminant for repeatability and happens to suppress observer disagreement, which is a different variable and was never the point.
The failure mode is to conclude from a pass under standard conditions that a thing works. A standard condition is standard because it is reproducible, not because it is representative, and the two are frequently in tension.
There is a last, slightly deflating observation about how much any of this matters relative to the other things a booth controls. The illuminance, the surround’s lightness and the sample’s immediate background all change what an assessor reports by amounts that are hard to compare with a colour difference and are certainly not small. A booth’s specification controls all three tightly, for good reasons found by experiment over decades.
Against that, a factor of 1.34 in the observer term is a modest addition to a well-controlled situation rather than a discovery that the situation is uncontrolled. The reason it is worth stating is that it runs in the opposite direction from the others: the booth’s other controls make an assessment more representative of ordinary viewing, and this one makes it less.
Who found it, and when
Viewing conditions for colour appraisal have been standardised since the 1960s and the current standards specify the illuminant, its metamerism index, the surround’s reflectance and the illuminance. The observer’s variation is not among the quantities they control, and the observers themselves are usually required only to have normal colour vision.
That the daylight condition minimises observer spread does not appear to be stated anywhere as a reason for it, and it is very likely never to have been a consideration — the standards are about reproducing daylight because daylight is what things are seen in, and the rest is a happy accident.
Where the ladder goes next
The audit has now reached the room a colour is judged in. What is left is the colour itself, and one class of colour carries the term more than any other: a brand colour is chosen once and seen by everybody, which is the largest population any specification faces.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A brand colour for a population acceptability · individual variation · observer metamerism · quality control · specification
- A lamp switched on is not the lamp measured illuminant · quality control · specification · viewing condition · white led
- A tolerance is a probability acceptability · individual variation · observer metamerism · quality control · specification
- A tolerance needs a second number acceptability · individual variation · observer metamerism · quality control · specification
- An observer is a contract acceptability · individual variation · observer metamerism · quality control · specification
- The instrument is one observer exactly individual variation · measurement condition · observer metamerism · quality control · specification
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AcceptabilityIlluminantIndividual variationLuminaireMeasurement conditionObserver metamerismQuality controlSpecificationViewing conditionWhite LED