The reference lamp must not move
Assumes The rods' route is priced by the lamp, The rods are a fourth curve and A gain is not an observer.
The rods’ route is priced by the lamp found that how much it matters whether a rod signal reaches the S-cone pathway depends almost entirely on the lamp: 15 per cent under daylight, a factor of 2.75 under a phosphor white LED. It then drew a design principle out of that, and the principle is the thing to test.
To measure an uncertain weight, use the condition in which the answer depends on it most, which is the opposite of the condition a standard is written for.
The experiment it proposed is a set of asymmetric matches — a surface seen under one lamp, matched by adjusting a field under another — with the S weight fitted to the settings. And an asymmetric match is not a reading under one condition. It is a difference of two readings, and the principle as stated is about one of them.
Daylight belongs in the experiment, on the other side
The three most sensitive pairs of the five lamps all have daylight in them. Daylight against a phosphor LED moves a match by 0.58 ΔE₀₀ as the S weight goes from nothing to equal; tungsten against a fluorescent tube, the pair the principle as stated would reach for, moves it by 0.09.
- A reference that moves with the weight cancels the signal. What an observer sets is where the two displacements agree, so only their difference is measurable.
- The best pair of two sensitive lamps is worse than every pair with daylight in it: phosphor against the tube reaches 0.22 against daylight-and-phosphor’s 0.58.
- The signal is ordered by the gap between the two lamps’ rod-to-S-cone catch ratios, at a rank correlation of 0.76 over the ten pairs — computable from the spectra and the four sensitivity curves before any observer is recruited.
- The cost is quadratic in that choice. At half a colour difference of scatter per setting, resolving the weight to a tenth of its range takes 75 settings on the best pair and 2,804 on the worst.
- That is about two sessions against seventy-one. The question is answerable in a morning or not at all, and which of the two is decided before the first observer sits down.
Why a match is a difference
The quantity the earlier essay computed is a departure: how far apart the reference observer and the rod-carrying observer land on the same surface under the same lamp. Under a phosphor LED that departure grows from 0.45 to 1.24 ΔE₀₀ as the S weight opens, which is a large change and is what made the LED look like the condition to measure in.
But nobody can read a departure. An observer does not report where they land; they report when two fields look the same, and the two fields are under different lamps. If the rod-carrying observer is displaced by one vector under the test lamp and a second vector under the reference lamp, the match they set is displaced by the difference — and the reference lamp’s displacement subtracts.
The five curves here are the reason the principle looked settled. Daylight’s is nearly flat, rising 15 per cent across the whole range; the phosphor LED’s runs steadily from 0.45 to 1.24, and every quarter step of the weight adds about a fifth of a colour difference. Read as five separate conditions they say unambiguously that the LED is where the answer lives and daylight is where it does not.
Read as candidates for the two halves of one match, they say something else. A pair’s signal is roughly the vertical gap between two of these curves at an S weight of one, minus the gap at nought — and two curves that both rise steeply have almost no gap between their rises. Daylight’s flatness, which reads as uselessness when the curves are read one at a time, is exactly the property that makes the gap large.
Seven of the ten pairs sit under a quarter of a colour difference and three sit above a third, and all three of those have daylight in them. Under daylight the rod signal is 0.86, 0.94 and 0.90 of the three cone classes’ own catches — nearly a common gain, which a gain is not an observer says a white point divides out exactly. So the daylight field barely moves as the weight opens, and it holds still while the other field does the moving.
The pair the principle as stated would have picked is two sensitive lamps, because each of them individually is a condition in which the answer depends on the weight a great deal. Tungsten and the fluorescent tube are both such lamps — their S excesses are 1.66 and 1.57 — and they are the worst pair in the table, because their displacements are so nearly equal that the difference is almost nothing.
The gap, not the sensitivity
That suggests the quantity to maximise is not either lamp’s sensitivity but the distance between them, and it is.
The cloud rises with a rank correlation of 0.76. The widest gap in the table is daylight against the phosphor LED, at 1.36, and it produces the largest movement. The narrowest gaps — the tube against the three-emitter LED at 0.19, tungsten against the tube at 0.26 — produce the smallest.
The useful part is that both axes are computable from the lamps. The S excess of a lamp is the rod curve integrated against that lamp’s spectrum, divided by each cone class’s own integral, and nothing about surfaces, observers or matches enters it. A session can therefore be designed from two spectra and four curves, which is what a design principle ought to deliver and what the stated one did not.
The three bars per lamp are where the excess comes from. Daylight’s are 0.86, 0.94 and 0.90 — within a tenth of each other, which is the near-gain. The phosphor LED’s are 0.58, 0.64 and 1.97, and the excess is the third minus the mean of the first two. What the pair table then asks is how far apart two lamps’ thirds bars sit once their first two are accounted for, and daylight sits at one end of that range on its own: every other lamp here has an S excess above 1.4 and daylight’s is 0.01.
That is worth stating as a fact about the lamp rather than about the experiment. Daylight is not a neutral reference because it is a standard; it is a neutral reference because it is smooth across the blue, and a lamp is not a blackbody is where the structure the other four have was priced for what it does to rendering. A second smooth lamp of a different colour temperature would serve as well, and a warm LED dimmed in the evening would not serve at all.
There is one caveat the scatter makes visible. The correlation is 0.76 and not 1.0, and the residual is the surfaces: a pair’s movement is a median over forty-two surfaces, and how much a particular surface moves depends on how much of the rod band it reflects. Two pairs with the same excess gap can differ by a factor of two in their medians, so the gap chooses a shortlist and the median over surfaces ranks it.
What the session costs
An experiment’s real currency is settings, and settings go as the inverse square of the signal.
The range is a factor of thirty-seven, from 75 settings to 2,804. At forty settings a session — which is a comfortable hour of asymmetric matching with the adaptation time these lamps need — that is two sessions against seventy-one.
The declared numbers are worth stating plainly because the conclusion is sensitive to them. Half a colour difference is the scatter assumed on one setting and a tenth of the range is the band the weight is wanted to; both are stated rather than measured here. Halving the assumed scatter divides every count by four and does not change the ordering; asking for a twentieth rather than a tenth multiplies every count by four and does not change it either. The ordering is what this figure is for, and it is a property of the lamps.
The inverse square is the part that makes the design choice matter more than it looks. A pair that moves the match half as far does not cost twice as much; it costs four times. So the difference between a well-chosen pair and a plausible one is not a marginal saving — it is the difference between an answerable question and an unanswerable one, and nothing about the second pair announces itself as hopeless.
What the observer is being asked to do
A few things about the experiment are decided by the same arithmetic and are worth saying, because they are where a design of this kind usually goes wrong.
The luminance has to be mesopic and stated. The whole effect exists because rods contribute, and the eye that has no colour is the regime description: below about three candelas a square metre the rods are running and the cones still are. The rod fraction the model uses — a tenth of each cone’s peak — stands for a particular point in that range, and the fitted weight is conditional on it.
The surfaces matter and the model says which. The movement is a median over forty-two constructed surfaces and its spread is wide, so a session that uses three samples chosen for convenience may land anywhere in that distribution. The samples to choose are the ones whose reflectance is high across the rod band and differs between the two lamps’ own strong regions.
And the observers have to be screened for the thing that would confound it. An observer’s own lens density and macular pigment both move the S-cone catch, and the observer has no age is the reminder that the standard observer has no such parameter. The macular is a band, not a filter is the sharper version of the same warning: a macular pigment absorbs over a band centred near 460 nanometres rather than cutting everything below a wavelength, so two observers differing in it differ most in exactly the region this experiment is reading. Two observers of different ages fitted to a common S weight would return the average of two different answers; the honest design fits a weight per observer and reports the spread.
There is a second confound that the pair choice creates rather than removes. Daylight and a phosphor LED differ in correlated colour temperature as well as in structure, so an observer moving between them is adapting across a large white-point change, and two lamps do not average is the reminder that an adaptation model is doing work in the middle of that. The signal computed here already has CAT16 in it, so the count of settings is conditional on the adaptation transform being right — which is a much stronger assumption than the rod model is. A pair of lamps matched in colour temperature and differing in structure would remove it, and the two constructed lamps that come closest are the phosphor LED and the three-emitter LED, whose pair sits sixth in the table at 0.22.
How the signal was computed
For each surface and each lamp, the reference observer’s CIELAB reading and the rod-carrying observer’s reading are taken with CAT16 adaptation to the lamp, and the departure is the difference of the two vectors. The match’s displacement is the test lamp’s departure minus the reference lamp’s, and the signal is the colour difference between where that displacement puts the match at an S weight of nought and at an S weight of one — which is the distance between the two hypotheses the experiment is deciding between.
The observer model is the site’s pigment-template eye at its median values, and the rod curve is the rhodopsin template through the ocular media without macular pigment, normalised to its peak and scaled to a tenth of each cone class’s own peak. The lamps are the census’s five: a 2856 K tungsten lamp, a 6500 K thermal radiator, a phosphor white LED, a three-emitter LED and a fluorescent tube with mercury lines. The three-laser projector is left out because on the five-nanometre grid its lines collapse to one and every observer agrees about it exactly.
The number of settings is the scatter divided by the band times the signal, squared — the standard error of a mean falling as the root of the count, with the signal converting a colour difference into a unit of weight.
Where this stops
The signal is linear in the weight and the test is not. The count of settings assumes the match moves proportionally as the weight opens, which the earlier sweep showed it very nearly does under most lamps and does not under tungsten, where a small S signal first offsets the tilt the L and M signals give before dominating it. A fit on tungsten pairs would need the non-monotone part modelled.
Half a colour difference per setting is a plausible number and not a measured one. Mesopic matching is harder than photopic matching and an asymmetric match is harder than a side-by-side one, so the true scatter is probably larger, which multiplies every count by the same factor and leaves the ordering alone.
The five lamps are constructed spectra. They are analytic sums of Gaussians and a Planckian, not measured lamps, so a real luminaire’s excess would have to be computed from its own measurement before a session was designed around it — which is the point of the excess being computable rather than a difficulty with it. The lamp in the shop decides is the standing reminder that the lamp a specification meets is the one on sale rather than the one in a table.
And a fitted weight is a fitted weight. The three routes bracket an uncertainty in a linear model of rod–cone interaction; the real circuit is nonlinear and its weights are themselves adaptation-dependent. What this experiment would return is the weight that best predicts a set of matches at one light level, which is a useful number and is not a measurement of a gap junction.
Still open: whether one field can carry both halves
An asymmetric match needs two lamps and two adapted states, which is what makes it slow: an observer has to sit in each room long enough to adapt, and a session of forty settings is mostly waiting.
There is a cheaper arrangement if the difference can be put inside one field. A bipartite field with the two lamps on the two halves has one adaptation state rather than two, and the match is set across the join in seconds. The question is what the observer is then adapted to — something between the two lamps, weighted by their areas — and whether the displacement difference survives being read under a common adaptation instead of two separate ones.
The computation is this one with a single CAT16 adaptation to the area-weighted mixture of the two lamps, rather than one adaptation per side. The prediction is that the signal falls, because a common adaptation removes part of what distinguishes the two lamps, and that it falls least for the daylight-against-LED pair, whose difference is in the shape of the spectrum rather than in its white point. If it falls by less than the factor of four that a halved signal costs, the bipartite field is the better experiment; if by more, the waiting buys something.
A design principle names a condition, and an experiment has two
The habit is about where a sensitivity analysis is pointed.
The principle the earlier essay derived is sound and it is about a reading: to estimate a parameter, observe where the observation depends on it most. Applied to a comparison, it has to be applied twice and in opposite directions, because what a comparison measures is a difference. The sensitive condition belongs on the side that varies and the insensitive one on the side that does not, and a principle stated as “use the sensitive condition” quietly recommends putting it on both.
The move is to write down what the observer actually reports before choosing the conditions. Here it is a setting at which two fields agree, so the estimand is a difference of two model outputs, and the design question is about the difference rather than about either term. That is a sentence of arithmetic and it reverses the answer.
It generalises past this experiment in a specific way. Any measurement that reports an agreement rather than a value has this shape — a nulling instrument, a bipartite field, a substitution weighing — and in every one of them the design question is about the difference of two model outputs. Two observers and one metamer made the neighbouring point about what a departure is: every one of these observer departures is a pairing of an observer’s deviation with a stimulus’s, so a departure is already a difference and a match is a difference of differences. Reading the sensitivity of the first and designing on it is one level short.
The failure mode is designing a comparison as though it were a reading. Both terms get chosen for the same property, the property cancels, and the experiment measures a residue of what it was aimed at — with nothing in its design to say so, because every individual choice in it was the sensitive one.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A dark background moves every difference and no match corresponding colours · modelling assumption
- A fourth emitter spends the gap it fills spectral power distribution · white led
- A lamp has a direction spectral power distribution · white led
- A lamp has a waveform spectral power distribution · white led
- A lamp switched on is not the lamp measured spectral power distribution · white led
- A limit written in energy charges the reds modelling assumption · spectral power distribution
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Cone fundamentalsCorresponding coloursMeasurement uncertaintyMesopicModelling assumptionObserver variabilityPsychophysicsRodsSpectral power distributionWhite LED