Two sensors disagree about deep red, not lines
Assumes Flicker sorts lamps the wrong way, Two matrices do not reach a white LED and Three numbers cannot see a line.
Two matrices do not reach a white LED found that a camera profile’s two matrices, blended by colour temperature, are about twice as wrong under a lamp with lines in its spectrum as a matrix fitted under that lamp. Whether a lamp has lines is exactly what a white balance reading cannot see, since three numbers cannot see a line. So a device that wants the right matrix needs some other way to classify its lamp.
Flicker sorts lamps the wrong way tested the first candidate and found it measured how a lamp’s power is delivered rather than how its spectrum is shaped. It named the next: a phone already carries an ambient-light sensor with channels quite unlike the camera’s, and two instruments whose sensitivities are not linear transforms of each other disagree about a light by an amount that depends on its spectrum. That is observer metamerism with instruments in place of people, and the essay predicted it would work and be fragile.
It works on every lamp it was tried on, and the reason it works is not the one proposed.
Perfect on the lamps, for a different reason
With a red, green, blue and clear ambient sensor, every one of eight structured lamps disagrees with the camera by at least 0.133 and every one of six smooth lamps by at most 0.057. Choosing a matrix by that number costs nothing on the twelve fixtures where choosing by flicker cost 0.435 colour differences.
- The separation is mostly one channel’s: 59 per cent of the squared disagreement over the structured lamps is in the clear channel, which collects the deep red behind the camera’s infrared-cut filter.
- So the classifier reads deep red, and constructed lamps prove it. Daylight and a 3000 K radiator with their light above 660 nm rolled off score 0.168 and 0.290 and are called structured. A white LED and a three-emitter source with a far-red band added score 0.086 and 0.068 and are called smooth.
- Without the clear channel the classes still separate, by a third of the margin, and a one per cent calibration error breaks it. With it, five per cent does.
- In a room lit by an LED and a window the number is a share, not a class: it rises steadily with the LED’s share of the light and crosses the line at about 62 per cent.
A camera, a second sensor, and a map between them
The camera is the modelled silicon sensor — colour dyes on silicon behind an infrared-cut filter — whose profile work these essays have used throughout. The ambient-light sensor is modelled rather than measured, because no measured pair of sensitivities for a real phone’s two sensors is in hand here. It is bare silicon under gaussian filters with no infrared cut, in three designs a phone might carry: red, green, blue and clear channels; red, green and blue alone; and a clear channel with one photopic-like channel of the kind a lux meter uses.
The statistic is a prediction and its failure. A linear map from the camera’s white to the sensor’s reading is fitted over smooth lights only — radiators from 2500 to 6500 K and daylights from 4000 to 10000 K — so it learns how the two instruments relate on spectra without structure. A lamp’s disagreement is the relative residual when that map predicts the sensor’s reading from the camera’s. None of the fourteen lamps scored is among the lights the map was fitted on, and the fitted lights themselves leave a residual of at most 0.011.
The threshold sits at the geometric middle of the gap between the classes, which is the most favourable place it can go, since it was placed using the lamps it is scored on. Everything this essay finds wrong is wrong at that flattering threshold.
The fourteen lamps
The figure above is the result as a classifier would report it.
The four smooth radiators and daylights score between 0.002 and 0.011, essentially the fit’s own residual. The two tinted radiators — a 4000 K radiator filtered towards green and one towards pink, both off the daylight locus where the map was never fitted — score 0.053 and 0.057, the highest of the smooth lamps and still well clear. The structured lamps start at 0.133 for the three-emitter source, run through the three phosphor LEDs and the halophosphate tube at about 0.14 and the broadband tube at 0.155, and end at 0.195 for the triphosphor tube and 0.226 for the LED with a red phosphor.
No threshold anywhere between 0.057 and 0.133 makes a mistake, which is the result flicker could not reach at any threshold. And the regret follows. Each of the twelve fixtures the flicker essay used carries a price for each matrix — the blend and a third matrix fitted under a white LED — which is a matrix is fitted under one light applied to a choice of two, and a classifier that is right everywhere hands every fixture its better matrix.
Choosing by the sensor costs 0.000; choosing by flicker, 0.435; always using the third matrix, 0.419; always blending, 0.511. On the cases that exist, the prediction that a second sensor would work is borne out completely.
That is where a cross-tabulation would stop, and it is where the flicker essay’s habit said to stop — list the cases, mark each with its true class and with what the proxy says, and look at the off-diagonal. Here the off-diagonal is empty. The question left is whether it is empty because the proxy is the property, or because on these fourteen lamps the proxy and the property happen to agree.
Where the disagreement comes from
A residual has components, one per sensor channel, and they need not be equal.
Over the eight structured lamps, 59 per cent of the squared disagreement is in the clear channel, 25 in green, 9 in blue and 7 in red. The clear channel is the one with no counterpart in the camera. The camera’s red channel is at 45 per cent of its peak at 660 nm, 7 per cent at 680 and under 1 per cent at 700, because the infrared-cut filter in front of every camera sensor is there to stop exactly that light. Bare silicon is near its peak there. Whatever a lamp emits between 680 and 780 nm reaches the ambient sensor’s clear channel and never reaches the camera at all.
So the map from camera to sensor has to guess the deep red, and it guesses it the way the smooth lights it was fitted on taught it: from the slope of the red end of the spectrum the camera can see. A radiator or a daylight continues smoothly past 680 nm and the guess is good. A phosphor LED or a fluorescent tube does not — its red falls away before the deep red, and on every one of the eight structured lamps the clear channel collects less than the camera’s red predicts. The shortfall is largest for the LED with a red phosphor, whose extra red near 650 nm raises the camera’s red and so the prediction, while adding little past 680.
The fourteen lamps sort by structure because, among them, the structured lamps are also the lamps with less deep red than their visible red implies. That is a property of how white LEDs and tubes happen to be made, not of spectral lines.
Lamps built to separate the two
A proxy that agrees with the property on every case available can be tested only by making cases where they come apart. Deep red and spectral structure are independent properties, so each can be changed without the other.
Roll off a 3000 K radiator’s light above 720 nm and it scores 0.116, above the threshold; daylight crosses at 700 nm. Cut both at 660 nm and they score 0.290 and 0.168, higher than most of the real LEDs. Neither spectrum has acquired a line. Each is exactly as smooth as it was, and the luminance each lost is under one per cent.
Add a band of far red at 720 nm to a neutral white LED, carrying about a sixth of its power and about a tenth of one per cent of its luminance, and it scores 0.086, just on the smooth side of the line; a three-emitter source with the same addition scores 0.068. Neither has lost a line. The LED still has its blue pump and its phosphor hump, and the three-emitter source still has three narrow peaks, and each is now classified as a lamp the blended matrix suits.
Both constructions have ordinary counterparts. Horticultural and some full-spectrum fixtures add a far-red emitter near 730 nm to white LEDs, which is the right-hand panel. A glass or diffuser that trims the far red from daylight is the left, and whether a particular glazing trims enough is a measurement on that glazing rather than something this model can say. In both, the classifier is confidently wrong in the direction that costs: a structured lamp handed the blend, a smooth one handed the LED matrix.
That is the difference between this failure and flicker’s. Flicker was wrong on lamps that exist in every kitchen. This classifier is right on every lamp of that kind and wrong on a class that is rarer but not exotic, and nothing in its own output distinguishes the two.
The margin is the clear channel’s
If the clear channel carries most of the separation, a sensor without one should do worse, and the three designs measure how much.
All three designs separate the fourteen lamps. With red, green, blue and clear the gap runs from 0.057 to 0.133; with a clear channel and a photopic-like one, from 0.024 to 0.102. With red, green and blue alone it runs from 0.011 to 0.035, a gap of 0.024 against 0.076 and 0.078.
The design without a clear channel is reading something closer to structure. Adding far red to the LED and the three-emitter source leaves both on the structured side, at 0.047 and 0.027 against its threshold of 0.019, because none of its channels reaches 720 nm. It is not immune — a smooth lamp cut at 660 nm still crosses its line, because its red filter reaches into that range — but its errors are no longer about light the camera cannot see.
And it pays for that with its margin. The clear channel does two things at once: it widens the gap between the classes by a factor of three, and it makes the gap a measurement of deep red. The design choice that makes the classifier robust is the design choice that makes it read the wrong property.
What calibration error spends
The essay that proposed this classifier predicted fragility, on the reasoning that its statistic is a difference of two calibrations. The margin is where that shows.
The four-channel design and the clear-and-lux design make no mistake until each channel’s calibration is 5 per cent out in the worst combination; at 7.5 per cent the four-channel design gets five lamps wrong and at 10 per cent seven. The red, green and blue design makes its first mistake at 1 per cent, and at 3 per cent gets nine of the fourteen wrong.
A per-channel gain error is the simplest thing that happens to a sensor between the factory and a photograph: unit-to-unit variation in filter thickness, the dark current’s dependence on temperature, and the slow change of a dye with light and age all look like one. Whether a production ambient sensor holds its channels relative to a production camera to five per cent across all of those is a property of parts that are not modelled here. It is a much looser requirement than one per cent, and it is the requirement the design reading the right property imposes.
So the prediction of fragility is half right. The design that would classify by structure is fragile. The design that is not fragile classifies by something else.
A room is a mixture
The fourteen lamps were scored one at a time, and a room with two lights has no white. The sensor, mounted on the front of a phone, reads the whole room.
The score rises at every step, from 0.005 with daylight alone to 0.139 with the LED alone, and crosses the threshold at about 62 per cent LED. A room lit mostly by the window is called smooth and handed the blend; a room lit mostly by the LED is called structured and handed the LED matrix.
That is a sensible rule if the colour matrix’s error in a mixed room falls between its errors under the two lamps in proportion to the shares. It is not the rule the classifier was designed to implement, which was to detect whether a structured lamp is present. And the camera may not be pointing at the part of the room the sensor reads: two lamps decide what one lamp could not found surfaces in a two-lamp room lit in different proportions, and a phone’s front sensor faces the photographer rather than the subject.
How the instruments were modelled
The camera sensitivities are the silicon quantum efficiency times gaussian dye transmittances centred at 605, 535 and 455 nm times the infrared-cut filter. The ambient sensor’s channels are the same silicon quantum efficiency under gaussian filters at 615, 535 and 465 nm with widths of 40, 45 and 35 nm, a clear channel of silicon alone, and a photopic-like channel at 555 nm with a width of 45 nm. Lamps are the visible spectra from the two-matrix work, from 380 to 780 nm.
Each instrument’s reading is reduced to shares of its own channels, so that the lamp’s brightness cancels. The map is a least-squares fit, one row per sensor channel, from the camera’s three shares to the sensor’s, over seventeen radiators and thirteen daylights. The disagreement is the length of the prediction’s error divided by the length of the reading.
A calibration error multiplies each sensor channel by one plus, minus or zero times the stated fraction, and every combination is tried, which is 81 for four channels and 27 for three. The mixture is the sum of the two lamps’ spectra, each normalised to equal power across the visible band, in the stated proportion.
What this leaves out
The sensitivities are modelled, and the result depends on one feature of them: that the ambient sensor has a channel reaching past the camera’s infrared cut. Real ambient sensors are built that way, often with an explicit infrared channel beside the clear one, but the numbers here are for gaussian filters on a generic silicon curve.
The spectra stop at 780 nm. Real radiators and daylight emit far more beyond that than inside it, and a real clear channel collects some of it. That would move smooth thermal lamps further from the camera’s prediction rather than closer, and an ambient sensor with an explicit infrared channel sorts lamps into thermal and non-thermal — a third property, which again coincides with structure on most lamps that exist and not on a filtered daylight or an LED with a near-infrared emitter.
The map is linear and fitted on one family. A map fitted with some structured lamps in it would reduce their residuals and the gap with them; a nonlinear map could do better or worse. The flattering threshold is a best case and a deployed classifier would sit somewhere worse.
And the camera has its own metamers: two lamps the camera records identically may differ to the sensor for reasons that have nothing to do with either class. The scoring here asks only whether the fourteen lamps sort; it does not ask how many different lamps share each score.
Still open: whether a third channel can tell deep red from structure
The design without a clear channel read something nearer to structure and paid for it in margin. The design with one bought margin and read deep red. The natural question is whether a sensor can have both — a margin from somewhere other than the region the camera cannot see.
The candidate is a channel placed inside the camera’s range but between its peaks, where phosphor LEDs and tubes have their characteristic dips: near 480 nm, between the blue pump and the phosphor, and roughly 570 to 600 nm for tubes. A narrow channel there sees structure directly, and a smooth lamp cut in the deep red does not move it.
The calculation is the same one run with a fourth design and the constructed lamps included in the test. It should report three numbers: the margin on the fourteen lamps, the calibration error at which it first fails, and whether the far-red and trimmed-red constructions still cross the line. If the margin survives without the clear channel, the proxy has become the property. If it does not, the choice between a robust classifier and a correct one is a choice a sensor designer makes rather than one arithmetic removes.
A cross-tabulation passes on the cases that exist
The habit is about what a clean result on real cases can and cannot show.
The flicker essay’s move was to list the cases and look at the off-diagonal before building anything, and it found the flaw at once. The same move here finds nothing: fourteen lamps, no mistakes, no regret. The temptation is to read an empty off-diagonal as proof that the proxy measures the property.
It shows only that the two agree on the cases listed. When two properties are correlated in the population of things that get built — structured lamps also lack deep red, because of how phosphors are chosen — every real case agrees and the proxy looks perfect.
The move is to construct the cases the population does not contain: change the proxy without changing the property, and the property without the proxy, and see which one the classifier follows. Here that took two filters and an added band, and it said at once what the fourteen lamps could not.
The failure mode is to stop at the empty off-diagonal. A classifier validated only on the cases that exist is validated on the correlation that made them, and it inherits every lamp built differently.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A corner is corrected by one row calibration · colour matrix · identifiability · infrared
- The chart was measured by an observer too calibration · camera profile · colour matrix · observer metamerism
- A sensor has no lens calibration · camera profile · observer metamerism
- The chart decides the profile calibration · colour matrix · identifiability
- A brand colour for a population observer metamerism · spectral structure
- A fit can be exact and empty calibration · identifiability
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CalibrationCamera profileCamera sensitivityColour matrixFluorescentIdentifiabilityInfraredLED emissionObserver metamerismSpectral structure