What a camera does

Two sensors disagree about deep red, not lines

A phone's ambient-light sensor and its camera read the same lamp differently, and the difference sorts fourteen lamps into smooth and structured without a single mistake — where flicker made five. It is not reading their lines. Almost all of it comes from the sensor's clear channel collecting deep red the camera's infrared cut throws away, so daylight with its far red trimmed is called structured and a white LED with a far-red emitter is called smooth.

Assumes Flicker sorts lamps the wrong way, Two matrices do not reach a white LED and Three numbers cannot see a line.

Two matrices do not reach a white LED found that a camera profile’s two matrices, blended by colour temperature, are about twice as wrong under a lamp with lines in its spectrum as a matrix fitted under that lamp. Whether a lamp has lines is exactly what a white balance reading cannot see, since three numbers cannot see a line. So a device that wants the right matrix needs some other way to classify its lamp.

Flicker sorts lamps the wrong way tested the first candidate and found it measured how a lamp’s power is delivered rather than how its spectrum is shaped. It named the next: a phone already carries an ambient-light sensor with channels quite unlike the camera’s, and two instruments whose sensitivities are not linear transforms of each other disagree about a light by an amount that depends on its spectrum. That is observer metamerism with instruments in place of people, and the essay predicted it would work and be fragile.

It works on every lamp it was tried on, and the reason it works is not the one proposed.

Perfect on the lamps, for a different reason

With a red, green, blue and clear ambient sensor, every one of eight structured lamps disagrees with the camera by at least 0.133 and every one of six smooth lamps by at most 0.057. Choosing a matrix by that number costs nothing on the twelve fixtures where choosing by flicker cost 0.435 colour differences.

  • The separation is mostly one channel’s: 59 per cent of the squared disagreement over the structured lamps is in the clear channel, which collects the deep red behind the camera’s infrared-cut filter.
  • So the classifier reads deep red, and constructed lamps prove it. Daylight and a 3000 K radiator with their light above 660 nm rolled off score 0.168 and 0.290 and are called structured. A white LED and a three-emitter source with a far-red band added score 0.086 and 0.068 and are called smooth.
  • Without the clear channel the classes still separate, by a third of the margin, and a one per cent calibration error breaks it. With it, five per cent does.
  • In a room lit by an LED and a window the number is a share, not a class: it rises steadily with the LED’s share of the light and crosses the line at about 62 per cent.

A camera, a second sensor, and a map between them

The camera is the modelled silicon sensor — colour dyes on silicon behind an infrared-cut filter — whose profile work these essays have used throughout. The ambient-light sensor is modelled rather than measured, because no measured pair of sensitivities for a real phone’s two sensors is in hand here. It is bare silicon under gaussian filters with no infrared cut, in three designs a phone might carry: red, green, blue and clear channels; red, green and blue alone; and a clear channel with one photopic-like channel of the kind a lux meter uses.

The statistic is a prediction and its failure. A linear map from the camera’s white to the sensor’s reading is fitted over smooth lights only — radiators from 2500 to 6500 K and daylights from 4000 to 10000 K — so it learns how the two instruments relate on spectra without structure. A lamp’s disagreement is the relative residual when that map predicts the sensor’s reading from the camera’s. None of the fourteen lamps scored is among the lights the map was fitted on, and the fitted lights themselves leave a residual of at most 0.011.

The threshold sits at the geometric middle of the gap between the classes, which is the most favourable place it can go, since it was placed using the lamps it is scored on. Everything this essay finds wrong is wrong at that flattering threshold.

The fourteen lamps

How far each lamp's sensor reading is from what the camera predicts, with RGB + clearFourteen lamps, six smooth and eight structured, each scored by how far an ambient-light sensor with red, green, blue and clear channels reads from what the camera's white predicts through a map fitted on smooth radiators and daylights. The smooth lamps score at most 0.057 and the structured at least 0.133; the dashed line is the threshold at the gap's geometric middle, 0.087. The lights the map was fitted on score at most 0.0113. Every lamp falls on its own side of the line.thresholda 3000 K radiator0.002a 4000 K radiator0.0034000 K daylight0.0115000 K daylight0.005a radiator tinted green0.053a radiator tinted pink0.057a cool white LED0.143a neutral white LED0.139a warm white LED0.138an LED with a red phosphor0.226a halophosphate tube0.140a triphosphor tube0.195a broadband tube0.155a three-emitter source0.133relative disagreement between sensor and cameralight bars: smooth · dark bars: structureda camera and an ambient-light sensor, modelled
Fig. 1 Fourteen lamps scored by how far a four-channel ambient sensor reads from what the camera’s white predicts, with the threshold at the middle of the gap.

The figure above is the result as a classifier would report it.

The four smooth radiators and daylights score between 0.002 and 0.011, essentially the fit’s own residual. The two tinted radiators — a 4000 K radiator filtered towards green and one towards pink, both off the daylight locus where the map was never fitted — score 0.053 and 0.057, the highest of the smooth lamps and still well clear. The structured lamps start at 0.133 for the three-emitter source, run through the three phosphor LEDs and the halophosphate tube at about 0.14 and the broadband tube at 0.155, and end at 0.195 for the triphosphor tube and 0.226 for the LED with a red phosphor.

No threshold anywhere between 0.057 and 0.133 makes a mistake, which is the result flicker could not reach at any threshold. And the regret follows. Each of the twelve fixtures the flicker essay used carries a price for each matrix — the blend and a third matrix fitted under a white LED — which is a matrix is fitted under one light applied to a choice of two, and a classifier that is right everywhere hands every fixture its better matrix.

What each way of choosing a matrix costs over the twelve fixtures. Mean regret — how many colour differences more than the better of the two matrices each choice costs — over the twelve fixtures the flicker classifier was scored on. Choosing by the ambient sensor's disagreement costs 0.000; by flicker, 0.435; always using the third matrix, 0.419; always blending, 0.511. On these fixtures the sensor classifies every lamp correctly, so its regret is nothing — because on these twelve, deep red and spectral structure happen to go together.
Fig. 2 Mean regret over the twelve fixtures for four ways of choosing between the blended matrix and a third matrix fitted under a white LED.

Choosing by the sensor costs 0.000; choosing by flicker, 0.435; always using the third matrix, 0.419; always blending, 0.511. On the cases that exist, the prediction that a second sensor would work is borne out completely.

That is where a cross-tabulation would stop, and it is where the flicker essay’s habit said to stop — list the cases, mark each with its true class and with what the proxy says, and look at the off-diagonal. Here the off-diagonal is empty. The question left is whether it is empty because the proxy is the property, or because on these fourteen lamps the proxy and the property happen to agree.

Where the disagreement comes from

A residual has components, one per sensor channel, and they need not be equal.

The camera's three channels beside an ambient sensor with RGB + clear. Solid: the camera's red, green and blue sensitivities, silicon behind colour dyes and an infrared-cut filter, each scaled to the largest. Dashed: an ambient-light sensor's red, green, blue and clear channels on bare silicon with no infrared cut. Shaded is 680 to 780 nm. The camera's red channel is at 45 per cent of its peak at 660 nm, 7 at 680 and under 1 at 700; at 740 nm the clear channel is at 100 per cent of its own peak. No combination of the camera's channels can predict what the clear channel collects there.
Fig. 3 The camera’s three channel sensitivities beside the ambient sensor’s four, with 680 to 780 nm shaded.

Over the eight structured lamps, 59 per cent of the squared disagreement is in the clear channel, 25 in green, 9 in blue and 7 in red. The clear channel is the one with no counterpart in the camera. The camera’s red channel is at 45 per cent of its peak at 660 nm, 7 per cent at 680 and under 1 per cent at 700, because the infrared-cut filter in front of every camera sensor is there to stop exactly that light. Bare silicon is near its peak there. Whatever a lamp emits between 680 and 780 nm reaches the ambient sensor’s clear channel and never reaches the camera at all.

So the map from camera to sensor has to guess the deep red, and it guesses it the way the smooth lights it was fitted on taught it: from the slope of the red end of the spectrum the camera can see. A radiator or a daylight continues smoothly past 680 nm and the guess is good. A phosphor LED or a fluorescent tube does not — its red falls away before the deep red, and on every one of the eight structured lamps the clear channel collects less than the camera’s red predicts. The shortfall is largest for the LED with a red phosphor, whose extra red near 650 nm raises the camera’s red and so the prediction, while adding little past 680.

The fourteen lamps sort by structure because, among them, the structured lamps are also the lamps with less deep red than their visible red implies. That is a property of how white LEDs and tubes happen to be made, not of spectral lines.

Lamps built to separate the two

A proxy that agrees with the property on every case available can be tested only by making cases where they come apart. Deep red and spectral structure are independent properties, so each can be changed without the other.

Deep red moves a lamp across the line, and nothing about its lines does. Left: a 3000 K radiator and 5000 K daylight with their light above a stated wavelength rolled off, scored by an ambient sensor with red, green, blue and clear channels. Both start far below the threshold, 0.087, and both cross it once the cut reaches 700 nm or below; cut at 660 nm they score 0.290 and 0.168. Right: a neutral white LED and a three-emitter source with a band of deep red at 720 nm added, as a share of each spectrum's mean. Both start above the threshold and both end below it, at 0.086 and 0.068. Neither change adds or removes a line: the smooth lamps stay smooth and the structured lamps keep every peak.
Fig. 4 Left, two smooth lamps with their light above a stated wavelength rolled off; right, two structured lamps with a band of far red at 720 nm added, both scored against the threshold.

Roll off a 3000 K radiator’s light above 720 nm and it scores 0.116, above the threshold; daylight crosses at 700 nm. Cut both at 660 nm and they score 0.290 and 0.168, higher than most of the real LEDs. Neither spectrum has acquired a line. Each is exactly as smooth as it was, and the luminance each lost is under one per cent.

Add a band of far red at 720 nm to a neutral white LED, carrying about a sixth of its power and about a tenth of one per cent of its luminance, and it scores 0.086, just on the smooth side of the line; a three-emitter source with the same addition scores 0.068. Neither has lost a line. The LED still has its blue pump and its phosphor hump, and the three-emitter source still has three narrow peaks, and each is now classified as a lamp the blended matrix suits.

Both constructions have ordinary counterparts. Horticultural and some full-spectrum fixtures add a far-red emitter near 730 nm to white LEDs, which is the right-hand panel. A glass or diffuser that trims the far red from daylight is the left, and whether a particular glazing trims enough is a measurement on that glazing rather than something this model can say. In both, the classifier is confidently wrong in the direction that costs: a structured lamp handed the blend, a smooth one handed the LED matrix.

That is the difference between this failure and flicker’s. Flicker was wrong on lamps that exist in every kitchen. This classifier is right on every lamp of that kind and wrong on a class that is rarer but not exotic, and nothing in its own output distinguishes the two.

The margin is the clear channel’s

If the clear channel carries most of the separation, a sensor without one should do worse, and the three designs measure how much.

Three sensor designs: the gap between the classes, and how much calibration error it survives. For each of three ambient-sensor designs, the highest score among the six smooth lamps and the lowest among the eight structured, with the gap between them. RGB + clear: 0.057 against 0.133, a gap of 0.076, first misclassifying a lamp when its channels are each 5.0 per cent out; RGB only: 0.011 against 0.035, a gap of 0.024, first misclassifying a lamp when its channels are each 1.0 per cent out; clear + lux: 0.024 against 0.102, a gap of 0.078, first misclassifying a lamp when its channels are each 5.0 per cent out. Every design separates the fourteen lamps. The two with a clear channel do it with three times the room, and the room is what calibration error spends.
Fig. 5 For each of three sensor designs, the highest-scoring smooth lamp and the lowest-scoring structured one.

All three designs separate the fourteen lamps. With red, green, blue and clear the gap runs from 0.057 to 0.133; with a clear channel and a photopic-like one, from 0.024 to 0.102. With red, green and blue alone it runs from 0.011 to 0.035, a gap of 0.024 against 0.076 and 0.078.

The design without a clear channel is reading something closer to structure. Adding far red to the LED and the three-emitter source leaves both on the structured side, at 0.047 and 0.027 against its threshold of 0.019, because none of its channels reaches 720 nm. It is not immune — a smooth lamp cut at 660 nm still crosses its line, because its red filter reaches into that range — but its errors are no longer about light the camera cannot see.

And it pays for that with its margin. The clear channel does two things at once: it widens the gap between the classes by a factor of three, and it makes the gap a measurement of deep red. The design choice that makes the classifier robust is the design choice that makes it read the wrong property.

What calibration error spends

The essay that proposed this classifier predicted fragility, on the reasoning that its statistic is a difference of two calibrations. The margin is where that shows.

How many of the fourteen lamps a calibration error misclassifies. The worst case, over every combination of each sensor channel reading high, low or right by a stated fraction, of how many of the fourteen lamps land on the wrong side of the threshold. RGB + clear makes its first mistake at ±5.0 per cent; RGB only makes its first mistake at ±1.0 per cent; clear + lux makes its first mistake at ±5.0 per cent. The margin is spent by any error in the sensor's channels relative to the camera's — between units, with temperature, or with age — so the classifier is only as good as the calibration that holds the two instruments together.
Fig. 6 The worst number of the fourteen lamps misclassified, over every combination of each sensor channel reading high, low or right by a stated fraction.

The four-channel design and the clear-and-lux design make no mistake until each channel’s calibration is 5 per cent out in the worst combination; at 7.5 per cent the four-channel design gets five lamps wrong and at 10 per cent seven. The red, green and blue design makes its first mistake at 1 per cent, and at 3 per cent gets nine of the fourteen wrong.

A per-channel gain error is the simplest thing that happens to a sensor between the factory and a photograph: unit-to-unit variation in filter thickness, the dark current’s dependence on temperature, and the slow change of a dye with light and age all look like one. Whether a production ambient sensor holds its channels relative to a production camera to five per cent across all of those is a property of parts that are not modelled here. It is a much looser requirement than one per cent, and it is the requirement the design reading the right property imposes.

So the prediction of fragility is half right. The design that would classify by structure is fragile. The design that is not fragile classifies by something else.

A room is a mixture

The fourteen lamps were scored one at a time, and a room with two lights has no white. The sensor, mounted on the front of a phone, reads the whole room.

A room lit by a white LED and a window, scored as the LED's share rises. The disagreement for a room lit by a neutral white LED and 5000 K daylight, at shares of the room's visible power from the LED running from none to all, with red, green, blue and clear channels. It rises steadily, from 0.005 to 0.139, and crosses the threshold at about 62 per cent LED. A classifier reading this number is not deciding whether a structured lamp is present; it is deciding whether the structured lamp is most of the light, which is a different question with a different answer for the colour matrix.
Fig. 7 The disagreement for a room lit by a neutral white LED and 5000 K daylight, as the LED’s share of the light runs from none to all.

The score rises at every step, from 0.005 with daylight alone to 0.139 with the LED alone, and crosses the threshold at about 62 per cent LED. A room lit mostly by the window is called smooth and handed the blend; a room lit mostly by the LED is called structured and handed the LED matrix.

That is a sensible rule if the colour matrix’s error in a mixed room falls between its errors under the two lamps in proportion to the shares. It is not the rule the classifier was designed to implement, which was to detect whether a structured lamp is present. And the camera may not be pointing at the part of the room the sensor reads: two lamps decide what one lamp could not found surfaces in a two-lamp room lit in different proportions, and a phone’s front sensor faces the photographer rather than the subject.

How the instruments were modelled

The camera sensitivities are the silicon quantum efficiency times gaussian dye transmittances centred at 605, 535 and 455 nm times the infrared-cut filter. The ambient sensor’s channels are the same silicon quantum efficiency under gaussian filters at 615, 535 and 465 nm with widths of 40, 45 and 35 nm, a clear channel of silicon alone, and a photopic-like channel at 555 nm with a width of 45 nm. Lamps are the visible spectra from the two-matrix work, from 380 to 780 nm.

Each instrument’s reading is reduced to shares of its own channels, so that the lamp’s brightness cancels. The map is a least-squares fit, one row per sensor channel, from the camera’s three shares to the sensor’s, over seventeen radiators and thirteen daylights. The disagreement is the length of the prediction’s error divided by the length of the reading.

A calibration error multiplies each sensor channel by one plus, minus or zero times the stated fraction, and every combination is tried, which is 81 for four channels and 27 for three. The mixture is the sum of the two lamps’ spectra, each normalised to equal power across the visible band, in the stated proportion.

What this leaves out

The sensitivities are modelled, and the result depends on one feature of them: that the ambient sensor has a channel reaching past the camera’s infrared cut. Real ambient sensors are built that way, often with an explicit infrared channel beside the clear one, but the numbers here are for gaussian filters on a generic silicon curve.

The spectra stop at 780 nm. Real radiators and daylight emit far more beyond that than inside it, and a real clear channel collects some of it. That would move smooth thermal lamps further from the camera’s prediction rather than closer, and an ambient sensor with an explicit infrared channel sorts lamps into thermal and non-thermal — a third property, which again coincides with structure on most lamps that exist and not on a filtered daylight or an LED with a near-infrared emitter.

The map is linear and fitted on one family. A map fitted with some structured lamps in it would reduce their residuals and the gap with them; a nonlinear map could do better or worse. The flattering threshold is a best case and a deployed classifier would sit somewhere worse.

And the camera has its own metamers: two lamps the camera records identically may differ to the sensor for reasons that have nothing to do with either class. The scoring here asks only whether the fourteen lamps sort; it does not ask how many different lamps share each score.

Still open: whether a third channel can tell deep red from structure

The design without a clear channel read something nearer to structure and paid for it in margin. The design with one bought margin and read deep red. The natural question is whether a sensor can have both — a margin from somewhere other than the region the camera cannot see.

The candidate is a channel placed inside the camera’s range but between its peaks, where phosphor LEDs and tubes have their characteristic dips: near 480 nm, between the blue pump and the phosphor, and roughly 570 to 600 nm for tubes. A narrow channel there sees structure directly, and a smooth lamp cut in the deep red does not move it.

The calculation is the same one run with a fourth design and the constructed lamps included in the test. It should report three numbers: the margin on the fourteen lamps, the calibration error at which it first fails, and whether the far-red and trimmed-red constructions still cross the line. If the margin survives without the clear channel, the proxy has become the property. If it does not, the choice between a robust classifier and a correct one is a choice a sensor designer makes rather than one arithmetic removes.

A cross-tabulation passes on the cases that exist

The habit is about what a clean result on real cases can and cannot show.

The flicker essay’s move was to list the cases and look at the off-diagonal before building anything, and it found the flaw at once. The same move here finds nothing: fourteen lamps, no mistakes, no regret. The temptation is to read an empty off-diagonal as proof that the proxy measures the property.

It shows only that the two agree on the cases listed. When two properties are correlated in the population of things that get built — structured lamps also lack deep red, because of how phosphors are chosen — every real case agrees and the proxy looks perfect.

The move is to construct the cases the population does not contain: change the proxy without changing the property, and the property without the proxy, and see which one the classifier follows. Here that took two filters and an added band, and it said at once what the fourteen lamps could not.

The failure mode is to stop at the empty off-diagonal. A classifier validated only on the cases that exist is validated on the correlation that made them, and it inherits every lamp built differently.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationCamera profileCamera sensitivityColour matrixFluorescentIdentifiabilityInfraredLED emissionObserver metamerismSpectral structure