What a camera does

The converter can choose except where it matters

A converter rebuilding a clipped highlight has to say whether the surface was matt and over-exposed or glossy and carrying a reflection, and the evidence is whether its raw chromaticity slides towards the lamp's on the way up. Read through the sensor's own noise, the slide is clear on most of the chart from the few tens of pixels a specular highlight holds — and on the surfaces where it takes half a frame, confusing the two models costs more than the median. A ratio cannot see a common scale, and an exposure supplies one.

Assumes Filling in a highlight is a claim about the surface, Clipped noise does not average away and The highlight is the white balance.

Filling in a highlight is a claim about the surface showed that a converter rebuilding a clipped channel has to decide what the highlight was. A matt surface over-exposed keeps its own ratio between channels; a glossy one carries a reflection of the lamp on top of its body colour, so its raw chromaticity slides towards the lamp’s as the reflection strengthens. It ended by saying the choice rests on seeing whether the slide is there in real pixels, and that how reliably it can be measured through noise is a question about images rather than about the arithmetic.

It is, and the arithmetic of the answer is a signal-to-noise ratio.

The statistic a converter would read, under each modelThe log ratio of the two channels that are still open, against how bright the surface is, under the two models of what a highlight is. A matt surface over-exposed keeps its own ratio exactly — the line is flat, and it must be, because scaling every channel by the same amount leaves a ratio alone. A glossy surface carries a reflection of the lamp on top of its body colour, so its ratio slides towards the lamp's as the reflection strengthens: -0.086 of a log unit between a quarter of full scale and nine tenths. That slide is the whole of the evidence a converter has for choosing between them.0.250.500.751.000.430.510.59log of the open channels' ratiobrightness, as a share of full scalea glossy reflectiona matt over-exposuresurface 21a highlight through the sensor's noise
Fig. 1 The log ratio of the two channels that are still open, against how bright the surface is, under the two models. A matt over-exposure keeps its ratio exactly; a glossy reflection slides.

Readable on most of the chart, and not where it counts

The slide is clear on most surfaces from the few tens of pixels a specular highlight holds, and needs most of a frame on the surfaces where confusing the two models costs at least the median — so the estimator is never weakest on a cheap mistake.

  • The slide is 0.307 of a log unit between a quarter of full scale and nine tenths under the glossy model, and exactly zero under the matt one.
  • No single pixel decides anything: the best surface reaches a signal-to-noise of 0.86 on one reading, and most are near half.
  • The median surface needs 20 pixels averaged before the slide is two standard deviations clear of zero. The hardest needs 413,718 — most of a frame.
  • The hardest surface costs 8.34 colour differences to confuse, against a median of 6.42 over the chart.
  • What makes it hard is that the two models differ there by a common scale in the open channels — factors of 0.8052 and 0.8041 — and a ratio discards a common scale by construction, while an unknown exposure supplies one.

The statistic, and why it is a ratio

A converter reading a highlight has two channels that have not clipped and one that has. Which one clipped is decided by the lamp rather than the surface — it is the channel the lamp drives hardest — so the converter knows which pair it is reading.

What it can read from that pair is their ratio, and a ratio is the right statistic for a reason that is also its limitation. A matt surface over-exposed is its own body colour scaled, so every channel is multiplied by the same number and the ratio does not move — exactly, not approximately. A glossy one has a share of the lamp’s own raw white added on top, and the lamp’s ratio is not the surface’s, so the sum’s ratio moves towards the lamp’s as the share grows.

So the test is a test of one number against zero, and zero is what the matt model predicts with no free parameters at all. That is a clean hypothesis test and it is the reason the question has a numerical answer rather than an argument.

The noise on it is the sensor’s, and clipped noise does not average away is the standing caution about what that noise does once it meets a clamp — which is a different part of the same pipeline and does not bear on a statistic read below the clip. A log ratio’s variance is the sum of the two channels’ relative variances, and a relative variance is the noise over the signal — so the estimator is noisiest in the channel carrying least light, which at a highlight is the one furthest from the lamp’s own colour.

How many pixels each surface needs before the two models can be told apart. Each of the chart's twenty-four surfaces, with the number of pixels that must be averaged before the slide is two standard deviations clear of zero, on a logarithmic scale. The median is 10 pixels, which a small specular highlight already holds. The worst is 139,574, which is most of a frame. The spread is four orders of magnitude over twenty-four surfaces of one chart, and it is not noise: it is the surfaces whose body colour already sits where the lamp's does in the two channels that are still open.
Fig. 2 Each of the chart’s twenty-four surfaces, with the number of pixels that must be averaged before the slide is two standard deviations clear of zero.

The median surface needs 20 pixels and the spread is four orders of magnitude, which is the shape a mean is not a worst case is about, from 6 to 413,718. That spread is not noise in the measurement; it is the surfaces themselves, and which surfaces are which has a cause.

Which channel clips, and why that is the lamp’s decision

Before anything can be read, something has to have clipped, and which channel it is is not a free variable.

A raw sensor’s three channels reach their ceiling at different signals because the lamp does not drive them equally. Under an incandescent lamp the blue channel is starved and the red is not, so the red clips first on almost every surface; under daylight the balance is different. The highlight is the white balance is where that observation entered this collection, and it does two things here.

It fixes which pair the statistic reads, which is useful: the converter does not have to discover the pair, it can compute it from the white it has already estimated. And it explains why the difficulty pattern is a property of the lamp. The pair the statistic reads is the lamp’s choice, so which surfaces are parallel to the lamp within that pair is the lamp’s choice too — which is why the rank correlation between difficulty and cost is +0.52 under one lamp and −0.18 under another, on the same chart with the same sensor.

The estimator’s blind spot is therefore a moving one, and it moves with the thing the converter is trying to correct for. That is worth knowing before anybody builds a table of hard surfaces: the table would be a table for one lamp.

Where the estimator fails, and what it costs there

The four orders of magnitude would be a curiosity if the hard surfaces were the ones where the decision did not matter. They are not.

The surfaces the statistic cannot read are the expensive ones. Each surface by how many pixels its slide needs across, and by what confusing the two models costs it up. The two are positively correlated — a rank correlation of 0.52 — which is the worst possible arrangement: the estimator is least able exactly where the mistake is most expensive. The hardest surface needs 413,718 pixels and costs 8.34 colour differences; the easiest needs 6 and costs 2.79. Nothing about the estimator can be improved to change that, because what it is short of is not precision.
Fig. 3 Each surface by how many pixels its slide needs across, and by what confusing the two models costs it up.

The hardest surface needs 413,718 pixels and costs 8.34 colour differences to get wrong; the easiest needs 6 and costs 2.79. Under an incandescent lamp the two are positively correlated at a rank correlation of 0.52.

That correlation is not the durable part and it is worth saying so before leaning on it. Under daylight it is −0.18: the relationship between difficulty and cost across the whole chart is a property of the lamp, because which channel clips and how starved it is are both the lamp’s doing. What holds under both lamps is the part that matters — the hardest surface’s cost is at least the median surface’s, so the estimator is never weakest on a cheap mistake.

A ratio cannot see a common scale

The mechanism is worth getting exactly right, because the obvious explanation is wrong.

The obvious explanation is that the two models nearly agree on the hard surfaces, so the statistic has little to find and little is at stake. The costs say otherwise: 8.34 colour differences is a large disagreement. So the two models differ substantially there and the statistic cannot see the difference, which means the difference is of a kind the statistic discards.

What the two models actually differ by in the open channels. Three surfaces — the easiest, a middling one and the hardest — with the factor by which the glossy model's reading exceeds the matt model's in each of the two channels that are still open. On the easiest surface the two factors are 0.967 and 0.697, plainly different, and their ratio is what the statistic reads. On the hardest they are 0.9375 and 0.9389 — the same number. The two models differ there by a common scale, a ratio discards a common scale by construction, and an unknown exposure supplies one anyway, so the two are confounded rather than merely hard to tell apart.
Fig. 4 Three surfaces — the easiest, a middling one and the hardest — with the factor by which the glossy model’s reading exceeds the matt model’s in each of the two open channels.

On the easiest surface the two factors are 0.873 and 0.628. Two different numbers, so their ratio is not one, so the ratio statistic reads the difference. On the hardest they are 0.8052 and 0.8041 — the same number to three places.

So on that surface the two models differ in the open pair by a common scale: the glossy reading is four fifths of the matt reading in both channels. A ratio discards a common scale by construction; that is what a ratio is for. And the difference is nonetheless real, because the clipped channel does not scale with them and the reconstruction depends on where the open pair sits absolutely.

The converter could in principle read the absolute level instead — and that is exactly what an unknown exposure supplies. A matt surface photographed a stop brighter and a glossy surface with a reflection produce the same brighter open pair. The two models are not merely hard to separate there; they are confounded with the exposure, which is a parameter the converter does not know either.

That is the same shape an image does not determine the light found one stage up, and it is why the answer is a statement about identifiability rather than about precision: no better estimator helps, because the information is not in the pixels.

Light is what is short, not gain

A signal-to-noise problem invites the response of turning something up, and it is worth being precise about which thing.

Light is what the estimator is short of. The pixels needed for the median surface, the easiest and the hardest, at three sensor states. A low amplification needs 1 pixels at the median and a very high one 43, because the statistic's noise is the photon count's and an amplification adds no photons. The hardest surface runs from 10,961 to 573,931 over the same range, which is the difference between a large region and more pixels than the frame has. The scale is logarithmic.
Fig. 5 The pixels needed for the median surface and for the hardest, at three sensor states.

The median surface needs 2 pixels at a low amplification and 95 at a very high one; the hardest runs from 31,503 to 1,776,452 over the same range. The statistic’s noise is the photon count’s, and raising the amplification adds no photons — it only makes the read noise smaller relative to a signal that is already small.

So the dial that helps is exposure, and the dial that does not is gain. A highlight is by definition the brightest part of the picture, so a converter reading a highlight is already reading the most photons it will get; the surfaces that stay hard at a low amplification stay hard everywhere, which is what the sweep shows.

How many pixels each surface needs before the two models can be told apart. Each of the chart's twenty-four surfaces, with the number of pixels that must be averaged before the slide is two standard deviations clear of zero, on a logarithmic scale. The median is 1 pixels, which a small specular highlight already holds. The worst is 10,961, which is most of a frame. The spread is four orders of magnitude over twenty-four surfaces of one chart, and it is not noise: it is the surfaces whose body colour already sits where the lamp's does in the two channels that are still open.
Fig. 6 The same twenty-four surfaces at a low amplification, where a large sensor gathers fifty times the photons. The shape of the curve is unchanged and the whole of it has moved down by about a decade: the hard surfaces are hard for a reason light does not fix.

What a converter could actually do

What a converter could average over, and whether it is enough. Four regions a converter might average the statistic over, with the signal-to-noise each reaches on the median surface. Two standard deviations is the line. One pixel reaches 0.78 and decides nothing; a small specular highlight reaches 6.0 and decides the median surface comfortably; a whole frame reaches 3475. So a converter can make this choice per surface on most of a picture — which is more than the earlier essay expected and less than it needs, because the surfaces it still cannot decide are the ones that cost most to get wrong.
Fig. 7 Four regions a converter might average the statistic over, with the signal-to-noise each reaches on the median surface. Two standard deviations is the line.

Correcting colour costs noise sets the scale for what a camera’s noise is worth downstream; here it is the whole budget.

The thresholds matter less than the ordering. Two standard deviations is a convention and a converter weighting the two fills by their posterior probability would not use a threshold at all; what survives either arrangement is that a highlight’s own pixels carry enough evidence on most surfaces and nowhere near enough on a few, and that the few are not the cheap ones.

One pixel reaches 0.50 and decides nothing. A small specular highlight of sixty pixels reaches 3.7 and decides the median surface comfortably. A face-sized region reaches 67 and a whole frame 2,117.

So the answer to the question the earlier essay asked is better than it expected in one respect and worse in another. A converter can choose its fill model per surface on most of a picture, from the pixels the highlight itself holds, which is more than a per-picture decision and much more than a fixed rule. And on the surfaces it cannot decide, it cannot decide at any region size short of the whole frame, and averaging the whole frame answers a different question — one decision for a picture containing surfaces that would individually have answered differently.

The practical arrangement that follows is a three-way one rather than a two-way one. Read the slide; act on it where it is significant; and where it is not, say so — use the fill the two models agree about most closely, or carry the clipped values and let a later stage decide, rather than committing to a model the pixels do not support. The site’s habit for this is the one a model is a claim about what can be known states: the honest output of an estimator that cannot decide is a refusal rather than a guess.

How the readings and the noise were computed

The camera is the modelled silicon sensor with its infrared-cut filter, and the surfaces are the chart’s twenty-four at chroma 0.45, each reduced to its raw ratio under the lamp. A matt reading at a stated brightness is that ratio scaled so the brightest channel reaches the stated share of full scale; a glossy one is a fixed body term with a share of the lamp’s own raw white added, normalised the same way, so that the two models are compared at equal exposure rather than at equal parameter.

The noise is the sensor model used here for the clipped-noise work: a read noise in electrons and a full-well count, combined as the root of the read noise squared plus the signal, all as shares of the largest channel’s white. Three sensor states bracket a large sensor at a low amplification and a small one pushed hard.

The pixel count is the square of the ratio of the statistic’s standard deviation to half the slide, at two standard deviations — which is the count at which a test of the observed slide against zero separates the two models at that confidence. The cost of confusing them is the colour difference between the two models’ full-scale readings put through the camera’s matrix at equal lightness, computed independently of the statistic that would make the choice.

What this leaves out

The mosaic is not modelled, although a grey edge arrives coloured is what it does at an edge and a highlight’s boundary is one. A real sensor reads one channel per photosite, so the two channels of the ratio come from different pixels and are demosaiced before anything else — which correlates their noise and costs some resolution, and both of those make the estimator worse rather than better.

The surfaces are the chart’s, which is a constructed family spanning hue at one chroma. A real scene’s highlights sit on skin, on paint, on plastic and on metal, and a metal’s highlight is not dichromatic at all: its reflection carries the metal’s own colour, which is a third model neither of these two covers. A surface that is not a multiplication takes that apart.

The two models are the two the earlier essay compared. A converter that fitted the reflection’s strength as a free parameter rather than choosing between two rules would have a different estimator with a different noise, and whether the confounding with exposure survives that is the same calculation on a larger model.

And the count is a threshold at two standard deviations. A converter does not have to decide at a fixed confidence; it could weight the two fills by their posterior probability, which is a better arrangement and does not change where the information is.

Still open: whether the confounding survives a second lamp

The hardest surfaces are hard because the glossy and matt readings differ there by a common scale in the open pair, and a scale is what an unknown exposure also supplies. That is a confounding between two parameters, and confoundings of that shape are usually broken by a second observation with a different geometry.

A photograph taken under two lamps at once is such an observation and is the ordinary indoor case. Two lamps put two different whites into the reflection, so the glossy model’s addition is no longer a single direction, and a surface whose body is parallel to one lamp’s white in the open pair need not be parallel to the other’s. The prediction is that the hard surfaces change identity rather than disappearing — each lamp has its own set — and that a scene lit by two lamps has fewer surfaces hard for both than for either.

The measurement is the same one run with a mixture, and the interesting output is not the median but the overlap: how many surfaces are above the pixel budget under both lamps at once. If the answer is near zero, a converter in a real room has a per-surface decision everywhere and the problem this essay describes is a property of a single-lamp studio.

An estimator’s failures are a set, and the set has a shape

The habit is about what to report when a measurement works sometimes.

The natural summary is a rate — how often the estimator succeeds, or the median case — and it is the wrong one whenever the successes and the failures are not interchangeable. What matters is whether the failures fall where the answer is cheap or where it is expensive, and that is a question about the joint distribution rather than about either margin.

The move is to compute the cost of being wrong independently of the statistic that decides, and then to plot one against the other. Independently is the load-bearing word: a cost derived from the same quantity the estimator reads will be correlated with it by construction, and the plot will show a relationship that is an artefact.

The failure mode is to report a median. Twenty pixels is the median here and it is a true and useless number, because the surfaces that need twenty are not the surfaces anybody would worry about.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Camera rawCamera sensorClippingDichromatic reflectionHighlight recoveryIdentifiabilityNoiseShot noiseSignal-to-noiseSpecular