What a camera does

One step has no choice

Four of a raw converter's operations can be arranged twenty-four ways. The reconstruction cannot be arranged at all — a colour matrix needs three numbers and a mosaic site has one, so filling in the mosaic is forced to the front by arithmetic rather than by convention. What is not forced is whether it happens in linear light or after the curve, and that decision costs 5.8 colour differences at an ordinary edge and nothing at all four sites away from it.

Assumes The order is not in the documentation, A grey edge arrives coloured and Raw is not a picture.

The previous rung counted twenty-four arrangements of four operations and found five outcomes. There is a fifth operation and it is not in that count, because it has no arrangements: the reconstruction that turns a mosaic of single measurements into three numbers per pixel is pinned to the front of the chain by arithmetic.

Where the mosaic is filled in, along one row through an edge. A Bayer row across a step from 0.9 to 0.08, in units of the sensor's own ceiling, reconstructed in linear light and reconstructed after the tone curve, with the second undone so the two are compared at the same point in the chain. Away from the edge they agree to 8.3e-14, because a constant interpolates to itself under any curve. At the edge they differ by 5.78 colour differences. Interpolating encoded values pulls an edge towards its dark side.
Fig. 1 A Bayer row across an edge, reconstructed in linear light and reconstructed after the tone curve. Away from the edge the two agree to the floating-point floor; at the edge they part by 5.8 colour differences.

The claim

The reconstruction’s place in the chain is decided by arithmetic and not by anybody, and the one freedom it has left costs 5.8 colour differences at an edge and exactly nothing away from one.

  • A colour matrix needs three numbers and a mosaic site has one, so nothing that mixes channels can precede the reconstruction. It is the only step in the pipeline whose position is forced.
  • Whether it happens in linear light or in the encoded variable is not forced, and both are implemented.
  • The two differ by 5.78 colour differences at the edge of a step from 0.9 to 0.08, and by 6 × 10⁻¹⁵ four sites away.
  • The cost scales with the edge’s contrast: 0.73 at a step to 0.5, 5.78 at 0.9, 14.33 at 1.35 where the bright side is over the ceiling.

Why it cannot be anywhere else

A three-by-three matrix takes a vector of three numbers to a vector of three numbers. That is not a stylistic preference; it is what the object is.

A sensor site behind a colour filter array has measured one number. It has an R value, or a G value, or a B value, and no information at all about the other two — the filter over it absorbed them. So there is nothing for a matrix to be applied to, and any operation that mixes channels is undefined until the missing two have been supplied from somewhere.

The four steps, and how much freedom each has. The arrangement every converter's documentation describes. The reconstruction is not among them because it has no choice: a colour matrix needs three numbers and a mosaic site has one, so filling in the mosaic is forced to the front by arithmetic. Of the four shown, the clip can move to 3 places without changing anything and the other three cannot move at all. The worst rearrangement lands 65 colour differences away.
Fig. 2 The four steps the previous rung permuted, with the reconstruction named as the thing that is not among them. Its absence from the list is the finding rather than an omission.

Supplying them is the reconstruction. It is an interpolation, it is an assumption about the scene wearing the costume of a measurement, and this collection has already measured what it does to a black-and-white edge: the naive version gives an achromatic scene a chroma of a hundred and forty-four per cent of its own local mean, and the standard repair — interpolating colour differences rather than channels — removes most of that.

None of that is in question here. What is in question is a step that comes after the reconstruction has been decided: at what point in the tone chain does it run.

The freedom that remains

A converter can reconstruct in linear light, which is what the sensor measured, or it can apply the tone curve to the sampled values first and reconstruct the encoded ones. Both are done, and the second is not a mistake — it is what any pipeline that treats the raw values as an image and applies its curve early ends up doing, and it has a real advantage: interpolating in a perceptually flatter variable produces edges that look better in the shadows, which is what the encoding was shaped for.

The measurement makes the two comparable by undoing the curve after the second reconstruction, so that what is compared is where the averaging happened rather than what came after it.

The curve, and what it does to a mid grey. The tone curve used throughout these essays: a smooth S applied in the encoded variable, at strength 0.70, plotted here against linear light. The diagonal is the identity. Its slope through the middle is 1.34 against 0.50 at the toe, which is what makes it a contrast control: the middle of the range is stretched and both ends are compressed. An eighteen per cent grey comes out at 16.9 per cent. Everything measured here follows from the curve being applied to each channel separately.
Fig. 3 The curve in question. Its slope varies by a factor of a few across the range, and an interpolation is an average — so an average taken through it is not the average of the same two values taken before it.

The mechanism is the round’s central one arriving in a camera. An average and a nonlinearity do not commute: the mean of two curved values is not the curve of their mean, and the gap is set by how much the curve bends between them — which is the same statement a mixture line makes. Where two neighbouring sites carry nearly the same value the gap is nothing; where they straddle an edge it is the whole of the curve’s bend across that step.

Why it is invisible away from the edge

The figure’s most useful feature is the flat part.

Four sample sites away from the edge the two reconstructions agree to 6 × 10⁻¹⁵ — floating-point noise, which is to say exactly. That is not an approximation and it has a one-line reason: a constant interpolates to itself under any monotone function whatever, because the average of two equal values is that value and the curve of it is the curve of it.

So the entire cost of this decision lives on edges, and on nothing else. A frame of smooth gradients is reconstructed identically either way. A frame of foliage against sky is not.

Where the mosaic is filled in, along one row through an edge. A Bayer row across a step from 0.5 to 0.08, in units of the sensor's own ceiling, reconstructed in linear light and reconstructed after the tone curve, with the second undone so the two are compared at the same point in the chain. Away from the edge they agree to 8.2e-14, because a constant interpolates to itself under any curve. At the edge they differ by 0.73 colour differences. Interpolating encoded values pulls an edge towards its dark side.
Fig. 4 The same measurement at a gentler edge — a step to half the ceiling rather than to nine tenths. The cost falls to 0.73 colour differences, because the curve bends less across a shorter interval.

The scaling with contrast is what the mechanism predicts. A step from 0.5 to 0.08 crosses less of the curve than a step from 0.9 to 0.08, so the discrepancy falls from 5.78 to 0.73 — a factor of eight for a factor of about two in the step’s size, which is the curvature entering quadratically.

Where the mosaic is filled in, along one row through an edge. A Bayer row across a step from 1.35 to 0.08, in units of the sensor's own ceiling, reconstructed in linear light and reconstructed after the tone curve, with the second undone so the two are compared at the same point in the chain. Away from the edge they agree to 7.2e-2, because a constant interpolates to itself under any curve. At the edge they differ by 14.33 colour differences. Interpolating encoded values pulls an edge towards its dark side.
Fig. 5 And at a step whose bright side is over the sensor’s ceiling. The edge cost rises to 14.33, and for the first time the flat region is not flat either, because both reconstructions are being clipped rather than merely curved.

Beyond the ceiling the argument changes character. A clipped value is not a curved value: the curve is invertible and the clip is not, so undoing the curve cannot undo the clip and the two reconstructions no longer differ only in where the averaging happened. The flat region departs from exactness by 0.072, which is small and is not zero, and its being non-zero is the signature of the clip rather than of the curve.

What this costs a picture

Translating an edge measurement into a statement about photographs needs care, because a frame is mostly not edges.

The mean over the whole row is 0.226 colour differences, against 5.78 at the worst site. On a real image the fraction of pixels within two sites of a high-contrast edge is small — a few per cent for most subjects, much more for foliage, text or a picket fence. So the average cost is negligible and the worst case is not, and which of those a viewer notices depends on whether the edges are where they are looking, which they usually are.

Three exchanges, two of which move the colour. The documented pipeline is a white balance, a colour matrix, a tone curve and a clip. Each bar is what happens when two neighbours change places, over 30 surfaces the modelled sensor captures: the filled bar is the mean and the tick is the worst patch. Exchanging the balance and the matrix costs 9.2 colour differences at the mean and 13.2 at the worst. Exchanging the curve and the clip costs exactly nothing, and that is a theorem rather than a small number: a monotone curve onto the unit interval commutes with clamping to it.
Fig. 6 The previous rung’s three exchanges, for scale. The reconstruction’s freedom costs more at an edge than the matrix-and-curve exchange costs anywhere, and less than the balance-and-matrix exchange costs everywhere.

Set beside the four permutable steps, the reconstruction’s freedom sits between them: worse at an edge than the live disagreement between competent converters, better everywhere else. It is a local defect rather than a global one, which is the opposite of the arrangement problem and needs a different kind of report.

The two domains, argued rather than measured

Before the measurement is used it is worth setting out the case each side makes, because both are good and the essay’s number is the price of the disagreement rather than a verdict.

The case for linear light is that the sensor measured linear light and interpolation is an averaging operation, so averaging the measured quantity is averaging the thing that was measured. A half-covered photosite receives half the photons; a pixel between two sites should carry the value the site would have measured if it had been there, and that value is the linear mean. Every physical argument points this way.

The case for the encoded variable is about noise and about what an interpolation is for. Sensor noise is roughly proportional to the square root of the signal, so in linear light it is much larger in the highlights than in the shadows; in the encoded variable it is far more nearly uniform, and an interpolator that assumes uniform noise — which every linear one does — behaves better there. And an interpolation is not really trying to reconstruct the photon count. It is trying to reconstruct what the scene looked like between two samples, which is a perceptual quantity and is closer to uniform in the encoded variable too.

Neither argument is decisive and the measurement is the cost of the argument, not its resolution. On a smooth frame the choice is free. On an edge it is 5.8 colour differences, and both parties would recognise their own reasoning in that: the linear camp would call it an error and the encoded camp would call it the shadows being reconstructed properly.

The mean and the worst, and which to report

A defect confined to edges puts a familiar problem in front of whoever has to report it, and this collection has met it before under a different name.

The mean over the row is 0.226 colour differences and the worst site is 5.78 — a ratio of twenty-six. Reporting the mean is defensible and useless: it describes a frame in which the defect is diluted by every pixel that does not have the defect, and diluting a local error by a global count is the standard way to make it disappear.

Reporting the worst is defensible and alarming: it is attained on real pictures, at every high-contrast edge, in every frame that has one.

The honest report is both, with the fraction attached. On this row four sites of thirty-two exceed one colour difference, which is twelve and a half per cent — a figure that would be much smaller on a photograph of a face and much larger on a photograph of a fence. A mean is not a worst case, and this collection has been through that argument on adaptation numbers; a defect whose distribution is bimodal needs the fraction as well as the two ends.

What a fence looks like

The measurement’s structure suggests where to look for it in a real image, which is the useful form for anybody trying to see it.

The cost rises with the contrast of the edge and falls to nothing away from one, so the subject that maximises it is a scene made almost entirely of high-contrast edges at the sampling limit: bare branches against a bright sky, a wire fence, text, a striped shirt. Those are precisely the subjects on which demosaicing artefacts are already known to be worst, and the two effects are not the same one.

A reconstruction artefact is a false colour where the scene had none, which is the effect measured on a grey edge. This one is a colour shift where the scene did have colour and the reconstruction reported a different amount of it, and it is present even when the reconstruction is perfect in the sense that essay means.

That distinction matters for anybody trying to attribute what they see. Turning up a converter’s demosaic quality setting addresses the first and cannot address the second, because the second is not about which interpolation is used.

The step that also cannot move, and does

There is a second operation with a forced position and it is worth naming because converters routinely violate it.

The clip against the sensor’s ceiling is a physical fact about the measurement: a site that saturated has no more information, and every arrangement in which a converter’s own clip precedes that one is arithmetic on a value the sensor never recorded. Highlight reconstruction is precisely the business of guessing what that value was, from the channels that did not saturate, and it is a well-founded guess rather than a measurement.

All twenty-four arrangements of four steps. Every permutation of the four operations, ranked by how far its picture lands from the documented one. 3 of the 24 are identical to it — the clip is a no-op wherever it follows the curve, and it is also a no-op before the matrix on values already inside the range. The rest run out to 64.8 colour differences at the mean and 98.1 at the worst patch. Nothing in a converter's documentation says which of the twenty-four it implements.
Fig. 7 The twenty-four arrangements again. The three that are identical to the documented one are identical only because nothing in the set is over the ceiling; a set with clipped highlights in it would separate them.

That has a consequence for the previous rung’s census. Its three identical arrangements are identical because every surface in the fit set is inside the range — the clip is a no-op wherever it is put. Repeat the census on a set with clipped highlights and those three separate, which is the subject two rungs further on.

Why a forced step is the dangerous kind

There is a reason to spend an essay on an operation that cannot be moved, and it is not the operation.

A step whose position is a decision gets documented eventually, because two implementations disagree and somebody has to adjudicate. A step whose position is forced never gets documented at all, because nobody ever had to choose — and what does not get documented does not get examined, so the assumptions it carries travel unexamined with it.

The reconstruction carries three. It assumes the scene’s hue varies more slowly than its luminance, which is what makes colour-difference interpolation work and is false at a coloured edge. It assumes the sampling lattice is the right one to average on, which is a statement about the optical low-pass filter in front of it. And it assumes the values it is averaging are the ones that should be averaged — which is the assumption this essay measures, and the only one of the three with a free parameter in it.

Being forced into first place is what let the third assumption go unnoticed. Everything after the reconstruction is a decision somebody makes, argues about and eventually writes down. The reconstruction is where the pipeline starts, and a pipeline’s first step reads as a given rather than as a choice.

What a converter could report, and what it could not

The specification question the previous rung left open takes a different form here.

Naming the arrangement is a four-word field and would settle the twenty-four. Naming the reconstruction’s domain is a one-word field — linear or encoded — and would settle this. Both are cheap and neither exists.

What could not be reported is the reconstruction itself. There are dozens of algorithms in use, most of them proprietary, several of them adaptive in ways that make them functions of the whole neighbourhood rather than of a fixed stencil, and no compact way to name one. So a converter’s output is not reproducible from its metadata even in principle, and the reconstruction is the reason.

That is worth separating from the rest of the round’s findings. Everything else measured here is a decision that could be recorded in a field. This one is a decision that could not be, and the honest statement about raw as an archival format is that it is archival up to the reconstruction and no further.

What was computed, and how

The row is thirty-two sites of a Bayer pattern alternating green and red, across a step at the midpoint. Each channel is interpolated linearly from its own samples, which is the naive reconstruction rather than the repaired one; using the colour-difference reconstruction changes the absolute numbers and not the comparison, because both runs use whichever reconstruction is chosen.

A grey edge, reconstructed from a Bayer row, arrives coloured. Above: an achromatic step through 24 sensor sites, with green sampled on the even ones and red on the odd. Interpolating each channel separately reconstructs them from data taken on either side of the edge, so their ratio moves. Below: the resulting chroma, peaking at 144 per cent of the local mean, and 112 per cent once colour differences are interpolated instead.
Fig. 8 The reconstruction itself, from an earlier round, with the colour-difference repair. This essay’s measurement sits on top of whichever of these is chosen; the choice between them is a different question with a much larger answer.

The curve-domain run applies the tone curve to the sampled values, interpolates, and then inverts the curve by bisection so that the two runs are compared at the same point in the chain. Without that inversion the comparison would be between a curved image and an uncurved one, which is a difference of tone rather than of reconstruction.

The matrix and the white balance are the previous rung’s, fitted on the same thirty surfaces, and the comparison is made in CIELAB after the whole chain, with the difference formula the round has used throughout.

Everything between the photons and the picture, and what each stage decides. The 8 stages of a camera pipeline. Only the second is physics; every one after it is a decision somebody made, and the reason two cameras pointed at the same scene disagree is that they made different ones.
Fig. 9 The pipeline with the mosaic highlighted, as this collection has drawn it since the imaging phase. The drawing puts the mosaic first, which is right, and does not say that it had to be.

Where the model stops

One dimension is enough for the mechanism and is not a real demosaic. A two-dimensional reconstruction uses neighbours in both directions and most modern ones are edge-directed, choosing an interpolation axis per pixel — which makes the operation nonlinear in its own right, and two nonlinearities compose in ways nothing here describes.

The row is noiseless. Real reconstruction decisions are made in the presence of noise, and a converter that reconstructs in the encoded variable is doing so partly because noise is more uniform there — correcting colour costs noise is the same trade seen from the matrix’s side — which is a real advantage this measurement does not price.

And the tone curve here is a fixed global curve. Every current converter applies a local one, whose slope varies with position as well as with value, and a local curve’s interaction with an interpolation is a subject this collection has no machinery for.

The generalisation

The habit is about distinguishing a constraint from a convention.

A pipeline diagram shows a sequence, and a reader cannot tell from the drawing which of the arrows are forced. Some are: the step needs an input the previous step is the only source of, or it needs a shape the previous step is the only producer of. Others are habit, or history, or the order somebody wrote the functions in.

The move is to ask, of each arrow, what would break if it were reversed. The answer is either nothing type-checks, which is a constraint, or the numbers change, which is a decision. Both are worth writing down and only the second is usually treated as interesting.

The failure mode is the reverse of the previous rung’s. There a decision was being read as a fact; here a fact is available to be read as a decision, and somebody attempting to reverse it produces code that runs and computes nothing. A matrix applied to one number is not an error anything reports — it is a matrix applied to whatever the other two channels happened to contain, which in a mosaic is whatever was left in the buffer.

Who found it, and when

Demosaicing has been a research subject since the Bayer patent of 1976 and the literature on it is enormous, almost all of it concerned with which interpolation to use rather than where in the chain to put it.

The domain question — linear or encoded — is discussed among converter authors and is usually framed as a matter of noise and of shadow quality rather than of colour accuracy. The colour cost measured here does not appear in that discussion, and the reason is probably that it is confined to edges, where every other reconstruction artefact also lives and where a colour error is hard to separate from an interpolation one.

Where the ladder goes next

The tone curve has now appeared in three measurements as the thing that makes an operation fail to commute. It deserves examining on its own terms, because a curve applied to each channel separately is a function of one number at a time and it nevertheless moves hue and raises chroma — so a contrast control is three controls, and two of them are unlabelled.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

ClippingColour managementDeclared inputInterpolationMeasurement errorSamplingSpatial frequencySpecificationStructural choiceTransfer function