What a camera does

Correcting colour costs noise

The matrix that turns a sensor's raw into XYZ has large off-diagonal terms of both signs, because that is what correcting a sensor which fails Luther's condition requires. Differences of large numbers are where noise grows, and the full correction amplifies photon noise by 1.66 times.

Assumes No matrix is right everywhere and A threshold is not a unit.

The colour matrix has been treated so far as the thing that fixes what it can. It is also the thing that makes something else worse, and the two are the same operation.

Correcting colour costs noise, and the two cannot both be least. Sweeping the colour matrix from its own diagonal — a white balance with no cross terms — to the full least-squares fit. Mean ΔE00 falls from 18.61 to 1.29; photon noise rises by a factor of 1.03. At 10000 photons per pixel.
Fig. 1 The sweep from a matrix with no cross terms to the full least-squares fit. Mean colour error falls; photon noise rises by a factor of 1.66. There is no setting on this axis where both are at their minimum, and that is not a manufacturing compromise.

The exposure the sweep is taken at is an argument, and it decides how much of the trade a photographer ever meets.

Correcting colour costs noise, and the two cannot both be least. Sweeping the colour matrix from its own diagonal — a white balance with no cross terms — to the full least-squares fit. Mean ΔE00 falls from 18.61 to 1.29; photon noise rises by a factor of 1.03. At 1000 photons per pixel.
Fig. 2 The same sweep at a tenth the light. The colour half of the trade is unchanged — a matrix is a matrix — and the noise half has grown, so the same correction costs more in a dark room than in a bright one.

At the other end of the exposure range the noise cost nearly disappears, which is why the same profile can be right on a tripod and wrong indoors.

Correcting colour costs noise, and the two cannot both be least. Sweeping the colour matrix from its own diagonal — a white balance with no cross terms — to the full least-squares fit. Mean ΔE00 falls from 18.61 to 1.29; photon noise rises by a factor of 1.03. At 100000 photons per pixel.
Fig. 3 And at ten times the light, where the noise cost nearly disappears. What a full correction is worth is therefore a property of the exposure rather than of the camera, which is why the same profile is right on a tripod and wrong indoors.
Correcting colour costs noise, and the two cannot both be least. Sweeping the colour matrix from its own diagonal — a white balance with no cross terms — to the full least-squares fit. Mean ΔE00 falls from 18.61 to 1.29; photon noise rises by a factor of 1.03. At 10000 photons per pixel.
Fig. 4 The original exposure on a more saturated set of surfaces. The colour the correction buys grows and the noise it costs does not, so how far along the sweep to stop depends on what is being photographed as well as on how much light there is.

The claim

The off-diagonal terms of a colour matrix are what correct a sensor that fails Luther’s condition, and they are also what amplify photon noise. Colour accuracy and signal-to-noise cannot both be maximised, and the trade-off is a consequence of the condition failing.

The second sentence is the part worth defending, because the trade-off is usually described as an engineering compromise between two independent desirable properties. It is not independent — though the dependence turns out to be on which basis the raw channels are in rather than on the condition itself, which a later section measures.

Where the amplification comes from

Photon arrival is Poisson. If a photosite collects NN photons on average, the standard deviation of the count is N\sqrt{N}, so the relative noise is 1/N1/\sqrt{N} — better in bright light, worse in dim, and irreducible by any amount of engineering because it is a property of the light rather than of the detector.

Crucially, the three channels’ noise is independent. Different photosites, different photons, no correlation. So the noise arriving at the matrix has a diagonal covariance.

A linear map turns that into MΣMTM \Sigma M^{\mathsf{T}}. For a diagonal Σ\Sigma with entries σi2\sigma_i^2, the variance of the $i$th output is

jmij2σj2\sum_j m_{ij}^2 \sigma_j^2

which is the squared row norm weighted by the input variances. A row with large entries of both signs has a large squared norm, even though the entries sum to something near one — the signs cancel in the signal and add in quadrature in the noise.

That is the whole mechanism. A typical row of a real matrix looks like 1.9R0.9G+0.0B1.9R - 0.9G + 0.0B, or 0.2R+1.6G0.4B-0.2R + 1.6G - 0.4B: the coefficients are much larger in magnitude than their sum, and the difference between the two is exactly the amplification.

Why the cross terms have to be there

They are not decoration. They are the manoeuvre that synthesises a response the sensor does not have.

The Luther fit has to go negative to reproduce the second lobe of xˉ\bar{x}, because no single filter transmittance has two lobes and a linear combination buys one by subtraction. The colour matrix is doing the same thing on the way to XYZ: it is constructing an approximation to a matching function out of three curves that individually are the wrong shape.

A silicon sensor's best possible impersonation of the standard observerThe 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 31.7 per cent of the matching functions' own magnitude, worst at 440 nm. Colour reproduction is exact if and only if this is zero.0x̄ ȳ z̄, behind — what a colorimeter needsthe sensor's best linear fit to the matching functionsthe fit goes negative, and has toworst at 440 nmwhat is left over — residual 31.7% overall400450500550600650700750wavelength / nmmodelled silicon sensorLuther–Ives 1927, the fit and its residual
Fig. 5 The subtraction, made visible. The fitted curves dip below zero, which no filter can do and a matrix can — and a matrix that does it is a matrix with large opposed coefficients, which is a matrix that amplifies noise.

So the sequence is: the sensor fails Luther, therefore the matrix needs cross terms, therefore the noise grows. A Luther-satisfying sensor’s matrix would be whatever fixed mixing matrix relates it to XYZ, with no fitting and no excess.

The sweep, and why it is a sweep

The measurement is made by interpolating between two matrices rather than comparing two points, because a single pair invites the reading that one of them is correct.

One end is the fitted matrix’s own diagonal: the same white balance with none of the cross terms, which is what a camera that declined to correct colour would actually apply. The other end is the full least-squares fit. Every point between is a matrix, and each one has a colour error and a noise gain.

Colour error falls monotonically along the sweep. Noise rises monotonically. The two curves cross nowhere useful, and the assertion in the library requires exactly that — that error falls and noise rises along the same axis — so a change that made them agree would stop the build.

There is no optimum on the axis without a further statement about how much a unit of noise is worth against a unit of ΔE, and that statement is a product decision rather than a measurement. Manufacturers place it differently, and the placement varies with the ISO setting: at high gain, cameras demonstrably back off the colour correction, because the noise term has grown and the trade has moved.

The one number that makes it concrete

At ten thousand photons per pixel — a reasonable mid-tone exposure — the full matrix amplifies noise by 1.66 times relative to leaving raw alone.

A factor of 1.66 in noise is a little under one and a half stops of exposure. Put the other way: a camera applying a full colour correction needs about 2.75 times as many photons to reach the same noise as one that does not, which is a substantial fraction of a sensor generation’s worth of improvement, spent on colour.

Nobody presents it that way, because the two quantities are measured by different people using different metrics and are reported in different sections of the same review.

Which sensors pay the penalty

Fitting the best matrix for four sensors and taking each row’s noise gain — its Euclidean norm over the sum of its entries, which is what turns independent equal-variance inputs into output variance:

sensor colour error mean noise gain
the colorimetric control as shipped 10⁻¹¹ 1.571
the same, mixed to the identity 10⁻¹² 1.000
the same, mixed for the best adaptation basis 10⁻¹³ 0.787
a real silicon sensor 2.03 1.025

The first three all satisfy Luther’s condition exactly and are all exact in colour. Their noise gains are 1.571, 1.000 and 0.787.

So the penalty is not paid for failing the condition. It is paid for the mixing matrix, and the condition says nothing at all about which mixing matrix. A Luther sensor whose channels are already XYZ needs the identity and pays exactly nothing. One whose channels are a plausible-looking long, medium and short mixture needs a matrix to undo that mixture, and pays 57 per cent. One mixed for the best adaptation basis pays less than nothing, its rows having smaller norms than their own sums.

And the real silicon sensor, which fails the condition by a Luther residual of 0.317, pays 1.025 — less than the control that satisfies it exactly.

That does not undo the trade; it relocates it. The off-diagonal terms really do amplify noise and a fit really does put them there. What decides how large they are is how far the raw channels sit from the output basis, not how far they sit from being a linear combination of the matching functions — and those two distances are unrelated, which is why one sensor here is exact and noisy while another is inexact and quiet.

What is done about it instead

Since the trade cannot be escaped, the engineering effort goes into changing the shape of the problem, and three approaches are worth naming because each is visible in ordinary photographs.

Denoise before the matrix. If the noise is reduced while the covariance is still diagonal, the amplification acts on a smaller quantity. This is why modern converters denoise early and why raw converters that denoise late produce noticeably worse colour noise.

Denoise the chroma preferentially. Human spatial acuity for chroma is far below that for luminance — the same fact the Bayer pattern exploits — so blurring the colour channels heavily while leaving luminance sharp removes most of the visible penalty at almost no cost in apparent detail. Every camera does this, aggressively, and it is why a high-ISO photograph looks grainy rather than blotchy.

Back off the correction when the noise is large. At high gain the matrix is moved towards its diagonal, trading colour accuracy for noise deliberately. It is the same sweep, chosen per exposure.

A less saturated chart is the case a manufacturer would reach for to make the trade look better, and the sweep prices what that buys.

Correcting colour costs noise, and the two cannot both be least. Sweeping the colour matrix from its own diagonal — a white balance with no cross terms — to the full least-squares fit. Mean ΔE00 falls from 16.56 to 0.85; photon noise rises by a factor of 1.03. At 10000 photons per pixel.
Fig. 6 The same sweep from the matrix’s own diagonal to the full least-squares fit, fitted against samples of chroma 0.3. Mean ΔE00 falls from 16.56 to 0.85 and the noise still rises by a factor of 1.03, so a duller chart lowers the error at both ends and does not change the price.

The same trade at the sensor

The matrix is the second place the trade appears and not the first. It is already present in the choice of the filters, and putting the two together explains a design decision that otherwise looks perverse.

Narrow, well-separated colour filters fit the matching functions better — they give a smaller Luther residual, so the matrix needs smaller cross terms, so the noise amplification is lower. They also pass less light, because a narrow filter transmits a narrower band, so the photon count NN falls and 1/N1/\sqrt{N} rises.

Broad, heavily overlapping filters do the opposite: more light, a worse residual, larger cross terms, more amplification of a smaller noise.

So the noise penalty is paid twice and in opposite directions, and the optimum is somewhere in the middle of a two-dimensional trade rather than at either end. That is why real sensitivity curves look the way they do: not because nobody has tried narrower filters, but because narrower filters lose more than they save.

It also explains the one design that escapes part of the trade. A sensor with a fourth channel, or with unfiltered “white” sites among the coloured ones, collects more light without narrowing anything — and several phone sensors do exactly this, trading the mosaic’s colour resolution for luminance signal. It is the Bayer argument applied again with a different weighting.

What a photographer sees of it

The trade is normally invisible because both of its ends are adjusted at once, and there are three places where it surfaces.

Colour gets worse as the light gets worse, and not only because the picture is noisier. At high gain the correction is deliberately reduced, so a high-ISO photograph is less colour-accurate in a way that has nothing to do with the noise being visible. Two exposures of the same subject at two sensitivities differ in colour, not just in grain.

Different converters disagree more at high ISO than at low. Each has its own view of where on the sweep to sit, and the views diverge as the noise term grows. A comparison of raw converters made on a base-ISO studio shot understates their disagreement considerably.

And a “colour-accurate” profile costs measurable image quality. Photographers who install a strictly colorimetric profile in place of the manufacturer’s often report the result as noisier, and it is: a profile fitted harder against a chart has larger cross terms, and larger cross terms amplify more.

That last one is worth stating plainly because it is usually attributed to something else. The accurate profile is not revealing noise that the pleasing one was hiding. It is making more of it, and the mechanism is the row norm.

Walking the sweep in twenty-one steps rather than eleven says whether the curve has any structure between the samples the coarser walk took.

Correcting colour costs noise, and the two cannot both be least. Sweeping the colour matrix from its own diagonal — a white balance with no cross terms — to the full least-squares fit. Mean ΔE00 falls from 18.61 to 1.29; photon noise rises by a factor of 1.03. At 10000 photons per pixel.
Fig. 7 Twenty-one settings between the diagonal and the fit. Mean ΔE00 still runs from 18.61 to 1.29 and the noise factor is still 1.03 at the far end, so the trade is monotone all the way along rather than turning anywhere in the middle.

Where the model stops

Shot noise is not the only noise. Read noise is additive and roughly constant, dark current grows with temperature and exposure time, and quantisation adds a floor. Shot noise dominates in the mid-tones and read noise dominates in the shadows, so the amplification computed here is the mid-tone case; in the shadows the arithmetic is similar with a different Σ\Sigma and the conclusion is unchanged.

The covariance stops being diagonal after demosaicing. Interpolating a channel from its neighbours correlates adjacent pixels, so the noise arriving at the matrix in a real pipeline is spatially correlated even if it started independent. That makes the exact number different and does not change the mechanism.

And the noise gain is computed at one signal level. It varies across the tonal range because σj\sigma_j does, and the figures state the photon count for that reason.

The generalisation

The transferable statement is about what a linear correction does to an error budget, and it is one of the more useful things to carry out of this field.

A transform that corrects a systematic error by taking differences of correlated quantities amplifies the uncorrelated error in those quantities, and the amplification is the row norm rather than the row sum. The signal cares about the sum; the noise cares about the norm; and a correction whose coefficients are large and opposed has a small sum and a large norm by construction.

That shape is everywhere. Background subtraction in a spectroscopic measurement, differencing two nearly equal survey estimates, a control system cancelling a disturbance by opposing a large term with another large term, a financial hedge constructed from two correlated positions: in every case the correction is real and the residual variance grows.

The general protection is to state both quantities in one place and sweep between them, which is what the figure at the top of this essay does and what almost nobody does in practice. Reporting the corrected accuracy alone is not dishonest and it is incomplete, and the incompleteness is systematic: the party who benefits from the correction reports it, and the party who pays for it in noise is usually downstream and unrepresented.

This site has already met the same shape once, from a different direction. Raising a check’s tolerance to make it pass leaves a check unable to distinguish a correct answer from a wrong one — a correction bought at the price of the thing the correction was for.

Why the noise is coloured, and why that matters

One consequence deserves separating because it changes what the noise looks like rather than how much of it there is.

Before the matrix, the three channels’ noise is independent and equal in relative terms, which reads as luminance grain — the kind of texture film had, and which people find acceptable and even attractive.

After the matrix it is neither. The three output variances differ, because the three rows have different norms, and the outputs are now correlated, because each shares inputs with the others. Correlated noise of unequal variance in three channels is chromatic noise: coloured blotches rather than grain.

And the eye is much more tolerant of luminance noise than of chroma noise, which is the same asymmetry the Bayer pattern exploits and the opponent channels explain: the chromatic pathways are low-pass, so chromatic noise appears as slow coloured drifts across areas rather than as fine texture, and slow coloured drifts across a face are far more objectionable than grain.

So the matrix does not simply increase noise; it converts a tolerable kind into an intolerable one. That is the real reason chroma denoising is applied as aggressively as it is, and it is why a photograph at high ISO can look grainy and clean at the same time.

Who noticed, and when

Poisson statistics of photon counting go back to the beginnings of quantum optics, and their application to imaging sensors was thoroughly worked out during the development of scientific CCDs in the nineteen-seventies and eighties, where the noise budget is the entire performance specification.

The interaction with colour correction is more recent as a stated design consideration. It appears in the digital camera literature from around the turn of the century, usually under “colour correction noise amplification” or as part of a broader treatment of the imaging chain’s error propagation, and the standard metric is the amplification factor computed exactly as above.

What is worth noticing is how little the framing has changed. It is nearly always presented as a trade-off between two things a camera would like to have, rather than as a consequence of a condition stated in 1927 — and stating it the second way makes it predictable, because it says that the amplification is largest for sensors with the largest Luther residual. That is a testable relationship between two published numbers and it is not, as far as this collection can tell, routinely tested.

What this predicts, and how it could be checked

A claim worth stating separately, because it is the kind this collection prefers: the argument above is not merely an explanation but a prediction, and it is testable with numbers that are already published.

If the amplification is caused by the Luther residual, then sensors with larger residuals should require matrices with larger row norms, and should therefore show a larger noise penalty from colour correction. That is a relationship between two quantities that different communities already measure — sensitivity metamerism indices from the colour-science literature, and colour-corrected signal-to-noise from imaging measurement — for overlapping sets of cameras.

The prediction has a caveat that would need handling. A sensor with broad filters has both a worse residual and a higher photon count, and those pull the measured noise in opposite directions, so the comparison has to be made at matched light collection rather than at matched exposure. Without that control the relationship could easily come out flat or backwards, and a flat result would be an artefact of the design.

This site cannot run that test: it has one modelled sensor rather than a population of real ones, and a prediction about a population needs a population. What it can do is state the prediction precisely enough that somebody with the data could refute it, which is the standard every claim here is held to.

Where the ladder goes next

Downward, this rung sits on no matrix is right everywhere, which is where the matrix is optimised without regard to its cost, and on a threshold is not a unit, which is where a ΔE stopped being a fixed quantity.

Upward, a blown highlight turns is the other end of the sensor’s range, and a photograph is not a measurement collects this and the rest of the field into one answer.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 22 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Camera rawColour matrixΔELeast-squaresLuther conditionQuantisationShot noiseSignal-to-noiseSpectral sensitivityThreshold