Raw is not a picture
Assumes A camera is a fourth observer and A hex code is not a colour.
Two people open the same raw file in two converters and get visibly different pictures. The usual explanation is that one of the converters is better, and the usual response is to argue about which.
Both are the wrong shape. The file does not contain a picture that one program renders more faithfully than the other; it contains a set of measurements and a list of questions, and the two programs answered the questions differently.
Every other stage in that chain is a decision of the same kind, and marking three more of them says what a raw file actually leaves open.
The largest term in the chain is the next one along, and unlike the mosaic it is estimated from the picture rather than computed from it.
The stage a raw file is furthest from recording is the first one, and it is the only one that is a property of the camera rather than of the software.
The claim
A raw file contains three integrals per sensor site and nothing that determines a colour. Six further decisions stand between it and an image, each has a defensible answer, and none has a correct one.
The word doing the work is determines. The measurements constrain the picture heavily — a raw file of a sunset will not render as a snowfield under any settings — but they do not pick one out, and the gap between constrained and determined is where the whole argument about converters lives.
What is actually in the file
Less than most people expect.
At each photosite there is one number: the count of electrons collected, digitised, with a black-level offset that has to be subtracted and an analogue gain that has to be divided out. One number, not three. The site had one colour filter over it, so it measured one of the three integrals and knows nothing about the other two.
Alongside the pixel data there is metadata, and the metadata is where the interesting omissions are. There is an exposure time, an aperture, an ISO setting, a camera model. There is usually a white balance the camera guessed and almost always a thumbnail rendered with the manufacturer’s own choices — the JPEG preview, which is what a photographer sees on the back of the camera and is not what the raw data says.
What is generally not in the file is the thing this whole field turns on: the spectral sensitivities of the sensor that took it. They are proprietary, they vary between production batches, and the number that stands in for them is a matrix somebody fitted, which is not the same object at all.
Stage by stage, and what each one decides
Black level and gain are the least controversial and are not free. A sensor’s zero is not zero; there is a pedestal, it drifts with temperature, and it is estimated from masked pixels at the edge of the array. Get it slightly wrong and the deepest shadows acquire a colour cast, because the three channels’ pedestals differ and a small absolute error is a large relative one down there.
Demosaicing invents two thirds of the data. Each site measured one channel and the other two are interpolated from neighbours, which is an assumption about how scenes behave. It is a good assumption and it fails at edges, where it produces colour that was not in the scene — an achromatic edge comes off the mosaic coloured, and that is a fact about the sampling rather than about the algorithm.
White balance is three multipliers, and choosing them is choosing what the light was. There is no measurement of the illuminant in the file — the camera did not have a spectroradiometer — so the number is an estimate, produced by an algorithm making an assumption about the scene. Grey-world assumes the average surface is neutral; max-RGB assumes something in frame is white; neither wins on every scene, and a camera that has chosen one has chosen which photographs to be wrong about.
The colour matrix takes raw to XYZ. It is the stage that would be exact if the sensor satisfied Luther’s condition, and since it does not, the matrix is a least-squares fit whose error depends on which surfaces it was fitted to. Two converters ship different matrices for the same camera. Both are defensible.
The tone curve is the largest single contributor to how the picture looks and the one least often described as a decision. A linear rendering of a scene looks flat and dark, and the S-curve applied to fix that is not a correction of anything — it is an appearance model in disguise, compensating for the fact that a photograph is viewed in a different surround, at a different absolute luminance, and subtending a different angle than the scene did.
The encoding turns a number that is a ratio into a code value, at which point it becomes the kind of thing a hex code names and stops being a measurement of anything.
The six-decision estimate
Six decisions, each with several defensible answers, is a combinatorial space rather than a disagreement. If white balance has three reasonable estimators, the matrix two reasonable fits, the tone curve four plausible shapes and demosaicing three algorithms, that is seventy-two renderings of one file, every one of them produced by competent people following defensible rules.
This is the honest answer to “which converter is right”. None of them, and the question does not have the form it appears to have. What can be asked instead — and this field spends most of its time asking it — is how far apart two defensible answers can be, and for which scenes the gap is largest.
The decision that is not on the list
There is a seventh thing a converter does that is not a stage in the chain, and it is the one that changes a photograph most: it decides what to do about the parts of the scene the sensor could not hold.
A sensor has a floor and a ceiling. Below the floor is read noise; above the ceiling the well is full and the count stops rising. Both ends are where the file stops being a measurement, and the ceiling is the interesting one because scenes routinely exceed it — a window in an interior, a sky above a landscape, a specular highlight on anything at all.
A clipped channel is a missing measurement, and every response to it is an invention. Clip all three to white and the highlight goes neutral, which is often what a person expects. Leave them and the hue turns, which is what the numbers say. Reconstruct the missing channel from the ratio of the two survivors and the result is a plausible colour that was never measured — defensible, frequently the best-looking option, and not a measurement.
The essay on what a blown highlight does takes this apart properly. What belongs here is only that the list of six decisions above is the list for the pixels that were in range, and that a converter is additionally deciding what to say about the ones that were not.
The camera’s own JPEG is not the answer either
A common resolution is to take the manufacturer’s rendering as the reference: the camera made a JPEG, the camera knows its own sensor, so that is what the file “really” looks like.
It is a reasonable instinct and it does not survive inspection. The manufacturer’s rendering is the same six decisions taken by people optimising for the same thing every other converter optimises for, which is that a photograph should look good on a screen rather than that it should agree with a colorimeter. Their tone curve is chosen for pleasingness, their white balance is deliberately incomplete — cameras leave a warm cast under tungsten on purpose, because full adaptation is not what the eye does either — and their saturation is boosted well past unity.
The manufacturer’s rendering has one genuine advantage over any third party’s: it was made by people who have the spectral sensitivities. It has one genuine disadvantage: it was made by people whose objective is not accuracy.
Two files, one scene, and what actually differs
It is worth being concrete about the size of the disagreement, because “converters differ” is compatible with anything from imperceptible to unrecognisable.
The white-balance decision is the largest of the six by a long way, and it is largest exactly where the scene breaks an estimator’s premise. On a scene that is mostly one hue — foliage, a green wall, a room lit through a single coloured blind — grey-world’s estimate is 43 degrees of angular error from the true illuminant, while an estimator that uses the specular highlights instead is under one. Forty-three degrees is not a nuance; it is a photograph of a forest rendered magenta.
The matrix decision is smaller and more insidious: a mean ΔE of one or two, spread unevenly, largest on the most saturated surfaces. Nobody notices it on a portrait and everybody notices it on a red car.
The tone curve dominates the impression and contributes almost nothing to colorimetric error, because it is a monotone map applied after the colour has been decided. That is the reason arguments about converters are so hard to settle: the stage a viewer reacts to first is not the stage that determines whether anything is accurate.
Where the model stops
Nothing above is about the optics. Lens flare, veiling glare and chromatic aberration all change the spectrum arriving at a photosite, and glare in particular is a colour problem — it adds a fraction of the scene’s average to every pixel, which desaturates everything and does so more in a bright scene. Longitudinal chromatic aberration has an exact analogue in the eye and a different consequence in a camera, because a camera can be refocused per channel and an eye cannot.
The stages are not really sequential. Modern converters demosaic and denoise jointly, apply white balance before demosaicing so the interpolation happens on more nearly equal channels, and fold the matrix and tone curve into one lookup. Drawing them as a chain is a description of what has to be decided rather than of what order the arithmetic happens in.
And nothing here is about compression. A raw file is usually losslessly compressed and sometimes not; a lossy raw format exists and discards information in a way that is invisible until one of the six decisions is changed afterwards, which is the point of raw in the first place.
The generalisation
The shape of this argument is not specific to photographs, and it recurs wherever an instrument’s output is presented as an image.
A measurement that has been made viewable has had decisions applied to it, and the decisions are usually invisible in the result. A false-colour satellite image, a medical scan windowed to a display range, an astronomical exposure stretched to show faint structure: in every case someone chose the mapping, in every case the choice is defensible, and in every case the picture is routinely reasoned about as though the mapping were part of the data.
The specific protection is not to distrust the picture but to keep the measurement and the rendering as separate objects, which is exactly what a raw file is for. The value of raw is not image quality. It is that the six decisions remain reversible, and remain visible as decisions.
That is also this site’s own practice, stated in a different vocabulary. Every figure here is generated from a rule rather than drawn once and saved, for the same reason: a rendering that cannot be re-derived has lost the distinction between what was measured and what was chosen.
What raw actually buys
Given all of the above, it is fair to ask what the point of keeping the file is, since none of the six decisions goes away.
The answer is that they stay reversible, and reversibility is worth more than any particular answer. A rendered JPEG has had the six decisions applied and then quantised to eight bits per channel, which is where the loss becomes permanent: a white balance applied and baked in has multiplied one channel up and clipped it, or multiplied another down into the quantisation floor, and no later correction recovers what was there. The eight-bit encoding is chosen to be perceptually efficient at one adaptation state, and changing the adaptation state afterwards is exactly the operation it cannot survive.
The second thing raw buys is more surprising and is a direct consequence of this field’s opening claim. Because the six decisions are separable, a photograph can be re-rendered under a different set of assumptions later — including assumptions that did not exist when the file was made. Camera profiles improve; demosaicing algorithms improve considerably; a file from fifteen years ago opened today is measurably better than the same file was then, and it is better because the measurements were kept apart from the interpretations.
There is no analogue of this for the eye. A person cannot re-render a memory of a colour under a corrected assumption about the light, which is why constancy is a mechanism rather than a setting and why the comparison between a camera and an observer stops being flattering to the observer at exactly this point.
Seventy-two counts four of the seven
The combinatorial estimate is offered as the size of the space the six decisions open, and it multiplies four of them: three white balances, two matrices, four tone curves, three demosaicing algorithms. Black level and encoding are on the list and not in the product.
Putting them in at two plausible options each gives 288. Adding the seventh decision the essay names separately — what to do about a clipped channel, which has at least three defensible answers — gives 216 on the four factors alone, or well over eight hundred on all seven.
So seventy-two is a floor and not a count. That does not weaken the point; it strengthens the shape of it, because the answer to which converter is right gets worse as the space gets larger and the essay’s own list is longer than its arithmetic.
The more useful correction is about which factor carries what. The essay says the tone curve dominates the impression and contributes almost nothing to colorimetric error, and the white balance is the largest colorimetric term by a wide margin. So the factor with the most options — four tone curves — is the one with the least colour in it, and the factor with the widest spread has three.
Read that way the seventy-two is four impressions of eighteen colours. A reader comparing two converters is nearly always comparing tone curves, which is the axis where the disagreement is largest to look at and smallest to measure, and that is why the argument never settles: the thing being argued about and the thing being measured are different axes of the same product.
A very short shutter is the case that shows the effect at its worst, and a twenty-thousandth of a second is inside what a modern camera offers.
What the transfer function is actually worth
The encoding stage is described as the point where a number stops being a measurement, and the size of what it buys and costs can be stated exactly.
At eight bits, the sRGB curve’s smallest step is at a linear value of 3.04 × 10⁻⁴; an eight-bit linear encoding’s smallest step is 3.92 × 10⁻³. So the curve reaches 12.9 times deeper into the shadows for the same number of bits. At the top the position reverses: sRGB’s last step spans 0.0089 of linear range against linear’s 0.0039, so it is 2.3 times coarser in the highlights.
Thirteen times better at one end and 2.3 times worse at the other is the whole bargain, and it is a good one because scene content and visual sensitivity are both concentrated at the bottom.
Put in the units the comparison is usually argued in, eight bits of sRGB behave like 11.7 linear bits in the shadows and 6.8 in the highlights. That is the reason a rendered file survives being looked at and does not survive being re-rendered: an operation that moves the shadows up into the highlights — which is what raising exposure is — moves data out of the eleven-bit region and into the seven-bit one, and the bits it needed are not there.
And it prices the raw file’s advantage precisely. A fourteen-bit linear raw resolves five times deeper than an eight-bit sRGB file — 6.1 × 10⁻⁵ against 3.04 × 10⁻⁴ — while carrying its precision uniformly rather than where one particular viewing condition wanted it. That uniformity is the property that makes the six decisions reversible: a linear file has no opinion about which part of the range matters, so changing one’s mind about the range costs nothing.
Who noticed, and when
Raw formats predate the argument. Early digital cameras wrote proprietary sensor dumps because they had no processing power to do anything else, and the files were an implementation detail rather than a philosophy.
The philosophy arrived with the converters. When independent software began opening manufacturers’ raw files around the turn of the century — reverse-engineered, since the formats were undocumented — the fact that two programs produced different pictures from one file became visible to everybody at once, and the argument this essay is about started immediately. Adobe’s DNG specification in 2004 is the clearest artefact of it: an attempt to standardise not the picture but the list of decisions, including a place to record the two illuminant-specific colour matrices that a converter would otherwise have to guess.
The interesting part is what DNG could not standardise. It has a field for the matrices and none for the sensitivities they were fitted from, because those are the manufacturer’s and always have been.
Where the ladder goes next
Downward, this rung rests on a camera is a fourth observer, which is where the sensor’s three curves are introduced, and on a hex code is not a colour, which is the same argument at the other end of the pipeline.
Upward, each stage has its own rung. Demosaicing and the matrix are next, white balance from the physics rather than from an assumption is two rungs further, and the whole question of what a photograph establishes closes the field.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The corner of the frame has another filter camera raw · colour filter array · colour matrix · spectral sensitivity · white balance
- A camera cannot record the excitation camera raw · colour matrix · spectral sensitivity · white balance
- A corner is corrected by one row camera raw · colour matrix · spectral sensitivity · white balance
- A matrix is fitted under one light camera raw · colour matrix · the icc profile · white balance
- Correcting colour costs noise camera raw · colour matrix · quantisation · spectral sensitivity
- Fitted to an eye nobody has camera raw · colour matrix · spectral sensitivity · white balance
What links here
The 8 essays that link to this one and share the most of its objects, of 11 that link here.
The objects this essay names
Each one links to every other essay that touches it.
Camera rawColour filter arrayColour matrixDemosaicThe ICC profileQuantisationSpectral sensitivityTone curveTransfer functionWhite balance