What a camera does

The coincidence was a mechanism

Two sensitivities from two libraries with no shared code came out two per cent apart, and the claim made about them was that they share a mechanism rather than a number. That claim has a colour-difference formula inside it, so it can be tested by changing the formula — and both curves move together across the whole menu, from 0.5 to 1.15.

Assumes A chart decides what a camera scores, Saturation is nearly everything and The weighting is the disagreement.

A round ago two numbers arrived from opposite ends of this collection and were two per cent apart. A claim was made about why, and the claim turns out to be testable.

Two sensitivities from two libraries, under every unit. Two quantities that share no code, no test set and no physical question: how much the adaptation census's residual depends on how saturated its surfaces are, and how much a camera profile's reported error depends on how saturated its test chart is. The first is a mean over fourteen changes of light built from cosine combinations; the second is one number about one silicon sensor scored on Gaussian bumps. Under the published unit they sit at 0.687 and 0.656. Across the whole menu they move together, from about 0.5 under the appearance unit to about 1.15 under plain CIELAB, staying within 12 per cent of each other at the worst point. Two numbers agreeing once is a coincidence; two curves agreeing at six points across a factor of two and a half is a shared mechanism, and the mechanism is the compression the unit applies to a chroma difference.
Fig. 1 Two sensitivities that share no code, no test set and no physical question, plotted against the six colour-difference formulae the collection could have published in. They move together across a factor of two and a half, staying within twelve per cent of each other everywhere.

The claim

Two independent measurements agreeing once is a coincidence; the same two agreeing at six points across a factor of two and a half is a shared mechanism, and the mechanism is the compression the unit applies to a chroma difference.

  • The two quantities are unrelated. One is how much the adaptation census depends on how saturated its test surfaces are; the other is how much a camera profile’s error depends on how saturated its test chart is.
  • They agreed to two per cent in the published unit, which is what prompted the claim that the agreement is about the pairing of a spectral mismatch with a compressive metric rather than about either measurement.
  • Vary the metric and both move, from about 1.15 under plain CIELAB to about 0.5 under the appearance unit.
  • They stay together the whole way. The largest gap between the two curves is 12 per cent, at the far end; four of the six units have them inside 6.
  • And the agreement is not a rescaling. The two curves are not proportional to each other; they track through a factor of 2.4 with a nearly constant small offset.

The two quantities, and how little they share

Worth restating, because the strength of the result depends entirely on their independence.

The census’s elasticity. Fourteen changes of light, each scored as what an adapted observer is left with, averaged over 125 constructed surfaces built as sums of three cosines. Multiply the surfaces’ modulation depths by 1.25 and by 0.8 and take the log-ratio of the answers: that is the elasticity, and it is the largest sensitivity anywhere in this collection at 0.687.

The camera’s elasticity. One silicon sensor with a colour filter array and an infrared-cut filter, a 3×3 fitted by least squares on twelve desaturated Gaussian-bump reflectances, then scored on twelve Gaussian bumps of a different saturation. Sweep the scoring chart’s chroma with the matrix held and the reported error moves by a factor of 5.4 across the full range; over this collection’s standard ±25 per cent span about the reported chroma it is 0.656.

They share the words test set and saturation and nothing else. Different files, different physics — one is a failure of a diagonal gain and the other a failure of the Luther condition — different families of reflectance, different reference points, different white points in the arithmetic. There is no code path connecting them, which is why the agreement was worth noticing rather than debugging.

What the claim was, and why it is testable

The sentence written when the two numbers arrived was:

That is not a coincidence and it is not a shared bug — there is no shared code to carry one. It is what happens when two quantities are both a failure to handle spectral structure, reported in a compressive colour-difference metric.

Read carefully, that names two ingredients and asserts that the agreement belongs to their pairing.

A failure to handle spectral structure is the physics, and it genuinely is common to both: a von Kries gain cannot see structure narrower than a cone, and a 3×3 from camera raw to tristimulus cannot either. In both cases a more saturated test surface has more spectral structure, so the failure grows.

A compressive difference metric is the unit, and it is where the two ingredients can be pulled apart. If the agreement is a property of the pairing, then removing the compression should change both numbers and change them together. If it is a coincidence between two unrelated quantities that happened to land near each other, removing the compression should scatter them.

That is a prediction, and it is one the arithmetic could refuse. The round that made the claim had no way to test it, because a unit was not a variable then.

What happens

unit census camera apart
CAM16-UCS 0.532 0.472 11.9%
ΔE2000 0.687 0.656 4.8%
ΔE*94 0.835 0.792 5.3%
ΔE*uv 1.024 1.120 8.9%
Oklab 1.071 1.189 10.4%
ΔE*ab 1.105 1.151 4.0%

Both curves run from about 0.5 to about 1.15, in the same order, with the same steps. The prediction holds.

And it holds in a stronger form than the original claim asserted. The original said the agreement was a property of the pairing. What the table shows is that both quantities are, individually, functions of the unit’s compression with nearly the same functional form — so it is not merely that they agree at one point, it is that they are the same curve measured through two different pieces of machinery.

The ordering is the collection’s now-familiar one. The two units that divide a chroma difference by the chroma it was measured at give the lowest elasticities; the three that do not give elasticities above one. Above one means an answer that moves more than proportionally with the set it was measured over, which is a bad place for a published figure of merit to be and is where both quantities sit under three of the six formulae.

How much the census depends on its test set, under each unit. Each column is a unit and each dot is one of the fourteen census rows: the elasticity of that row's residual to how saturated the test surfaces are, over the same ±25 per cent span every sensitivity in this collection uses. An elasticity of one means a set half again as saturated gives an answer half again as large. In ΔE2000 the fourteen run 0.49 to 0.91 about a mean of 0.69 — already the largest sensitivity measured anywhere in this collection. In ΔE76, which does not weight chroma at all, every one of the fourteen rises, the mean goes past one to 1.11, and the spread narrows from 0.42 to 0.25. The bar across each column is its mean.
Fig. 2 The census’s side of the pair, per row rather than averaged. The spread across the fourteen rows narrows as the weighting strengthens, so the mean the coincidence uses is a fair summary in the weighted units and a slightly generous one in the unweighted.
A camera's spectral sensitivities, after the infrared-cut filter. Silicon quantum efficiency times the colour-filter dye times the infrared-cut filter, per channel, on a grid running to 1100 nm rather than to 780. With the filter removed, 68 per cent of the area under the three curves lies beyond the visible band, and all three curves are the same curve out there.
Fig. 3 The camera’s side, at its root: a silicon sensor’s three spectral sensitivities, which are not a linear combination of the colour-matching functions and cannot be made into one. Every surface with spectral structure the three curves weight differently is a surface the profile gets wrong, and a more saturated chart has more of them.

Why the two curves are not identical

They agree to between 4 and 12 per cent rather than exactly, and the residual disagreement is informative.

The camera is below the census under the two weighted units and above it under the three unweighted ones. That is a crossing rather than an offset, and it happens between ΔE*94 and CIELUV.

The reason is the shape of the test surfaces. The census’s are sums of two cosines on a level, so a more saturated member has broader structure — more modulation across the whole visible range. The camera chart’s are Gaussian bumps, so a more saturated member has narrower structure — a deeper, sharper band. A chroma weighting acts on the resulting colour difference and does not care which; but the two families reach a given chroma by different spectral routes, and how much of the mismatch survives to become a chroma difference differs.

So the crossing says the two mechanisms are the same to first order and differ in the second. Which is the correct amount of agreement to find: two identical curves from unrelated machinery would suggest a shared bug after all, and this collection has learned to distrust an agreement that is too good.

The published value was 0.672 and this essay says 0.656

Both are right and the difference is a construction, which is worth stating rather than smoothing over.

The earlier figure was taken over the whole chroma sweep, from 0.10 to 1.00 — a factor of ten in the input — and reported as an elasticity by taking the log-ratio of the endpoints. The figure here is taken over this collection’s standard ±25 per cent span about the chroma the profile is reported at, which is the span every other elasticity in the collection uses and is what makes the two numbers in this essay comparable.

The two constructions differ by 2.4 per cent, which is smaller than the effect being measured and larger than nothing. The standard span is the right one here because the whole point is a comparison against the census’s elasticity, and a comparison between two elasticities taken over different spans is not a comparison. Where a quantity is being compared, the construction has to match; where it is being reported alone, the wider span says more about the function’s shape.

That an elasticity depends on the span it is taken over is not a defect — it is why this collection reports finite differences rather than derivatives, since nobody is asking what an infinitesimally different test set would give.

A silicon sensor's best possible impersonation of the standard observerThe 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 31.7 per cent of the matching functions' own magnitude, worst at 440 nm. Colour reproduction is exact if and only if this is zero.0x̄ ȳ z̄, behind — what a colorimeter needsthe sensor's best linear fit to the matching functionsthe fit goes negative, and has toworst at 440 nmwhat is left over — residual 31.7% overall400450500550600650700750wavelength / nmmodelled silicon sensorLuther–Ives 1927, the fit and its residual
Fig. 4 The camera’s fit, which is held fixed in every column of the sweep. Only the surfaces it is scored on change, so the elasticity is a property of the chart rather than of the profile.
The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.
Fig. 5 The other side’s underlying data: the census in six units. The levels move by up to a factor of two; the sensitivities move by more, which is why a sensitivity is the better instrument for detecting a shared mechanism.

What a tracking pair of curves is worth

More than two agreeing numbers, and the reason generalises past this subject.

Two numbers agreeing at one point can agree by chance. With two quantities both plausibly between 0.3 and 1.2, landing within two per cent of each other has a probability of a few per cent — small, and not small enough to publish a mechanism on.

Two curves agreeing at six points cannot. Each unit is a separate opportunity for the two to part, and they do not. The joint probability under any reasonable null is negligible, and more to the point the ordering is identical: both curves rank the six units the same way, which is one of 720 orderings.

And the shared variable is the one named in the claim. The claim did not say “these two agree”; it said “they agree because both are reported in a compressive metric”. Varying exactly the named ingredient and getting both to move together is the strongest form of confirmation available without an experiment. This collection’s habit is to give every claim a test it could fail, and this is one where the test arrived a round after the claim.

What it says about the two literatures

The result has a use beyond confirming a sentence, and it is about how the two figures of merit should be quoted.

A camera profile’s error and an adaptation residual are both means over a test set, reported in a unit, and both are quoted in the trade as though they were properties of the device or of the change of light. The tracking curves say something sharper than “both depend on the test set”: they say the dependence itself is a property of the reporting convention shared by the two fields, and is therefore transferable.

That gives a rule of thumb worth having. An error quoted at one chart saturation and wanted at another is out by the ratio to the power of the elasticity, and the elasticity is now known as a function of the unit rather than as one number: about 0.66 in ΔE2000, 0.79 in ΔE*94, 1.15 in plain CIELAB. A camera error measured on a desaturated chart and quoted for a saturated one, in CIELAB, is out by nearly the full ratio.

It also says which correction is safe. Correcting between charts of similar saturation is reliable in any unit; correcting across a factor of two of chroma is reliable only if the unit is stated, because the exponent changes by 1.7 across the menu. The same caution attends every extrapolated exponent in this collection, and this one has the unusual merit of being measured at six points rather than assumed.

Where the model stops

The camera side is one sensor. A different colour filter array, a different infrared-cut, or a sensor that satisfies the Luther condition would give a different elasticity — the Luther-satisfying control gives an error of zero and no elasticity at all, which is the degenerate case that proves the effect is about mismatch.

The census side is one basis and one family of surfaces. Both have been audited separately and neither is measured, so the elasticity is a property of a construction rather than of adaptation in general.

And “compressive” is doing work that has not been pinned down. Three of the six units compress, in three different ways — a linear divisor, a logarithm inside a model, and nothing at all except the space’s own cube root. The claim as tested is that more compression gives a lower elasticity, and the six points are consistent with it; a mechanism stated as a functional form rather than as a direction would need a dial rather than a menu, and the one dial available covers only two of the six.

The census along the line from ΔE76 to ΔE94, and past it. ΔE94 is ΔE76 with two weighting constants in it, and at zero those constants make every weight exactly one, so the two formulae are joined by a line rather than separated by a choice. The horizontal axis is how much of the published weighting is applied: 0 is exactly ΔE76, 1 is exactly ΔE94, and 3 is three times more weighting than anybody has proposed. The falling curve is the census's mean elasticity to how saturated its test set is, which drops from 1.11 to 0.74 — most of the fall happening before the published value is reached. The other curve is Kendall's τ against ΔE2000's ranking, and it peaks at w = 0.5, not at 1: the weighting that best reproduces the published ordering is about half the published weighting. There is no value of this dial that reaches ΔE2000, whose rotation term is not on this line at all.
Fig. 6 The dial that does exist, running from no weighting to three times the published weighting. Both elasticities in this essay would follow this curve if the dial covered them; it covers the census’s and not the camera’s, because the camera’s is computed in a different library.
What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 7 Six published quantities across the menu, for scale. The two in this essay are sensitivities rather than levels, and sensitivities move further than levels do — which is what makes them the better detector.

The last figure is the physical content of both halves of the coincidence, and it is worth naming as one thing. A camera fails because three sensitivities that are not a linear combination of the matching functions cannot separate two spectra a person separates. An adapted observer fails because three gains cannot undo a multiplication that acted on structure finer than three channels can resolve. Both are a three-dimensional instrument meeting a higher-dimensional world, and in both the amount of world that leaks through grows with how much spectral structure the test surfaces have.

Who found it, and when

The idea that a quantity common to two independent measurements can be tested by varying it is the oldest move in experimental design and needs no attribution. Its application here is unusual only in that the common quantity is a reporting convention rather than a physical parameter, and conventions are not usually variables.

The specific mechanism — that a compressive metric reduces the apparent sensitivity of a figure of merit to the saturation of its test set — is, as far as this collection can tell, not stated anywhere, and it is the sort of thing that would be obvious to anybody who computed it and is computed by nobody. Both halves of that sentence are the ordinary condition of a result that lives between two literatures: colour-difference formulae are studied by people fitting them to observers, and test-set design is studied by people specifying charts, and the interaction belongs to neither.

Three and three, not two and three

The account of the crossing given above is right about where it happens and wrong about how many units sit on each side, and the miscount is worth correcting because it changes which property of a unit the crossing is attributed to.

Subtracting the two columns gives the camera minus the census, unit by unit: −0.060 under CAM16-UCS, −0.031 under ΔE2000, −0.043 under ΔE*94, then +0.096 under CIELUV, +0.118 under Oklab, +0.046 under CIELAB. Three below and three above, with the sign change between ΔE*94 and CIELUV — which is exactly where the crossing was located, so the description of the boundary was never in doubt. What was misfiled is ΔE*94 itself, counted with the unweighted units on one line and used as the lower edge of the crossing on the next.

ΔE*94 divides a chroma difference by 1 + 0.045C. That is a chroma weighting, milder than ΔE2000’s and applied without the hue-dependent terms, but it is the same construction: an absolute chroma difference reported as a fraction of the chroma it was measured at. So the split is not two weighted against three unweighted with one unaccounted for; it is three weighted below the line and three unweighted above it, and the boundary between the two groups and the boundary between the two signs are the same boundary. That is a cleaner result than the one claimed, and it survives without the extra assumption that the mildest weighting behaves like no weighting at all.

The pair is not an offset and not a ratio

The other loose end is the shape of the relation between the two curves, described above as tracking with a nearly constant small offset. Fitted three ways over the six points, that is the worst of the three descriptions available:

relation parameter rms residual
camera = k × census k = 1.038 0.064
camera = census + b b = 0.021 0.070
camera = m × census + c m = 1.284, c = −0.227 0.035

A constant offset does slightly worse than a constant ratio, and the two-parameter line does twice as well as either. What the fitted slope says is that the camera’s elasticity is about a quarter more responsive to the choice of unit than the census’s is — the census’s six values span a factor of 2.08 and the camera’s span 2.52, over the same six units. The correlation across the six is 0.992, so the two really are one curve seen twice; they are not the same curve at the same scale.

That distinction matters for the transfer rule the closing section proposes. Reading one elasticity off the other requires the slope as well as the ratio, and a rule that assumes a constant offset will be out by 0.06 at the compressive end — which is the whole of the disagreement the earlier section set out to explain.

One smaller thing, for anyone recomputing the table: the apart column is a symmetric percentage, the gap over the mean of the two rather than over either one. Taken against the census instead it reads 11.3, 4.5, 5.1, 9.4, 11.0, 4.2 — the same story, and up to 0.7 points different at the ends.

Where the ladder goes next

The camera profile in this essay is fitted by a linear least-squares solve in tristimulus space. That is an objective, it decides the matrix, and it is not on the menu of six — nobody proposed it as a colour-difference formula and nobody would. Refitting the same 3×3 under each of the six that were proposed turns a linear solve into a nine-parameter search, improves the fit in every one of them, and moves the matrix.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationCamera profileChromaChromatic adaptationColour differenceElasticityLuther conditionMetamerismSensitivityTest set