The coincidence was a mechanism
Assumes A chart decides what a camera scores, Saturation is nearly everything and The weighting is the disagreement.
A round ago two numbers arrived from opposite ends of this collection and were two per cent apart. A claim was made about why, and the claim turns out to be testable.
The claim
Two independent measurements agreeing once is a coincidence; the same two agreeing at six points across a factor of two and a half is a shared mechanism, and the mechanism is the compression the unit applies to a chroma difference.
- The two quantities are unrelated. One is how much the adaptation census depends on how saturated its test surfaces are; the other is how much a camera profile’s error depends on how saturated its test chart is.
- They agreed to two per cent in the published unit, which is what prompted the claim that the agreement is about the pairing of a spectral mismatch with a compressive metric rather than about either measurement.
- Vary the metric and both move, from about 1.15 under plain CIELAB to about 0.5 under the appearance unit.
- They stay together the whole way. The largest gap between the two curves is 12 per cent, at the far end; four of the six units have them inside 6.
- And the agreement is not a rescaling. The two curves are not proportional to each other; they track through a factor of 2.4 with a nearly constant small offset.
The two quantities, and how little they share
Worth restating, because the strength of the result depends entirely on their independence.
The census’s elasticity. Fourteen changes of light, each scored as what an adapted observer is left with, averaged over 125 constructed surfaces built as sums of three cosines. Multiply the surfaces’ modulation depths by 1.25 and by 0.8 and take the log-ratio of the answers: that is the elasticity, and it is the largest sensitivity anywhere in this collection at 0.687.
The camera’s elasticity. One silicon sensor with a colour filter array and an infrared-cut filter, a 3×3 fitted by least squares on twelve desaturated Gaussian-bump reflectances, then scored on twelve Gaussian bumps of a different saturation. Sweep the scoring chart’s chroma with the matrix held and the reported error moves by a factor of 5.4 across the full range; over this collection’s standard ±25 per cent span about the reported chroma it is 0.656.
They share the words test set and saturation and nothing else. Different files, different physics — one is a failure of a diagonal gain and the other a failure of the Luther condition — different families of reflectance, different reference points, different white points in the arithmetic. There is no code path connecting them, which is why the agreement was worth noticing rather than debugging.
What the claim was, and why it is testable
The sentence written when the two numbers arrived was:
That is not a coincidence and it is not a shared bug — there is no shared code to carry one. It is what happens when two quantities are both a failure to handle spectral structure, reported in a compressive colour-difference metric.
Read carefully, that names two ingredients and asserts that the agreement belongs to their pairing.
A failure to handle spectral structure is the physics, and it genuinely is common to both: a von Kries gain cannot see structure narrower than a cone, and a 3×3 from camera raw to tristimulus cannot either. In both cases a more saturated test surface has more spectral structure, so the failure grows.
A compressive difference metric is the unit, and it is where the two ingredients can be pulled apart. If the agreement is a property of the pairing, then removing the compression should change both numbers and change them together. If it is a coincidence between two unrelated quantities that happened to land near each other, removing the compression should scatter them.
That is a prediction, and it is one the arithmetic could refuse. The round that made the claim had no way to test it, because a unit was not a variable then.
What happens
| unit | census | camera | apart |
|---|---|---|---|
| CAM16-UCS | 0.532 | 0.472 | 11.9% |
| ΔE2000 | 0.687 | 0.656 | 4.8% |
| ΔE*94 | 0.835 | 0.792 | 5.3% |
| ΔE*uv | 1.024 | 1.120 | 8.9% |
| Oklab | 1.071 | 1.189 | 10.4% |
| ΔE*ab | 1.105 | 1.151 | 4.0% |
Both curves run from about 0.5 to about 1.15, in the same order, with the same steps. The prediction holds.
And it holds in a stronger form than the original claim asserted. The original said the agreement was a property of the pairing. What the table shows is that both quantities are, individually, functions of the unit’s compression with nearly the same functional form — so it is not merely that they agree at one point, it is that they are the same curve measured through two different pieces of machinery.
The ordering is the collection’s now-familiar one. The two units that divide a chroma difference by the chroma it was measured at give the lowest elasticities; the three that do not give elasticities above one. Above one means an answer that moves more than proportionally with the set it was measured over, which is a bad place for a published figure of merit to be and is where both quantities sit under three of the six formulae.
Why the two curves are not identical
They agree to between 4 and 12 per cent rather than exactly, and the residual disagreement is informative.
The camera is below the census under the two weighted units and above it under the three unweighted ones. That is a crossing rather than an offset, and it happens between ΔE*94 and CIELUV.
The reason is the shape of the test surfaces. The census’s are sums of two cosines on a level, so a more saturated member has broader structure — more modulation across the whole visible range. The camera chart’s are Gaussian bumps, so a more saturated member has narrower structure — a deeper, sharper band. A chroma weighting acts on the resulting colour difference and does not care which; but the two families reach a given chroma by different spectral routes, and how much of the mismatch survives to become a chroma difference differs.
So the crossing says the two mechanisms are the same to first order and differ in the second. Which is the correct amount of agreement to find: two identical curves from unrelated machinery would suggest a shared bug after all, and this collection has learned to distrust an agreement that is too good.
The published value was 0.672 and this essay says 0.656
Both are right and the difference is a construction, which is worth stating rather than smoothing over.
The earlier figure was taken over the whole chroma sweep, from 0.10 to 1.00 — a factor of ten in the input — and reported as an elasticity by taking the log-ratio of the endpoints. The figure here is taken over this collection’s standard ±25 per cent span about the chroma the profile is reported at, which is the span every other elasticity in the collection uses and is what makes the two numbers in this essay comparable.
The two constructions differ by 2.4 per cent, which is smaller than the effect being measured and larger than nothing. The standard span is the right one here because the whole point is a comparison against the census’s elasticity, and a comparison between two elasticities taken over different spans is not a comparison. Where a quantity is being compared, the construction has to match; where it is being reported alone, the wider span says more about the function’s shape.
That an elasticity depends on the span it is taken over is not a defect — it is why this collection reports finite differences rather than derivatives, since nobody is asking what an infinitesimally different test set would give.
What a tracking pair of curves is worth
More than two agreeing numbers, and the reason generalises past this subject.
Two numbers agreeing at one point can agree by chance. With two quantities both plausibly between 0.3 and 1.2, landing within two per cent of each other has a probability of a few per cent — small, and not small enough to publish a mechanism on.
Two curves agreeing at six points cannot. Each unit is a separate opportunity for the two to part, and they do not. The joint probability under any reasonable null is negligible, and more to the point the ordering is identical: both curves rank the six units the same way, which is one of 720 orderings.
And the shared variable is the one named in the claim. The claim did not say “these two agree”; it said “they agree because both are reported in a compressive metric”. Varying exactly the named ingredient and getting both to move together is the strongest form of confirmation available without an experiment. This collection’s habit is to give every claim a test it could fail, and this is one where the test arrived a round after the claim.
What it says about the two literatures
The result has a use beyond confirming a sentence, and it is about how the two figures of merit should be quoted.
A camera profile’s error and an adaptation residual are both means over a test set, reported in a unit, and both are quoted in the trade as though they were properties of the device or of the change of light. The tracking curves say something sharper than “both depend on the test set”: they say the dependence itself is a property of the reporting convention shared by the two fields, and is therefore transferable.
That gives a rule of thumb worth having. An error quoted at one chart saturation and wanted at another is out by the ratio to the power of the elasticity, and the elasticity is now known as a function of the unit rather than as one number: about 0.66 in ΔE2000, 0.79 in ΔE*94, 1.15 in plain CIELAB. A camera error measured on a desaturated chart and quoted for a saturated one, in CIELAB, is out by nearly the full ratio.
It also says which correction is safe. Correcting between charts of similar saturation is reliable in any unit; correcting across a factor of two of chroma is reliable only if the unit is stated, because the exponent changes by 1.7 across the menu. The same caution attends every extrapolated exponent in this collection, and this one has the unusual merit of being measured at six points rather than assumed.
Where the model stops
The camera side is one sensor. A different colour filter array, a different infrared-cut, or a sensor that satisfies the Luther condition would give a different elasticity — the Luther-satisfying control gives an error of zero and no elasticity at all, which is the degenerate case that proves the effect is about mismatch.
The census side is one basis and one family of surfaces. Both have been audited separately and neither is measured, so the elasticity is a property of a construction rather than of adaptation in general.
And “compressive” is doing work that has not been pinned down. Three of the six units compress, in three different ways — a linear divisor, a logarithm inside a model, and nothing at all except the space’s own cube root. The claim as tested is that more compression gives a lower elasticity, and the six points are consistent with it; a mechanism stated as a functional form rather than as a direction would need a dial rather than a menu, and the one dial available covers only two of the six.
The last figure is the physical content of both halves of the coincidence, and it is worth naming as one thing. A camera fails because three sensitivities that are not a linear combination of the matching functions cannot separate two spectra a person separates. An adapted observer fails because three gains cannot undo a multiplication that acted on structure finer than three channels can resolve. Both are a three-dimensional instrument meeting a higher-dimensional world, and in both the amount of world that leaks through grows with how much spectral structure the test surfaces have.
Who found it, and when
The idea that a quantity common to two independent measurements can be tested by varying it is the oldest move in experimental design and needs no attribution. Its application here is unusual only in that the common quantity is a reporting convention rather than a physical parameter, and conventions are not usually variables.
The specific mechanism — that a compressive metric reduces the apparent sensitivity of a figure of merit to the saturation of its test set — is, as far as this collection can tell, not stated anywhere, and it is the sort of thing that would be obvious to anybody who computed it and is computed by nobody. Both halves of that sentence are the ordinary condition of a result that lives between two literatures: colour-difference formulae are studied by people fitting them to observers, and test-set design is studied by people specifying charts, and the interaction belongs to neither.
Three and three, not two and three
The account of the crossing given above is right about where it happens and wrong about how many units sit on each side, and the miscount is worth correcting because it changes which property of a unit the crossing is attributed to.
Subtracting the two columns gives the camera minus the census, unit by unit: −0.060 under CAM16-UCS, −0.031 under ΔE2000, −0.043 under ΔE*94, then +0.096 under CIELUV, +0.118 under Oklab, +0.046 under CIELAB. Three below and three above, with the sign change between ΔE*94 and CIELUV — which is exactly where the crossing was located, so the description of the boundary was never in doubt. What was misfiled is ΔE*94 itself, counted with the unweighted units on one line and used as the lower edge of the crossing on the next.
ΔE*94 divides a chroma difference by 1 + 0.045C. That is a chroma weighting, milder than ΔE2000’s and applied without the hue-dependent terms, but it is the same construction: an absolute chroma difference reported as a fraction of the chroma it was measured at. So the split is not two weighted against three unweighted with one unaccounted for; it is three weighted below the line and three unweighted above it, and the boundary between the two groups and the boundary between the two signs are the same boundary. That is a cleaner result than the one claimed, and it survives without the extra assumption that the mildest weighting behaves like no weighting at all.
The pair is not an offset and not a ratio
The other loose end is the shape of the relation between the two curves, described above as tracking with a nearly constant small offset. Fitted three ways over the six points, that is the worst of the three descriptions available:
| relation | parameter | rms residual |
|---|---|---|
| camera = k × census | k = 1.038 | 0.064 |
| camera = census + b | b = 0.021 | 0.070 |
| camera = m × census + c | m = 1.284, c = −0.227 | 0.035 |
A constant offset does slightly worse than a constant ratio, and the two-parameter line does twice as well as either. What the fitted slope says is that the camera’s elasticity is about a quarter more responsive to the choice of unit than the census’s is — the census’s six values span a factor of 2.08 and the camera’s span 2.52, over the same six units. The correlation across the six is 0.992, so the two really are one curve seen twice; they are not the same curve at the same scale.
That distinction matters for the transfer rule the closing section proposes. Reading one elasticity off the other requires the slope as well as the ratio, and a rule that assumes a constant offset will be out by 0.06 at the compressive end — which is the whole of the disagreement the earlier section set out to explain.
One smaller thing, for anyone recomputing the table: the apart column is a symmetric percentage, the gap over the mean of the two rather than over either one. Taken against the census instead it reads 11.3, 4.5, 5.1, 9.4, 11.0, 4.2 — the same story, and up to 0.7 points different at the ends.
Where the ladder goes next
The camera profile in this essay is fitted by a linear least-squares solve in tristimulus space. That is an objective, it decides the matrix, and it is not on the menu of six — nobody proposed it as a colour-difference formula and nobody would. Refitting the same 3×3 under each of the six that were proposed turns a linear solve into a nine-parameter search, improves the fit in every one of them, and moves the matrix.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- Three choices reached calibration · chromatic adaptation · colour difference · elasticity · sensitivity · test set
- A choice with no magnitude calibration · colour difference · elasticity · sensitivity · test set
- The census in six units calibration · chromatic adaptation · colour difference · metamerism · test set
- The objective nobody chose calibration · camera profile · colour difference · luther condition · test set
- The disagreement is at the near end calibration · chroma · colour difference · test set
- The input nobody declared chromatic adaptation · elasticity · sensitivity · test set
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CalibrationCamera profileChromaChromatic adaptationColour differenceElasticityLuther conditionMetamerismSensitivityTest set