Difference and uniformity

Saturation is nearly everything

The set of test surfaces has three numbers describing it, and only one of them matters. How saturated the surfaces are carries an elasticity of about 0.7 on every result computed over them; how bright they are carries 0.10. A test chart's chroma range decides its answer and its lightness range does not.

Assumes A mean has a set under it, Which measurement is worth making and A width nobody varied.

The region the test surfaces come from is described by three numbers. Two of them are worth measuring properly and one of them can be left alone, and knowing which is which is worth more than any of the three.

How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one.
Fig. 1 Every change of light in the adaptation census against the three numbers that describe the region its test surfaces are drawn from. The bars for how saturated the surfaces are run four to nine times longer than the bars for how bright.

The claim

Every result computed over the test set responds to how saturated its members are, and almost not at all to how bright they are.

  • Saturation carries an elasticity of 0.49 to 0.91, averaging 0.69: multiply the modulation depths by 1.25 and every published residual rises by about a fifth.
  • The L¹ reach — how far two modulations may go together — carries 0.51 to 0.99, averaging 0.79. It is the same quantity by another name.
  • Brightness carries 0.078 to 0.147, averaging 0.104. A twenty-five per cent change in how bright the surfaces are moves the answer by two per cent.
  • The ratio is about seven. Effort spent choosing a chart’s lightness range is effort spent on the wrong axis.
  • The same elasticity appears in a completely different part of this site. A camera profile’s reported error responds to its test chart’s saturation at 0.67, against the census’s 0.69, from machinery that shares nothing.

What the three numbers are

The test surfaces are built from three basis functions and are indexed by three parameters, and the region they are drawn from is fixed by three declarations:

what it fixes as declared what it is
the brightness five levels from 0.10 to 0.54 how much light the surfaces return
the saturation seven modulation depths, ±0.6 how far each of two spectral undulations goes
the reach the two depths sum to at most 0.7 how far both may go at once

None of the three is quoted from anything. They are round numbers that make a lattice of a convenient size out of surfaces that stay physical, and nothing in the collection has ever moved them.

Because they are scalars, they take the instrument the round before this one built for the five declared population widths: a central difference in the logarithm, taken over a stated finite span rather than as a derivative, over the same factor of 1.25 up and 0.8 down so that the two elasticity tables can be read against one another. That comparability is the point of using the same span, and is why it is not chosen per parameter.

The measurement

change of light saturation reach brightness
the macular pigment 0.911 0.993 0.147
daylight to tungsten 0.816 0.913 0.113
daylight to D100 0.758 0.848 0.100
daylight to D50 0.745 0.834 0.099
a red wall 0.741 0.827 0.106
daylight to D40 0.736 0.823 0.099
daylight to halogen 0.734 0.817 0.105
daylight to a white LED 0.695 0.789 0.104
an older lens 0.671 0.741 0.101
daylight to a three-primary display 0.606 0.693 0.095
a green wall, one bounce 0.602 0.677 0.096
a green wall, two bounces 0.599 0.673 0.098
a triphosphor tube 0.516 0.598 0.089
daylight to a blackbody 0.494 0.515 0.078

Two columns of large numbers and one column of small ones. Nothing is close to the boundary between them: the smallest saturation elasticity, 0.494, is more than three times the largest brightness elasticity, 0.147.

The objects the average is over, and the region they come from. Two panels. On the left, eight of the 125 reflectance spectra in the test set, drawn as reflectance against wavelength from 380 to 780 nanometres — smooth, broad curves between about 0.02 and 0.9, with at most two gentle undulations each, because each is a level times a combination of two cosines. None of them has a narrow feature, because the family has no basis function that could make one. On the right, the region those surfaces come from, drawn in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the constraint that the two depths may not exceed 0.7 in sum, and 25 lattice points inside it. Five levels of each of those pairs is the whole test set. The square's four corners — the most saturated surfaces the two cosines could make — are outside the diamond and are not in the set at all.
Fig. 2 The region in its own coordinates. Scaling the brightness moves the five levels; scaling the saturation stretches the diamond. Only the second changes what the census measures.

Why brightness does almost nothing

Because the residual is measured in a lightness–chroma space and the operation is a multiplication.

Making a surface brighter multiplies its whole reflectance by a constant. Under either light that multiplies the tristimulus by the same constant, so the two stimuli move together, and the CIELAB coordinates move together too except for the compression the cube root imposes. The residual is a difference between two adapted coordinates, and a common multiplicative factor almost cancels out of a difference.

What is left is the cube root’s curvature — a bright surface and a dark one sit at different places on it, so the same tristimulus ratio maps to slightly different lightness differences. That is the whole of the 0.10, and it is a second-order effect by construction.

So a test chart’s lightness range is not a design parameter for this measurement. That is a useful negative result: it means a chart can be built at whatever lightness is convenient for the instrument, and the answer will not move.

No one surface carries the answer, and the set is smaller than it looks. A falling bar chart of the 125 surfaces in the test set, ordered by how much each contributes to the published mean for the macular pigment. The tallest bar is 2.40 per cent of the total, so the mean is not a few awkward objects with a crowd behind them and a leave-one-out would move it by well under a per cent. The tail is the other half of the story: 5 surfaces contribute essentially nothing, because a flat grey is a surface an adaptation gain handles exactly. Counting the set by how evenly it contributes rather than by how many members it has gives 94.4 effective surfaces out of 125, which is what "a mean over a hundred and twenty-five surfaces" is really worth.
Fig. 3 The macular pigment has the largest saturation elasticity in the table at 0.911, and this is why: its answer is concentrated in the most modulated surfaces, so the region’s reach decides almost all of it.

Why saturation does nearly everything

Because a residual after adaptation is exactly a failure to handle spectral structure, and the modulation depth is the spectral structure.

A flat surface has none, and a flat surface’s residual is exactly zero — an adapted observer’s gain is the ratio of the whites, and a grey under a changed light is the white scaled. Everything the census measures comes from departures from flat, so scaling the depths scales the thing being measured almost directly.

An elasticity of 0.7 rather than 1.0 says the relationship is sublinear, and the reason is CIEDE2000. The metric compresses at high chroma — its chroma weighting function grows with chroma, so a fixed spectral perturbation on a saturated surface produces a smaller reported difference than the same perturbation on a neutral one. Deeper modulation gives more physical mismatch and less metric credit for it, and the two combine to 0.7.

The reach and the saturation are the same quantity, and both decide where the worst surface sits. The L¹ bound decides how many of the deepest surfaces are in the set at all; scaling it scales the set’s most modulated members. Its elasticity is uniformly a little higher than the depths’ because it acts only on the deep end, which is where the answer is.

What one change of light costs, surface by surface — daylight to tungsten. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to tungsten is 1.635 ΔE₀₀. The curve runs from 4.4e-14 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 3.058, which is 1.87 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 4 Where the answer is. The curve’s steep upper end is the deeply modulated surfaces, and the region’s declared reach decides how many of them the set contains.

The two columns are proportional, which the explanation above does not predict

The account just given makes the brightness column an artefact of the metric: a common multiplicative factor cancels out of a difference except for the cube root’s curvature, and what survives is the 0.10. That is an explanation in terms of something every row shares, so it predicts a brightness column that is the same number fourteen times.

It is not. The column runs from 0.078 to 0.147 — a factor of 1.88, which is very nearly the factor of 1.84 the saturation column spans. And the two are not varying independently:

r=0.86r = 0.86

Plotting one against the other gives a straight line through very nearly the origin with a slope of 0.115, and the largest departure from it anywhere in the table is 0.019. Row by row, the brightness elasticity is a fixed fraction of the saturation elasticity — between 0.132 and 0.172 of it, averaging 0.145 — rather than a fixed quantity.

A metric artefact would be a constant. What is measured is a proportion. So the cube root is not the whole story: something about each change of light decides how sensitive it is to the test set’s amplitude in general, and both columns inherit that in a fixed ratio. The row that responds most to how saturated its surfaces are also responds most to how bright they are, and by the same multiple.

That does not overturn anything practical, and it strengthens the headline. The claim above says the ratio between the two columns is “about seven”; taking it row by row gives 5.8 to 7.6, averaging 6.9. Seven is not an average that happens to come out of fourteen scattered ratios — it is what every row separately reports, to within thirteen per cent. A conclusion of the form lightness range matters seven times less than chroma range is therefore transportable to a change of light not in this table, which the averaged version could not have licensed.

What is left open is what the shared factor is. The obvious candidate is the same one that governs the saturation column: how much of each row’s answer sits in the deeply modulated surfaces, since those are also the surfaces furthest from the cube root’s linear region. That would make both columns readings of one concentration, and it is testable against a statistic the collection already computes — the ratio of the worst surface to the mean. It does not survive the test: the rank correlation between that ratio and the saturation elasticity is +0.32, which on fourteen points is nothing. Two plausible measures of how concentrated a row’s answer is disagree, so at least one of them is measuring something else, and this essay cannot say which.

And the reach-minus-depth range is misquoted

One correction while the table is open. The section below says the reach elasticity runs 0.05 to 0.11 above the depth elasticity on every row; it runs 0.021 to 0.097, and both ends are wrong.

Thirteen of the fourteen rows sit between 0.070 and 0.097, which is the tight band the claim was describing. The fourteenth is a blackbody at daylight’s own temperature, at 0.021 — a third of the next-smallest gap and the row whose two elasticities are the lowest in the table by a clear margin.

That is consistent rather than anomalous. The reach acts only on the boundary of the region, so its extra leverage is worth something only where the answer is concentrated at the boundary. The blackbody row has the flattest distribution in the census — it treats every surface much alike — so it has the least to gain from a declaration that moves only the extreme members, and its two elasticities converge. The one row where the gap nearly vanishes is the one row where the boundary is not where the answer is.

The same number, from somewhere else

The strongest evidence that this is a property of the question rather than of one file is that it turns up, unchanged, in a part of the site that shares no machinery with it.

A camera profile is a 3×3 fitted on a set of surfaces and reported on a set of surfaces. Its test surfaces are Gaussian bumps of stated chroma, not cosine combinations; its error is a mean ΔE*₀₀ between recorded and true tristimulus, not a residual after adaptation; nothing in its code path touches the census.

Hold the fit fixed and sweep the chart’s chroma:

chart chroma reported mean error
0.10 0.373
0.25 0.758
0.55 1.185
0.85 1.661
1.00 2.002

A factor of 5.4 from one end to the other, with the camera and its matrix unchanged. Taken as an elasticity over the same span, that is 0.672, against the census’s mean of 0.687.

Two per cent apart, from two unrelated measurements over two unrelated constructions of a test set. That is not a coincidence and it is not a shared bug; it is what happens when two quantities are both a failure to handle spectral structure, reported in a compressive colour-difference metric. The elasticity is a property of that pairing.

A silicon sensor's best possible impersonation of the standard observerThe 1931 matching functions in outline, and the closest linear combination of the sensor's three sensitivities laid over them; underneath, what is left over at each wavelength. The residual is 31.7 per cent of the matching functions' own magnitude, worst at 440 nm. Colour reproduction is exact if and only if this is zero.0x̄ ȳ z̄, behind — what a colorimeter needsthe sensor's best linear fit to the matching functionsthe fit goes negative, and has toworst at 440 nmwhat is left over — residual 31.7% overall400450500550600650700750wavelength / nmmodelled silicon sensorLuther–Ives 1927, the fit and its residual
Fig. 5 The camera’s own fit. The matrix is the same in every row of the chroma sweep; only the surfaces it is scored on change.
What one change of light costs, surface by surface — daylight to a blackbody. A rising curve of 125 points, one per surface in the test set, sorted from the surface this change of light costs least to the one it costs most, with the published mean drawn across it as a horizontal line. The published residual for daylight to a blackbody is 0.263 ΔE₀₀. The curve runs from 0.0e+0 — 5 of the surfaces are flat greys, on which an adapted observer's gain is exactly right and the residual is exactly zero — to 0.457, which is 1.74 times the mean. The mean line crosses the curve about two thirds of the way along, so most surfaces cost less than the published number and a minority cost a great deal more. This is what a single published residual is a summary of.
Fig. 6 The row with the smallest saturation elasticity, 0.494. Its distribution is the flattest in the census — a blackbody at daylight’s own temperature is a mild change that treats every surface similarly — which is exactly what a low elasticity means.

What to do with it

Report a test set by its chroma range. A chart described as “twenty-four patches spanning the printable gamut” has said nothing that predicts what it will report. A chart described as “modulation depths to 0.6” or “chroma to 0.55” has said the only thing that matters.

Do not compare two published errors measured on different charts, for the reason a delivery tolerance has to name its condition. A factor of five is available between two defensible chart designs, and it swamps every real difference between two devices. This is the ordinary reason two laboratories report different numbers for the same camera, and the ordinary fix — agree on a chart — is the right one and works because the sensitivity to everything else is small.

And when the chart has to change, correct with the elasticity. An error measured at one chroma and quoted at another is out by the ratio to the power 0.7, which is a one-line correction and is better than nothing by a factor of about three.

Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy.
Fig. 7 The other lever on the same numbers: how the region is walked. Its effect is one to five per cent, against the tens of per cent the region’s saturation carries — so the shape of the region matters far more than the rule that samples it.
Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.
Fig. 8 What the elasticity does not touch: the ordering. Scaling the saturation moves every row together, so the ranking these gaps establish survives it entirely.

How this compares with what was declared

The round before this one ranked five declared population widths by how much of a conclusion’s uncertainty each carried, and the largest elasticity anywhere in that table was about a half.

The largest here is 0.99, and it belongs to a number that was never declared, never quoted with a range, and never varied. The comparison is this round’s own result and it deserves its own rung; the part that belongs here is the practical half, which is that the effort of an audit should go where the elasticity is, and the elasticity was not where the declared inputs were.

Two numbers or one

The saturation and the reach are two declarations doing one job, and separating them is worth a paragraph because the separation is what makes the elasticity actionable.

The depths declare how far a single modulation may go: ±0.6. The reach declares how far two may go together: 0.7 in sum. A surface at the maximum of one and zero of the other is admitted; a surface at 0.6 of each is not.

Their elasticities differ systematically — reach is 0.02 to 0.10 higher than depth on every row, and 0.07 to 0.10 on thirteen of the fourteen — and the reason is which surfaces each moves. Scaling the depths moves every surface in the set, including the mild ones near the centre, most of which contribute little. Scaling the reach moves only the boundary, which is where the answer is concentrated. A declaration that acts only on the extreme members is a more efficient lever on a mean whose upper tail carries it.

That has a practical reading for anyone designing a test set. The right thing to specify is not the maximum saturation of any patch but the shape of the region the patches fill, because two charts with the same maximum chroma and different distributions towards it will report different numbers. A chart described by its extreme is described by the one member the mean depends on least per unit of it.

What the elasticity does not license

One reading to refuse, because the exponent invites it.

It does not convert a residual measured on one set into a residual on a different construction. The correction (s₂/s₁)^0.7 is valid along the axis it was measured on — the same region, scaled — and says nothing about a set built to another rule. The clamped realistic family sits 1 to 52 per cent below the idealised one, which no saturation scaling reproduces, because the clamp changes the surfaces’ shape rather than their amplitude.

So the exponent is a within-family correction. It is exactly the right instrument for this chart was a quarter more saturated than that one, and exactly the wrong one for these two laboratories used charts of different kinds.

Where the model stops

The elasticities are computed on this region, about this operating point, over this span. They are local: the saturation elasticity at twice the declared depths would be different, because the metric’s compression grows and the surfaces start hitting the physical bounds.

And the brightness result is specific to a residual expressed in a lightness-normalised space. A quantity reported in absolute tristimulus rather than in CIELAB would respond to brightness at an elasticity near one, trivially, and the smallness here is a statement about the metric as much as about the surfaces. That is worth saying because it is the case where the actionable conclusion inverts: a chart’s lightness range is irrelevant to a ΔE and central to a radiometric error.

The elasticity’s own error

A number quoted to three decimal places invites a question about its precision, and the answer is worth having because it is not what the decimals suggest.

The elasticity is a finite difference between two means, each computed over a lattice, so it carries all of the quadrature bias that a level does — except that the bias is largely common-mode between the two, since both are lattices over scaled versions of the same region. Recomputing the saturation elasticity against the uniform measure rather than the lattice moves it by under two per cent on every row.

So the third decimal place is not meaningful and the second is. The right way to quote the table is 0.91 for the macular row and 0.49 for the blackbody row, and the useful statement is the ratio between the two columns rather than either entry: saturation is about seven times brightness, which is a factor that no plausible refinement of the measurement moves.

That is the same discipline the rest of the round arrives at. A quantity’s precision is set by the construction it was computed over, not by the arithmetic that produced it, and printing more digits than the construction supports is how a difference of four parts in ten thousand came to look like a step in a ranking.

Who found it, and when

The instrument belongs to the audit that preceded this one and is described there at length. What is added is applying it to a structural choice by finding the scalars inside the structure — a set has no magnitude, but the region it is drawn from has three declared numbers, and those do.

The cross-check against the camera profile was not planned. It was run to see whether this round’s thesis reached a second part of the site at all, and the answer coming back at 0.672 against 0.687 is the kind of agreement that makes a finding worth more than the sum of its two measurements. Two instruments agreeing to two per cent is evidence about the world; one instrument reporting 0.687 is evidence about a file.

How far each published number moves when a declared width does. A grid of bars, one row per published conclusion and one bar in each row per declared width of the population model. A bar's length is the elasticity — the proportional change in the conclusion for a proportional change in that width — so a bar of length one means an answer that doubles when the width doubles. The largest here is 0.92 and most are between a tenth and a half. One row is empty: the median member's distance from the standard observer does not respond to any width at all, exactly, because scaling a width leaves every median where it was. That row is the control and its being exactly flat is what says the rest are measuring spread.
Fig. 9 The earlier audit’s table, for comparison: the five declared population widths against the conclusions that rest on them. The longest bar there is about a half.

Where the ladder goes next

If the most elastic input in the collection is the one nobody declared, that is a statement about how audits go wrong rather than about colour, and it is worth stating on its own.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 12 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Camera profileChromaChromatic adaptationColour differenceElasticityLightnessReflectanceResidualSensitivityTest set