Where the model breaks

Two instruments and one ranking

A sampling error over a hundred and twenty-five surfaces and a change of colour-difference formula share no arithmetic at all, and they were asked the same question of the same table. Every adjacency the whole menu reverses had already been flagged as unresolved. And one the test set settles at nine standard errors is reversed by four of the five formulae, which is what makes them two instruments rather than one.

Assumes The error on a gap is not the errors at its ends, The census in six units and Four steps the test set cannot order.

The adaptation census’s fourteen rows have an order, and two instruments have now been pointed at it. Neither knows the other exists.

Which steps of the census ranking a change of unit reverses. Every adjacent pair in the published census ranking that at least one unit puts the other way round. The bar counts how many of the five other units reverse it. The marker on the left says whether the test set had already declared the pair unresolved — a gap smaller than twice its own paired standard error, which is a statement about sampling over 125 surfaces and shares no arithmetic with a change of ruler. The two pairs every unit reverses are both flagged, which is the agreement. The pair at the bottom is the disagreement: the test set resolves it at 9.1 standard errors and four of the five units reverse it anyway, because a sampling error cannot see a change of ruler and a change of ruler cannot see a sampling error.
Fig. 1 Every adjacent pair in the published census ranking that at least one colour-difference formula puts the other way round, with a marker saying whether the paired standard error over the test set had already declared the pair unresolved. Two pairs are reversed by all five formulae and both were flagged. One is resolved at nine standard errors and reversed by four.

The claim

Two instruments that share no arithmetic agree about which steps of a ranking are not real — and each finds something the other cannot.

  • The first instrument is sampling. The census’s residuals are means over 125 surfaces, so each gap between adjacent rows has a standard error, and a gap smaller than twice its own error is not established. Four of the thirteen steps are in that state.
  • The second is the unit. Recomputing the same fourteen rows under five other colour-difference formulae reverses between two and ten of the ninety-one pairs, depending on the formula.
  • Every adjacency all five formulae reverse was already flagged. Two of two, exactly, and there is no reason in the arithmetic why that should be so.
  • And one the test set resolves at t = 9.08 is reversed by four of the five. A sampling error cannot see a change of ruler, so this is new information rather than a contradiction.
  • The overlap is 2 of 6. Most of what the second instrument finds, the first could not have found.

Two errors that are not the same error

Before the comparison, the two quantities need separating, because both get called uncertainty and they are not the same kind of thing.

The first is a sampling error. Each census row is a mean over a set of 125 surfaces. The mean has a standard deviation, and the difference between two rows has one too — smaller than the sum of the two rows’ errors, because the same surfaces score both, by a factor averaging 1.78. A gap of less than twice its own paired error is a gap this test set does not establish. It is a statement about which surfaces are in the set.

The second is a modelling difference. Recompute every row under a different colour-difference formula, calibrate the formula onto the published one’s scale so that only the real disagreement is left, and see whether the order survives. It is a statement about what counts as far apart.

The two share the census’s numbers and nothing else. One is a variance over a finite set; the other is a deterministic recomputation with no variance in it at all. Nothing about either predicts the other, which is what makes the comparison worth making.

The overlap, which is exact

Six adjacencies in the published order are reversed by at least one formula. Two of them are reversed by all five.

adjacency formulae that reverse it flagged as unresolved? t
daylight to D40 → a three-primary display 5 of 5 yes 0.01
daylight to tungsten → a white LED 5 of 5 yes 0.81
a red wall → halogen 4 of 5 no 9.08
an older lens → a red wall 1 of 5 no 3.68
a three-primary display → an older lens 1 of 5 no 2.05
the macular pigment → daylight to D50 1 of 5 no 3.97

Both of the universally reversed pairs were already flagged, and no flagged pair is reversed by only some of the formulae. The two lists agree exactly at the top.

The first pair is the census’s shortest step: 0.0004 ΔE2000 between two rows printed to four figures as 0.9796 and 0.9800, with a paired standard error of 0.034, so t = 0.01. It is a difference the table has been printing that was never there, and it takes any perturbation whatever to reverse it.

The second is larger — a gap of 0.062 against an error of 0.076 — and is more interesting, because the two rows are a tungsten lamp and a white LED, which is one of the three-way ties this collection has already reported. Halogen, tungsten and a white LED are the three things a room is actually lit by, all three are broad smooth continua, and three cone channels cannot separate them finely. That the unit instrument reverses the same pair says the tie is a property of the physics rather than of the sample size.

The disagreement, which is the more useful half

If the two lists agreed completely, the second instrument would be a slower version of the first. They do not, and the case where they part is the best result here.

A red wall against halogen is separated by 0.226 ΔE2000, with a paired standard error of 0.025. That is t = 9.08 — the fourth-widest gap in the ranking, established beyond any doubt the test set could raise. Four of the five formulae put the two rows the other way round.

The mechanism is visible in the two rows’ composition. A red wall is a multiplication of the light by a broad reflectance, which pushes the whole test set towards the red and increases its chroma; halogen is a thermal source a few hundred kelvin from the reference, which shifts the set without saturating it. So the red-wall row’s residual is carried disproportionately by chroma differences and halogen’s by lightness and hue ones — and the single property separating these formulae is exactly whether a chroma difference is divided by its own chroma.

A sampling error is blind to that by construction. It is computed within one formula; every surface is scored by the same ruler, so the ruler cancels out of the variance entirely. Nine standard errors of separation is a completely correct statement and it is a statement about one ruler.

The converse holds too. The unit instrument is deterministic and cannot see that a gap of 0.0004 is smaller than the noise; under every one of the six formulae that pair is a definite ordering, and five of them happen to give the opposite definite ordering to the sixth. Only the sampling error says the pair is not ordered at all.

Which steps of the census ranking the test set actually resolves. A horizontal bar for each of the 13 adjacent pairs in the census's ranking, from the smallest mean residual to the largest. A bar's length is the gap between the two rows in ΔE₀₀; the whisker on its end is twice the standard error of that gap, computed as a paired difference because the same 125 surfaces score both rows. Where the whisker reaches back past zero the pair is not ordered by this test set, and 4 of the 13 are in that state — marked. The largest steps, at the two ends of the ranking, are twenty standard errors wide and are not in doubt at all. The smallest is four parts in ten thousand between two rows the table prints as different numbers.
Fig. 2 The thirteen steps of the census ranking with the paired error on each gap. Nine are resolved and four are not. The unit instrument reverses two of the four unresolved ones and one step that this chart shows as among the widest.
The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.
Fig. 3 The same ranking as six columns. The two universally reversed pairs are the crossings that appear in every column; the red-wall-against-halogen reversal is the crossing that appears in four.

What a t of nine is a statement about

The gap between a t of 9.08 and a reversal under four rulers deserves stating carefully, because it is a general point about error bars.

A standard error answers: if the same measurement were repeated with a different draw from the same population, how much would it move? Here the population is the set of test surfaces, the draw is a lattice rather than a random sample — which is its own complication — and the answer is 0.025 ΔE2000.

What it does not answer is: how much would this move if the measurement had been made differently? Every fixed choice inside the measurement is invisible to it. The unit is one such choice. So is the basis the gain is taken in, the set’s own three declared numbers, the pigment template under the observer, and the decision to average over a set at all.

An error bar is a lower bound on uncertainty and is routinely read as an estimate of it. That is not a criticism of error bars; it is what they are. The correct reading of 0.226 ± 0.025 is that repeating this calculation on different surfaces would give nearly the same answer, and the correct addition is that repeating it with a different ruler gives the opposite sign. Both sentences are true and the second is not smaller than the first.

How much overlap there is, and how much there should be

Two of six is the honest headline, and it is worth asking what number would have been surprising.

If the two instruments were measuring the same thing, all six reversals would be flagged and all four flagged pairs would be reversed. They are not: two flagged pairs are reversed by nobody, and four unflagged ones are reversed by somebody.

If they were entirely unrelated, the overlap would be about what chance gives. Four of thirteen adjacencies are flagged, so a reversal picked at random would be flagged about 31 per cent of the time; six reversals would give about two flagged by chance, which is what there are. On the count alone, the two instruments look independent.

The agreement is not in the count, it is in the strength. Both universally reversed pairs are flagged, and the universally reversed set is where the instruments agree exactly. Sorting the six reversals by how many formulae reverse them puts the two flagged pairs at the top and the four unflagged ones at the bottom, with no interleaving. That ordering is the result, and it is not something chance produces.

The error on a gap is not the two rows' errors added. Two bars for each of the 13 adjacent pairs in the census ranking. The upper, shorter bar is the standard error of the gap taken as a paired difference — the same 125 surfaces score both rows, so a surface that is awkward under one change of light is usually awkward under the other and the difference is quieter than either. The lower bar is the two rows' own errors added in quadrature, which is what comparing error bars by eye amounts to. Pairing is worth a factor of 1.78 on average and 3.36 on the pair it helps most, and it is the difference between 6 adjacencies unordered and 4. The gain is largest where the two rows are two daylights or two tungstens, because then the surfaces they find awkward are nearly the same surfaces.
Fig. 4 Why the paired error is the right one for this comparison. Adding two rows’ errors in quadrature — which is what a reader does by eye when a ranking is drawn with error bars on the rows — declares six adjacencies unresolved where four are, so the unpaired version would have over-agreed with the unit instrument for the wrong reason.
How much of the census ranking each unit keeps. Kendall's τ between each unit's ordering of the fourteen census rows and the published one; 1 would be perfect agreement. The number beside each bar is what τ is computed from — how many of the 91 pairs of rows the unit puts the other way round. The best agreement is ΔE′, the one unit here that is not a distance between two triples at all, at τ 0.956 and 2 discordant pairs. The worst is ΔEok at 0.780. A ranking that survives every unit is a ranking a reader can rely on; the ones that do not are listed in the figure that follows.
Fig. 5 Each formula’s total disagreement with the published ordering, as Kendall’s τ. Note that the number of reversed pairs and the number of reversed adjacencies are different counts: a formula can reverse ten pairs while reversing only four neighbours.

Kendall’s τ counts the reversals a unit produces. It does not say how far from the published unit a formula has to travel before it produces any, and that distance is measurable on its own.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 6 Each unit’s remaining disagreement with ΔE2000 once its own best scale factor has been divided out, over the reference pairs. A unit sitting at zero here would be the published one in different units and could reverse nothing at all; every reversal in this essay is bought with a departure on this chart.

Which formulae reverse what

The three formulae that reverse most are the three that do not weight chroma, and the ordering within them is informative.

Oklab reverses ten pairs of ninety-one and is the only formula to reverse the macular pigment against daylight-to-D50 — two of the mildest rows in the census, separated by 0.091 at t = 3.97. The macular pigment is a filter inside the observer that absorbs in the blue, so the surfaces it moves are moved along a direction where Oklab’s spacing differs most from CIELAB’s.

ΔE*ab reverses eight and is the only one to reverse an older lens against a red wall. CIELUV reverses seven and is the only one to reverse a three-primary display against an older lens. Each unweighted formula has one reversal of its own, and each of the three is a pair involving one of the two ocular filters. Those two rows are the ones whose residuals come from the blue end, and the blue end is where the three unweighted spaces differ from each other most.

CAM16-UCS reverses two, both of them the universally reversed pairs, and nothing else. It is the appearance model on the menu and its ranking of the census is the closest of the five to the published one, which is not the direction anybody would have predicted.

What to do with a ranking that has two kinds of hole in it

The practical question is what a reader should be told about a table like this, and there are three defensible answers of increasing cost.

Print the ranking with its unresolved steps marked, which is what this collection already does for the sampling half. That is cheap, it is honest about one instrument, and it is what the census’s own figure has shown since the standard errors were computed.

Print it with both marks, which costs a second computation of every row under a second formula and is what this essay argues for. The second computation is not expensive — the census under six units is under a second — and it catches a class of defect the first cannot: a step that is wide, well-established and dependent on a convention.

Or refuse to rank the middle at all. Nine of the thirteen steps are resolved by the test set and the two universally reversed ones are not; between them there is a block of rows the table orders and neither instrument fully supports. A ranking printed as three bands — the mild end, an unordered middle, the harsh end — would lose nothing that has been established and would stop inviting the reading that adjacent rows differ.

The third is the one this collection has not adopted, and the reason is worth stating rather than hidden: a banded table is much harder to compute a marginal yield from, and the marginal yield is the instrument that decides when a subject is finished. That is a reason of convenience, and naming it is the least that can be done about it.

Putting a number on “not something chance produces”

The essay’s strongest sentence is that sorting the six reversals by how many formulae reverse them puts the two flagged pairs at the top with no interleaving, and that that ordering is not something chance produces. It is worth computing, because the phrase invites a reader to imagine a small probability and the actual one is not that small.

Two nulls are available and the essay uses the first correctly. Four of the thirteen adjacencies are flagged and six are reversed by somebody, so drawing six adjacencies at random gives exactly two flagged with probability 0.44, and two or more with probability 0.66. The count carries no information at all, which is what the section above concludes.

The strength claim needs the second null: given that two of the six reversals are flagged, how likely is it that those two happen to be the two with the highest reversal counts? There are fifteen ways to choose two of six, so the answer is 1 in 15, or 0.067.

That is one observation at about the seven per cent level. It is worth reporting, it points the way the essay says it does, and not something chance produces is a stronger phrase than 1-in-15 supports. The honest form is that the agreement between the two instruments rests on a single coincidence that chance would supply about one time in fifteen.

A sharper version of the same observation

There is a reading of the same data that is more striking than the one the essay makes, and it comes from looking at the four flagged adjacencies rather than at the six reversals.

Of the four steps the test set cannot establish, two are reversed by all five formulae and two by none. Nothing lands in between. Of the nine steps the test set does establish, one is reversed by four formulae, three by one, and five by none.

So the flagged group is bimodal and the resolved group is not. That is a more specific pattern than flagged pairs get reversed more often — it says the flagged group contains two pairs that every ruler disagrees about and two that no ruler disagrees about, which is not what a general tendency would produce. The two unreversed flagged pairs are cases where the test set is short of surfaces and the rulers all agree anyway; the two reversed ones are cases where the gap is not there under any convention.

The sample is four, so this is an observation and not a result. It is the observation worth extending if the census ever grows, and it is a sharper question than the count of overlaps: among the steps one instrument cannot resolve, does the other split into all-or-nothing?

The rank correlation, and the one row that breaks it

The last thing the six reversals can be asked is whether the two instruments agree by degree rather than by threshold — whether a step with a smaller t is reversed by more formulae.

Over all six the rank correlation is −0.56, in the expected direction and weak. Removing the red wall against halogen — the row the essay is built around, resolved at nine standard errors and reversed by four formulae — takes it to −0.87.

That single row is therefore carrying most of the disagreement between the two instruments, and its removal is not a repair: it is the point. The five other reversals are consistent with the two instruments measuring one underlying thing at different sensitivities; the red-wall row is the evidence that they are not, and it is a sample of one.

So both halves of this essay’s headline rest on single observations. The agreement rests on two pairs landing at the top of a list of six, which chance supplies one time in fifteen. The disagreement rests on one row out of thirteen. Neither is weak evidence for the mechanism — the chroma-weighting account explains the red-wall row specifically and predicted it before it was looked for — but both are the kind of finding that a larger census would either confirm or dissolve, and neither is yet the settled fact the summary bullets read as.

Where the model stops

Six reversals is a small sample of reversals, and the exactness of the agreement at the top rests on two pairs. A census with more rows would give a better test of it, and the census is fourteen rows because that is how many changes of light this collection models.

The t values themselves depend on treating a lattice as though it were a sample, which it is not. A jackknife over a deterministic lattice is a statement about which members the lattice contains, which is a well-defined question and is not the same as a sampling variance. The comparison here is unaffected in ordering and the absolute t values should be read as indicative.

And neither instrument reaches the choices under both of them: the basis, the family of surfaces, the observer. A third instrument pointed at any of those could reverse a pair both of these agree about, and there is no reason to expect it would find the same pairs.

How much the census depends on its test set, under each unit. Each column is a unit and each dot is one of the fourteen census rows: the elasticity of that row's residual to how saturated the test surfaces are, over the same ±25 per cent span every sensitivity in this collection uses. An elasticity of one means a set half again as saturated gives an answer half again as large. In ΔE2000 the fourteen run 0.49 to 0.91 about a mean of 0.69 — already the largest sensitivity measured anywhere in this collection. In ΔE76, which does not weight chroma at all, every one of the fourteen rises, the mean goes past one to 1.11, and the spread narrows from 0.42 to 0.25. The bar across each column is its mean.
Fig. 7 A third view of the same object: how much each row’s residual depends on the test set, under each unit. The rows whose sensitivity moves most between units are not the rows whose ordering moves most, which is a fourth way of saying these instruments are not measuring one thing.

Who found it, and when

The statistical half is Student’s and Fisher’s and needs no introduction. The specific point that a paired comparison is the right one when both quantities are scored by the same sample is as old as the paired t-test, and this collection rediscovered it the hard way after publishing a table whose error bars were three times too wide.

The methodological half is newer and is not really from statistics. It is the observation that a computational result has two quite different kinds of uncertainty — the variance of its inputs and the arbitrariness of its construction — and that the literature reports the first and almost never the second. Under the name multiverse analysis the idea has had about ten years in the social sciences, where the constructions being varied are exclusions and covariates rather than colour-difference formulae. The finding there is the same as the finding here: the construction usually moves the answer more than the sampling does, and a paper reporting only a confidence interval is reporting the smaller of the two.

Where the ladder goes next

One of these six formulae was built on a space that this collection’s own uniformity instrument ranks last of the three it can rank. Whether the repair works is not a matter of opinion — MacAdam’s ellipses are a measurement of people, and turning that instrument on the units rather than on the spaces gives an answer with a sign nobody quotes.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationChromatic adaptationColour differenceJackknifeRank correlationResidualSamplingStandard errorTest setUncertainty