Two instruments and one ranking
Assumes The error on a gap is not the errors at its ends, The census in six units and Four steps the test set cannot order.
The adaptation census’s fourteen rows have an order, and two instruments have now been pointed at it. Neither knows the other exists.
The claim
Two instruments that share no arithmetic agree about which steps of a ranking are not real — and each finds something the other cannot.
- The first instrument is sampling. The census’s residuals are means over 125 surfaces, so each gap between adjacent rows has a standard error, and a gap smaller than twice its own error is not established. Four of the thirteen steps are in that state.
- The second is the unit. Recomputing the same fourteen rows under five other colour-difference formulae reverses between two and ten of the ninety-one pairs, depending on the formula.
- Every adjacency all five formulae reverse was already flagged. Two of two, exactly, and there is no reason in the arithmetic why that should be so.
- And one the test set resolves at t = 9.08 is reversed by four of the five. A sampling error cannot see a change of ruler, so this is new information rather than a contradiction.
- The overlap is 2 of 6. Most of what the second instrument finds, the first could not have found.
Two errors that are not the same error
Before the comparison, the two quantities need separating, because both get called uncertainty and they are not the same kind of thing.
The first is a sampling error. Each census row is a mean over a set of 125 surfaces. The mean has a standard deviation, and the difference between two rows has one too — smaller than the sum of the two rows’ errors, because the same surfaces score both, by a factor averaging 1.78. A gap of less than twice its own paired error is a gap this test set does not establish. It is a statement about which surfaces are in the set.
The second is a modelling difference. Recompute every row under a different colour-difference formula, calibrate the formula onto the published one’s scale so that only the real disagreement is left, and see whether the order survives. It is a statement about what counts as far apart.
The two share the census’s numbers and nothing else. One is a variance over a finite set; the other is a deterministic recomputation with no variance in it at all. Nothing about either predicts the other, which is what makes the comparison worth making.
The overlap, which is exact
Six adjacencies in the published order are reversed by at least one formula. Two of them are reversed by all five.
| adjacency | formulae that reverse it | flagged as unresolved? | t |
|---|---|---|---|
| daylight to D40 → a three-primary display | 5 of 5 | yes | 0.01 |
| daylight to tungsten → a white LED | 5 of 5 | yes | 0.81 |
| a red wall → halogen | 4 of 5 | no | 9.08 |
| an older lens → a red wall | 1 of 5 | no | 3.68 |
| a three-primary display → an older lens | 1 of 5 | no | 2.05 |
| the macular pigment → daylight to D50 | 1 of 5 | no | 3.97 |
Both of the universally reversed pairs were already flagged, and no flagged pair is reversed by only some of the formulae. The two lists agree exactly at the top.
The first pair is the census’s shortest step: 0.0004 ΔE2000 between two rows printed to four figures as 0.9796 and 0.9800, with a paired standard error of 0.034, so t = 0.01. It is a difference the table has been printing that was never there, and it takes any perturbation whatever to reverse it.
The second is larger — a gap of 0.062 against an error of 0.076 — and is more interesting, because the two rows are a tungsten lamp and a white LED, which is one of the three-way ties this collection has already reported. Halogen, tungsten and a white LED are the three things a room is actually lit by, all three are broad smooth continua, and three cone channels cannot separate them finely. That the unit instrument reverses the same pair says the tie is a property of the physics rather than of the sample size.
The disagreement, which is the more useful half
If the two lists agreed completely, the second instrument would be a slower version of the first. They do not, and the case where they part is the best result here.
A red wall against halogen is separated by 0.226 ΔE2000, with a paired standard error of 0.025. That is t = 9.08 — the fourth-widest gap in the ranking, established beyond any doubt the test set could raise. Four of the five formulae put the two rows the other way round.
The mechanism is visible in the two rows’ composition. A red wall is a multiplication of the light by a broad reflectance, which pushes the whole test set towards the red and increases its chroma; halogen is a thermal source a few hundred kelvin from the reference, which shifts the set without saturating it. So the red-wall row’s residual is carried disproportionately by chroma differences and halogen’s by lightness and hue ones — and the single property separating these formulae is exactly whether a chroma difference is divided by its own chroma.
A sampling error is blind to that by construction. It is computed within one formula; every surface is scored by the same ruler, so the ruler cancels out of the variance entirely. Nine standard errors of separation is a completely correct statement and it is a statement about one ruler.
The converse holds too. The unit instrument is deterministic and cannot see that a gap of 0.0004 is smaller than the noise; under every one of the six formulae that pair is a definite ordering, and five of them happen to give the opposite definite ordering to the sixth. Only the sampling error says the pair is not ordered at all.
What a t of nine is a statement about
The gap between a t of 9.08 and a reversal under four rulers deserves stating carefully, because it is a general point about error bars.
A standard error answers: if the same measurement were repeated with a different draw from the same population, how much would it move? Here the population is the set of test surfaces, the draw is a lattice rather than a random sample — which is its own complication — and the answer is 0.025 ΔE2000.
What it does not answer is: how much would this move if the measurement had been made differently? Every fixed choice inside the measurement is invisible to it. The unit is one such choice. So is the basis the gain is taken in, the set’s own three declared numbers, the pigment template under the observer, and the decision to average over a set at all.
An error bar is a lower bound on uncertainty and is routinely read as an estimate of it. That is not a criticism of error bars; it is what they are. The correct reading of 0.226 ± 0.025 is that repeating this calculation on different surfaces would give nearly the same answer, and the correct addition is that repeating it with a different ruler gives the opposite sign. Both sentences are true and the second is not smaller than the first.
How much overlap there is, and how much there should be
Two of six is the honest headline, and it is worth asking what number would have been surprising.
If the two instruments were measuring the same thing, all six reversals would be flagged and all four flagged pairs would be reversed. They are not: two flagged pairs are reversed by nobody, and four unflagged ones are reversed by somebody.
If they were entirely unrelated, the overlap would be about what chance gives. Four of thirteen adjacencies are flagged, so a reversal picked at random would be flagged about 31 per cent of the time; six reversals would give about two flagged by chance, which is what there are. On the count alone, the two instruments look independent.
The agreement is not in the count, it is in the strength. Both universally reversed pairs are flagged, and the universally reversed set is where the instruments agree exactly. Sorting the six reversals by how many formulae reverse them puts the two flagged pairs at the top and the four unflagged ones at the bottom, with no interleaving. That ordering is the result, and it is not something chance produces.
Kendall’s τ counts the reversals a unit produces. It does not say how far from the published unit a formula has to travel before it produces any, and that distance is measurable on its own.
Which formulae reverse what
The three formulae that reverse most are the three that do not weight chroma, and the ordering within them is informative.
Oklab reverses ten pairs of ninety-one and is the only formula to reverse the macular pigment against daylight-to-D50 — two of the mildest rows in the census, separated by 0.091 at t = 3.97. The macular pigment is a filter inside the observer that absorbs in the blue, so the surfaces it moves are moved along a direction where Oklab’s spacing differs most from CIELAB’s.
ΔE*ab reverses eight and is the only one to reverse an older lens against a red wall. CIELUV reverses seven and is the only one to reverse a three-primary display against an older lens. Each unweighted formula has one reversal of its own, and each of the three is a pair involving one of the two ocular filters. Those two rows are the ones whose residuals come from the blue end, and the blue end is where the three unweighted spaces differ from each other most.
CAM16-UCS reverses two, both of them the universally reversed pairs, and nothing else. It is the appearance model on the menu and its ranking of the census is the closest of the five to the published one, which is not the direction anybody would have predicted.
What to do with a ranking that has two kinds of hole in it
The practical question is what a reader should be told about a table like this, and there are three defensible answers of increasing cost.
Print the ranking with its unresolved steps marked, which is what this collection already does for the sampling half. That is cheap, it is honest about one instrument, and it is what the census’s own figure has shown since the standard errors were computed.
Print it with both marks, which costs a second computation of every row under a second formula and is what this essay argues for. The second computation is not expensive — the census under six units is under a second — and it catches a class of defect the first cannot: a step that is wide, well-established and dependent on a convention.
Or refuse to rank the middle at all. Nine of the thirteen steps are resolved by the test set and the two universally reversed ones are not; between them there is a block of rows the table orders and neither instrument fully supports. A ranking printed as three bands — the mild end, an unordered middle, the harsh end — would lose nothing that has been established and would stop inviting the reading that adjacent rows differ.
The third is the one this collection has not adopted, and the reason is worth stating rather than hidden: a banded table is much harder to compute a marginal yield from, and the marginal yield is the instrument that decides when a subject is finished. That is a reason of convenience, and naming it is the least that can be done about it.
Putting a number on “not something chance produces”
The essay’s strongest sentence is that sorting the six reversals by how many formulae reverse them puts the two flagged pairs at the top with no interleaving, and that that ordering is not something chance produces. It is worth computing, because the phrase invites a reader to imagine a small probability and the actual one is not that small.
Two nulls are available and the essay uses the first correctly. Four of the thirteen adjacencies are flagged and six are reversed by somebody, so drawing six adjacencies at random gives exactly two flagged with probability 0.44, and two or more with probability 0.66. The count carries no information at all, which is what the section above concludes.
The strength claim needs the second null: given that two of the six reversals are flagged, how likely is it that those two happen to be the two with the highest reversal counts? There are fifteen ways to choose two of six, so the answer is 1 in 15, or 0.067.
That is one observation at about the seven per cent level. It is worth reporting, it points the way the essay says it does, and not something chance produces is a stronger phrase than 1-in-15 supports. The honest form is that the agreement between the two instruments rests on a single coincidence that chance would supply about one time in fifteen.
A sharper version of the same observation
There is a reading of the same data that is more striking than the one the essay makes, and it comes from looking at the four flagged adjacencies rather than at the six reversals.
Of the four steps the test set cannot establish, two are reversed by all five formulae and two by none. Nothing lands in between. Of the nine steps the test set does establish, one is reversed by four formulae, three by one, and five by none.
So the flagged group is bimodal and the resolved group is not. That is a more specific pattern than flagged pairs get reversed more often — it says the flagged group contains two pairs that every ruler disagrees about and two that no ruler disagrees about, which is not what a general tendency would produce. The two unreversed flagged pairs are cases where the test set is short of surfaces and the rulers all agree anyway; the two reversed ones are cases where the gap is not there under any convention.
The sample is four, so this is an observation and not a result. It is the observation worth extending if the census ever grows, and it is a sharper question than the count of overlaps: among the steps one instrument cannot resolve, does the other split into all-or-nothing?
The rank correlation, and the one row that breaks it
The last thing the six reversals can be asked is whether the two instruments agree by degree rather
than by threshold — whether a step with a smaller t is reversed by more formulae.
Over all six the rank correlation is −0.56, in the expected direction and weak. Removing the red wall against halogen — the row the essay is built around, resolved at nine standard errors and reversed by four formulae — takes it to −0.87.
That single row is therefore carrying most of the disagreement between the two instruments, and its removal is not a repair: it is the point. The five other reversals are consistent with the two instruments measuring one underlying thing at different sensitivities; the red-wall row is the evidence that they are not, and it is a sample of one.
So both halves of this essay’s headline rest on single observations. The agreement rests on two pairs landing at the top of a list of six, which chance supplies one time in fifteen. The disagreement rests on one row out of thirteen. Neither is weak evidence for the mechanism — the chroma-weighting account explains the red-wall row specifically and predicted it before it was looked for — but both are the kind of finding that a larger census would either confirm or dissolve, and neither is yet the settled fact the summary bullets read as.
Where the model stops
Six reversals is a small sample of reversals, and the exactness of the agreement at the top rests on two pairs. A census with more rows would give a better test of it, and the census is fourteen rows because that is how many changes of light this collection models.
The t values themselves depend on treating a lattice as though it were a sample, which it is not. A jackknife over a deterministic lattice is a statement about which members the lattice contains, which is a well-defined question and is not the same as a sampling variance. The comparison here is unaffected in ordering and the absolute t values should be read as indicative.
And neither instrument reaches the choices under both of them: the basis, the family of surfaces, the observer. A third instrument pointed at any of those could reverse a pair both of these agree about, and there is no reason to expect it would find the same pairs.
Who found it, and when
The statistical half is Student’s and Fisher’s and needs no introduction. The specific point that a paired comparison is the right one when both quantities are scored by the same sample is as old as the paired t-test, and this collection rediscovered it the hard way after publishing a table whose error bars were three times too wide.
The methodological half is newer and is not really from statistics. It is the observation that a computational result has two quite different kinds of uncertainty — the variance of its inputs and the arbitrariness of its construction — and that the literature reports the first and almost never the second. Under the name multiverse analysis the idea has had about ten years in the social sciences, where the constructions being varied are exclusions and covariates rather than colour-difference formulae. The finding there is the same as the finding here: the construction usually moves the answer more than the sampling does, and a paper reporting only a confidence interval is reporting the smaller of the two.
Where the ladder goes next
One of these six formulae was built on a space that this collection’s own uniformity instrument ranks last of the three it can rank. Whether the repair works is not a matter of opinion — MacAdam’s ellipses are a measurement of people, and turning that instrument on the units rather than on the spaces gives an answer with a sign nobody quotes.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A dial through a discrete menu calibration · colour difference · rank correlation · residual · test set
- A mean has a set under it chromatic adaptation · colour difference · residual · sampling · test set
- A mean is not a worst case chromatic adaptation · colour difference · residual · sampling · test set
- The instrument named the pair that moved chromatic adaptation · residual · sampling · standard error · test set
- Three choices reached calibration · chromatic adaptation · colour difference · test set · uncertainty
- A partial correction is worth its fraction chromatic adaptation · colour difference · residual · test set
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CalibrationChromatic adaptationColour differenceJackknifeRank correlationResidualSamplingStandard errorTest setUncertainty