The ranking is not stable
Assumes The lens is worst under tungsten, A departure is straight in the excitations and A mean has a set under it.
Six bars on one axis invite exactly one misreading, and the previous round had to close it for its own four. This one is worse, because here the bars are within a factor of two of one another and the reader’s eye supplies an ordering the data do not support.
The claim
The ordering of the six observer departures is a property of the sample they were measured on, and it changes.
- On one red pigment: age 2.380, peaks 2.372, macular 1.714, rods 1.604, field 1.326, density 1.198 ΔE₀₀.
- Over forty-two surfaces, at the median: age 1.943, macular 1.610, peaks 1.392, rods 1.028, field 0.740, density 0.566. Two rows have swapped and the top two have separated.
- Every departure spans a factor of ten or more across the family, from 10 for the age to 35 for the density.
- And the ranges overlap comprehensively. The macular’s worst case is 4.24 and the age’s best is 0.66, so a reader picking two surfaces could produce almost any ordering they wanted.
The ordering comes from where each departure sits on the wavelength axis. The previous essay established that a departure is proportional to the sample’s deviation from the light it is seen under, in the observer’s own cone excitations. That fixes the scaling and leaves the direction free, and the direction is what produces an ordering.
Each departure has a place on the wavelength axis. The macular pigment absorbs in a band at 460 nanometres; the lens absorbs below about 500 with a long tail; the peaks slide the whole of each curve a nanometre or two; the density broadens rather than shifts. A sample whose spectral deviation from the light sits at 460 pairs strongly with the macular departure and weakly with the peaks.
So the ordering on any one sample is a statement about where that sample’s structure is. A bar chart of six departures is a bar chart of one sample’s spectral shape, seen through six filters, and it says as much about the sample as about the eye.
The distribution, and how wide it is
Over the forty-two-surface family — seven band centres by three widths by two depths — the six departures are:
| departure | smallest | median | 95th | largest | span |
|---|---|---|---|---|---|
| the age of the lens | 0.656 | 1.943 | 5.752 | 6.629 | 10× |
| the macular pigment | 0.255 | 1.610 | 4.081 | 4.239 | 17× |
| the pigment peaks | 0.349 | 1.392 | 2.401 | 2.931 | 8× |
| the rods | 0.311 | 1.028 | 1.748 | 1.859 | 6× |
| the field size | 0.207 | 0.740 | 2.477 | 3.154 | 15× |
| the cone optical density | 0.082 | 0.566 | 1.707 | 2.900 | 35× |
Two things about that table are more useful than the ordering.
The first is that the spans are of very different sizes. The rods span a factor of six and the density a factor of thirty-five, which means the two are quite differently predictable: quoting a single number for the rod contribution is defensible within a factor of two or three, and quoting one for the density is not.
The second is that the medians are much lower than the single-sample numbers. Every one of the six is smaller at the median of the family than it was on the red pigment, by between a quarter and a half. The red pigment is a fairly saturated sample and the family contains a good many pale ones, so the example this round has been quoting is above its own median throughout.
Why the density spans thirty-five and the rods span six
The extremes of the span column have an explanation and it is the same mechanism read at two settings.
The density departure is mostly a gain, and a gain on each cone is not an observer at all. What survives the normalisation is the broadening — the part of a self-screening change that is not a scaling — and the broadening’s effect depends very sharply on how much fine structure the sample has near the edges of each cone’s sensitivity. A pale broad sample has almost none and reads 0.082; a narrow-band sample has a great deal and reads 2.900.
The rod departure adds a fourth curve to all three channels. An addition cannot be normalised away and it does not care about the sample’s structure in the same way: it shifts every stimulus by a fixed amount in a fixed direction, so what varies across samples is only how far that shift is from being a scaling. That produces a much tighter distribution.
A multiplicative departure has a wide distribution and an additive one has a narrow one, which is a general fact about pairings and is worth carrying: what a normalisation removes is exactly what would have made the answer stable across samples.
Those two figures are the two ends of the span column drawn. The lens is a filter with a long tail and a wide distribution; the rods are an addition with a fixed shape and a narrow one. Nothing about the sizes of the two departures predicts that difference, and nothing about the parameters’ reported spreads does either.
The general rule that comes out of it is worth stating because it is not obvious: a departure’s spread across samples is decided by how localised it is on the wavelength axis, not by how large it is. A narrow band — the macular’s forty nanometres at 460 — is either hit or missed by a given sample and spans a factor of seventeen. A broad change is hit by everything and spans less.
What a specification should take from this
The practical question is how much of a tolerance to reserve for observer variation, and the distribution answers it differently from the ladder.
Six departures at their median, combined in quadrature on the assumption of independence, give about 3.0 ΔE₀₀. At the ninety-fifth percentile of the family they give about 8.2. At the smallest members, about 0.9.
That is a range of nine to one for the same population of observers, decided entirely by what is being looked at. A specification reserving three units is adequate for a median sample and inadequate for a saturated one by a factor of nearly three; one reserving eight is adequate everywhere and would be regarded as absurdly loose.
The resolution is not a better single number. It is to make the reserve depend on the sample, which every tolerance system already does for other reasons — a tolerance is a shape rather than a radius and it is already chroma-dependent. Adding an observer term that scales with the sample’s deviation from the illuminant is a small change to an existing structure and it is not made anywhere.
Why the median is the wrong summary too
The obvious response to a wide distribution is to quote its median, and this collection has an essay saying why that is not enough.
A mean is not a worst case, and for a quantity that decides whether a customer accepts a batch, the worst case is what matters. The medians above understate the ninety-fifth percentiles by a factor of between 1.7 and 3.0, and the worst cases by a little more.
But the deeper problem is the one the depth-08 round found about its own census: a median over a set is a statement about the set, and the set here is a construction. Forty-two analytic surfaces spanning seven band centres is a grid rather than a sample of anything, and no claim is made that real surfaces are distributed like it.
So the honest form of every number in this essay is conditional: over this family, the ordering is that one, the spans are those, and the quadrature sum at the median is three units. A different family — one drawn from a paint manufacturer’s catalogue, say, or from a set of skin tones — would give different medians and possibly a different ordering, and nothing here predicts by how much.
What the widest span means for a measurement programme
The density’s factor of thirty-five deserves a practical reading, because it inverts an obvious plan.
Anybody assembling a personalised observer has to decide which parameters to measure on the individual and which to take from a population mean. The natural criterion is the parameter’s reported spread, and by that criterion cone optical density is a strong candidate: its coefficient of variation is around nine per cent and it is genuinely variable between people.
The distribution says the criterion is wrong. The density’s median cost is the smallest of the six, so measuring it buys least on a typical sample — and its span is the widest, so on the sample where it matters it matters a great deal. A parameter with a low median and a wide span is the one whose value depends most on knowing the application, and it is the worst candidate for a general-purpose measurement and the best for a specific one.
The lens is the opposite: highest median, narrowest relative spread among the top three, and estimable from age alone with no measurement at all. It is the parameter to take from a population and the one it is least useful to measure.
Both factors at once
Changing the light moves the distributions as well as the ladders, and the two effects are not independent.
Under a tungsten lamp the macular departure’s median rises above the age’s, reversing the order that holds under daylight. That is the same reversal the single-sample table showed, and seeing it survive the move from one example to a family is what makes it a result rather than an artefact.
What does not survive is any single ordering. Across two lights and forty-two samples the six departures produce at least four different orderings of their top three, and no light-and-sample combination is more canonical than another. There is no fact of the matter about which observer parameter matters most, and the search for one is a search for a property of a pairing that belongs to neither of its factors.
The useful invariant is what is left when the ordering is abandoned: all six are of the same order, all six are above a delivery tolerance on a saturated sample, and none of them is dominant enough that fixing it alone would help — the same verdict the round’s opening ladder reached on one example, now supported by a set. That is a weaker statement than a ranking and it is the one the data support.
One more thing the distribution shows and the ladder cannot is that the six departures are not independent across samples. The surfaces on which the lens departure is largest are largely the surfaces on which the macular departure is largest, because both are blue absorptions and both pair with the same part of a sample’s structure. The correlation between those two columns across the forty-two members is strongly positive.
That matters for the quadrature sum quoted above, which assumed independence. Correlated departures add closer to linearly than in quadrature, so the reserve a specification needs on a blue-structured sample is larger than three units and closer to four. An extremum is not a sample and neither is a sum of medians; the honest quantity is the distribution of the combined departure over the family, and computing it means moving all six parameters at once, which this round does not do.
Why this had to be checked at all
A reader might reasonably ask why a ladder needed checking over a family when the previous essay established that the departure is linear in the sample’s deviation. If every departure scales with the same quantity, would they not all scale together and preserve their order?
They would if they scaled with the same quantity, and they do not. Each departure is an inner product with a different vector — the macular’s absorption band, the lens’s tail, the peaks’ derivative-like shift — so each one scales with the projection of the sample’s deviation onto its own direction. Six different projections of one sample, and a family of samples moves them by different amounts.
That is the whole reason a ranking is unstable while every individual departure is perfectly linear. Linearity in a scalar would preserve an order; linearity along six different directions does not, and the difference between those two situations is exactly the difference between a magnitude and a vector.
It is also why no reweighting of the family fixes the problem. A set of samples chosen to make the ordering stable would be a set chosen to have its deviation along one direction, which is a set of one sample repeated.
What was computed, and how
The family is seven band centres from 430 to 670 nanometres, three widths of 40, 80 and 140, and two depths of 0.35 and 0.70, on a base of 0.82. It is a grid, it is declared as one, and it is the same family the tabulation audit of this round uses — sharing it is what allows the two audits to be crossed.
Each departure is two observers differing in one parameter at its literature strength, computed through this collection’s usual CAT16 route, so the numbers are comparable with everything else here and carry that route’s own contribution.
Percentiles are computed by sorting and indexing rather than by interpolation, which for forty-two points means the ninety-fifth percentile is the fortieth value. That is coarse and it is stated rather than smoothed, because a smoothed percentile on forty-two points invents precision the set does not have.
Where the model stops
The family is analytic and every member is a single absorption band on a pale base. Real surfaces have several bands, and a sample with structure in two places pairs with two departures at once in a way no member of this family does.
The parameters are moved one at a time. Real observers differ in all six simultaneously, and the combination is not the quadrature sum used above unless the departures are orthogonal in stimulus space — which they are not, since the lens and the macular both absorb in the blue and their directions overlap substantially.
And the family is not weighted. Every one of the forty-two counts equally, which is a declaration that all forty-two are equally likely to be encountered, and that is certainly false for any real application.
One more caution belongs with the table and it is about the word spans. A factor of thirty-five sounds like an unstable measurement and it is not; every one of the forty-two values is computed exactly and reproducibly, and the spread is a property of the family rather than of the arithmetic. Measurement uncertainty and variation across a set are different quantities that a single error bar cannot distinguish, and reporting a span as though it were an uncertainty would be the same conflation the census had to unpick.
The generalisation
The habit is about what a bar chart claims.
A bar chart of k quantities makes an implicit claim that the quantities have single values. When each of them is really a distribution over some parameter the chart does not show, the ordering displayed is the ordering at one setting of that parameter, and the reader has no way to know how stable it is.
The repair is to draw the distribution rather than the value, which costs nothing and changes the reading completely. Here it turns “the lens is the largest departure” into “the lens has the highest median and the second-widest spread, and its lowest member is below the macular’s median”, which is a sentence nobody would mistake for a ranking.
The failure mode is to compute the chart on the sample that was to hand, publish the ordering, and have it repeated. A ladder that has not been checked over a set is a ladder of one example, and this round has now found four of them, including two of its own.
Where the ladder goes next
Both factors of the pairing have been varied and the ordering has come apart under each. What has not been varied is the arithmetic underneath, and one row of the light table is a zero that has nothing to do with eyes: the collection’s own grid removed an observer departure entirely, and it is the largest one in the table.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The observer has no age individual variation · macular pigment · measurement uncertainty · observer metamerism · specification · standard observer
- A field size is two changes individual variation · macular pigment · self-screening · specification · standard observer
- The instrument is one observer exactly individual variation · measurement uncertainty · observer metamerism · specification · standard observer
- The macular is a band, not a filter individual variation · macular pigment · observer metamerism · standard observer · test set
- Two yellow filters cancel on a slope individual variation · macular pigment · observer metamerism · standard observer · test set
- A cone absorbs its own light individual variation · observer metamerism · self-screening · standard observer
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Individual variationMacular pigmentMeanMeasurement uncertaintyObserver metamerismSelf-screeningSpecificationStandard observerTest setWorst case