What the eye does

The ranking is not stable

On a red pigment under daylight the six observer departures run from 2.38 down to 1.20 ΔE₀₀. Over forty-two surfaces two of them change places, the top two separate, and every one spans between a factor of ten and a factor of thirty-five. A chart of six bars is a chart of one example.

Assumes The lens is worst under tungsten, A departure is straight in the excitations and A mean has a set under it.

Six bars on one axis invite exactly one misreading, and the previous round had to close it for its own four. This one is worse, because here the bars are within a factor of two of one another and the reader’s eye supplies an ordering the data do not support.

Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.
Fig. 1 Each departure over forty-two surfaces, with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one spans more than a factor of three and two of them span more than fifteen.

The claim

The ordering of the six observer departures is a property of the sample they were measured on, and it changes.

  • On one red pigment: age 2.380, peaks 2.372, macular 1.714, rods 1.604, field 1.326, density 1.198 ΔE₀₀.
  • Over forty-two surfaces, at the median: age 1.943, macular 1.610, peaks 1.392, rods 1.028, field 0.740, density 0.566. Two rows have swapped and the top two have separated.
  • Every departure spans a factor of ten or more across the family, from 10 for the age to 35 for the density.
  • And the ranges overlap comprehensively. The macular’s worst case is 4.24 and the age’s best is 0.66, so a reader picking two surfaces could produce almost any ordering they wanted.

The ordering comes from where each departure sits on the wavelength axis. The previous essay established that a departure is proportional to the sample’s deviation from the light it is seen under, in the observer’s own cone excitations. That fixes the scaling and leaves the direction free, and the direction is what produces an ordering.

Each departure has a place on the wavelength axis. The macular pigment absorbs in a band at 460 nanometres; the lens absorbs below about 500 with a long tail; the peaks slide the whole of each curve a nanometre or two; the density broadens rather than shifts. A sample whose spectral deviation from the light sits at 460 pairs strongly with the macular departure and weakly with the peaks.

So the ordering on any one sample is a statement about where that sample’s structure is. A bar chart of six departures is a bar chart of one sample’s spectral shape, seen through six filters, and it says as much about the sample as about the eye.

The distribution, and how wide it is

Over the forty-two-surface family — seven band centres by three widths by two depths — the six departures are:

departure smallest median 95th largest span
the age of the lens 0.656 1.943 5.752 6.629 10×
the macular pigment 0.255 1.610 4.081 4.239 17×
the pigment peaks 0.349 1.392 2.401 2.931
the rods 0.311 1.028 1.748 1.859
the field size 0.207 0.740 2.477 3.154 15×
the cone optical density 0.082 0.566 1.707 2.900 35×

Two things about that table are more useful than the ordering.

The first is that the spans are of very different sizes. The rods span a factor of six and the density a factor of thirty-five, which means the two are quite differently predictable: quoting a single number for the rod contribution is defensible within a factor of two or three, and quoting one for the density is not.

The second is that the medians are much lower than the single-sample numbers. Every one of the six is smaller at the median of the family than it was on the red pigment, by between a quarter and a half. The red pigment is a fairly saturated sample and the family contains a good many pale ones, so the example this round has been quoting is above its own median throughout.

Six departures of the observer, each at a stated strength, under a 6500 K thermal radiator. Each bar is two observers differing in one argument, looking at the same sample under the same light, in ΔE₀₀. The strengths are the literature's: the working-age lens, two standard deviations of the reported macular and density spreads, the long-wavelength polymorphism, the CIE's own second observer, and a rod contribution of a tenth. They are within a factor of 5.3 of one another, which is the point: there is no single term to fix. Every one of them is above the ΔE of about one that a delivery tolerance is written in.
Fig. 2 The six departures on an interference filter rather than a red pigment. The ordering has changed again and the whole ladder has moved down.

Why the density spans thirty-five and the rods span six

The extremes of the span column have an explanation and it is the same mechanism read at two settings.

The density departure is mostly a gain, and a gain on each cone is not an observer at all. What survives the normalisation is the broadening — the part of a self-screening change that is not a scaling — and the broadening’s effect depends very sharply on how much fine structure the sample has near the edges of each cone’s sensitivity. A pale broad sample has almost none and reads 0.082; a narrow-band sample has a great deal and reads 2.900.

The rod departure adds a fourth curve to all three channels. An addition cannot be normalised away and it does not care about the sample’s structure in the same way: it shifts every stimulus by a fixed amount in a fixed direction, so what varies across samples is only how far that shift is from being a scaling. That produces a much tighter distribution.

A multiplicative departure has a wide distribution and an additive one has a narrow one, which is a general fact about pairings and is worth carrying: what a normalisation removes is exactly what would have made the answer stable across samples.

The three cone absorptances at two settings of the age of the lens. Solid and dashed are the same construction at the two ends of twenty years old against seventy. The curves are built from one pigment template through its ocular media, which is the same model its population of two hundred eyes is drawn from. The largest difference between the two sets is 25.5 per cent of the peak, and where it sits along the wavelength axis is what decides which stimuli the two observers disagree about — a departure concentrated in the blue is invisible on a sample with no blue in it.
Fig. 3 The three cone absorptances at twenty and at seventy years. The lens’s absorption is a long tail below 500 nanometres rather than a band, which is why its departure is largest on samples with structure anywhere in the short wavelengths.
The three cone absorptances at two settings of the rods. Solid and dashed are the same construction at the two ends of a tenth of the cone response, which is a dim room. The curves are built from one pigment template through its ocular media, which is the same model its population of two hundred eyes is drawn from. The largest difference between the two sets is 9.1 per cent of the peak, and where it sits along the wavelength axis is what decides which stimuli the two observers disagree about — a departure concentrated in the blue is invisible on a sample with no blue in it.
Fig. 4 The same three curves with and without a rod contribution of a tenth. What is added is one curve summed into all three channels, which is a different kind of change from a filter and produces a much narrower distribution across samples.

Those two figures are the two ends of the span column drawn. The lens is a filter with a long tail and a wide distribution; the rods are an addition with a fixed shape and a narrow one. Nothing about the sizes of the two departures predicts that difference, and nothing about the parameters’ reported spreads does either.

The general rule that comes out of it is worth stating because it is not obvious: a departure’s spread across samples is decided by how localised it is on the wavelength axis, not by how large it is. A narrow band — the macular’s forty nanometres at 460 — is either hit or missed by a given sample and spans a factor of seventeen. A broad change is hit by everything and spans less.

What a specification should take from this

The practical question is how much of a tolerance to reserve for observer variation, and the distribution answers it differently from the ladder.

Six departures at their median, combined in quadrature on the assumption of independence, give about 3.0 ΔE₀₀. At the ninety-fifth percentile of the family they give about 8.2. At the smallest members, about 0.9.

That is a range of nine to one for the same population of observers, decided entirely by what is being looked at. A specification reserving three units is adequate for a median sample and inadequate for a saturated one by a factor of nearly three; one reserving eight is adequate everywhere and would be regarded as absurdly loose.

The resolution is not a better single number. It is to make the reserve depend on the sample, which every tolerance system already does for other reasons — a tolerance is a shape rather than a radius and it is already chroma-dependent. Adding an observer term that scales with the sample’s deviation from the illuminant is a small change to an existing structure and it is not made anywhere.

Every departure under every light. Six departures across six lights, each cell the difference between two observers in ΔE₀₀, drawn as a bar whose length is the number. The rows are not multiples of one another: the lens is worst under tungsten and the pigment peaks are worst under a three-emitter LED, because a departure is a pairing and which light is being paired with decides it. The laser projector's row is empty, and that is not a fact about lasers — on this collection's five-nanometre grid a three-line spectrum is a one-line spectrum, and a single wavelength is a stimulus every observer agrees about exactly.
Fig. 5 Every departure under every light on an interference filter. The whole table is lower than the red pigment’s and its internal ordering is different again.

Why the median is the wrong summary too

The obvious response to a wide distribution is to quote its median, and this collection has an essay saying why that is not enough.

A mean is not a worst case, and for a quantity that decides whether a customer accepts a batch, the worst case is what matters. The medians above understate the ninety-fifth percentiles by a factor of between 1.7 and 3.0, and the worst cases by a little more.

But the deeper problem is the one the depth-08 round found about its own census: a median over a set is a statement about the set, and the set here is a construction. Forty-two analytic surfaces spanning seven band centres is a grid rather than a sample of anything, and no claim is made that real surfaces are distributed like it.

So the honest form of every number in this essay is conditional: over this family, the ordering is that one, the spans are those, and the quadrature sum at the median is three units. A different family — one drawn from a paint manufacturer’s catalogue, say, or from a set of skin tones — would give different medians and possibly a different ordering, and nothing here predicts by how much.

Each departure over forty-two surfaces rather than one. The same six departures measured over a family of forty-two analytic reflectances — an absorption band of stated centre, width and depth — with the smallest, the median, the ninety-fifth percentile and the largest marked. Every one of them spans more than a factor of three, and the ranking between them is not stable across the family: what decides a departure's size is which sample it is asked about, because a departure is a pairing and the sample is one of the two factors. Quoting any single number for what an observer's age is worth is quoting a choice of example.
Fig. 6 The same six distributions under a tungsten lamp. Every one has moved up and the top two have exchanged places, which is the light’s factor of the pairing moving the sample’s.

What the widest span means for a measurement programme

The density’s factor of thirty-five deserves a practical reading, because it inverts an obvious plan.

Anybody assembling a personalised observer has to decide which parameters to measure on the individual and which to take from a population mean. The natural criterion is the parameter’s reported spread, and by that criterion cone optical density is a strong candidate: its coefficient of variation is around nine per cent and it is genuinely variable between people.

The distribution says the criterion is wrong. The density’s median cost is the smallest of the six, so measuring it buys least on a typical sample — and its span is the widest, so on the sample where it matters it matters a great deal. A parameter with a low median and a wide span is the one whose value depends most on knowing the application, and it is the worst candidate for a general-purpose measurement and the best for a specific one.

The lens is the opposite: highest median, narrowest relative spread among the top three, and estimable from age alone with no measurement at all. It is the parameter to take from a population and the one it is least useful to measure.

Both factors at once

Changing the light moves the distributions as well as the ladders, and the two effects are not independent.

Under a tungsten lamp the macular departure’s median rises above the age’s, reversing the order that holds under daylight. That is the same reversal the single-sample table showed, and seeing it survive the move from one example to a family is what makes it a result rather than an artefact.

What does not survive is any single ordering. Across two lights and forty-two samples the six departures produce at least four different orderings of their top three, and no light-and-sample combination is more canonical than another. There is no fact of the matter about which observer parameter matters most, and the search for one is a search for a property of a pairing that belongs to neither of its factors.

The useful invariant is what is left when the ordering is abandoned: all six are of the same order, all six are above a delivery tolerance on a saturated sample, and none of them is dominant enough that fixing it alone would help — the same verdict the round’s opening ladder reached on one example, now supported by a set. That is a weaker statement than a ranking and it is the one the data support.

One more thing the distribution shows and the ladder cannot is that the six departures are not independent across samples. The surfaces on which the lens departure is largest are largely the surfaces on which the macular departure is largest, because both are blue absorptions and both pair with the same part of a sample’s structure. The correlation between those two columns across the forty-two members is strongly positive.

That matters for the quadrature sum quoted above, which assumed independence. Correlated departures add closer to linearly than in quadrature, so the reserve a specification needs on a blue-structured sample is larger than three units and closer to four. An extremum is not a sample and neither is a sum of medians; the honest quantity is the distribution of the combined departure over the family, and computing it means moving all six parameters at once, which this round does not do.

Why this had to be checked at all

A reader might reasonably ask why a ladder needed checking over a family when the previous essay established that the departure is linear in the sample’s deviation. If every departure scales with the same quantity, would they not all scale together and preserve their order?

They would if they scaled with the same quantity, and they do not. Each departure is an inner product with a different vector — the macular’s absorption band, the lens’s tail, the peaks’ derivative-like shift — so each one scales with the projection of the sample’s deviation onto its own direction. Six different projections of one sample, and a family of samples moves them by different amounts.

That is the whole reason a ranking is unstable while every individual departure is perfectly linear. Linearity in a scalar would preserve an order; linearity along six different directions does not, and the difference between those two situations is exactly the difference between a magnitude and a vector.

It is also why no reweighting of the family fixes the problem. A set of samples chosen to make the ordering stable would be a set chosen to have its deviation along one direction, which is a set of one sample repeated.

What was computed, and how

The family is seven band centres from 430 to 670 nanometres, three widths of 40, 80 and 140, and two depths of 0.35 and 0.70, on a base of 0.82. It is a grid, it is declared as one, and it is the same family the tabulation audit of this round uses — sharing it is what allows the two audits to be crossed.

Each departure is two observers differing in one parameter at its literature strength, computed through this collection’s usual CAT16 route, so the numbers are comparable with everything else here and carry that route’s own contribution.

Percentiles are computed by sorting and indexing rather than by interpolation, which for forty-two points means the ninety-fifth percentile is the fortieth value. That is coarse and it is stated rather than smoothed, because a smoothed percentile on forty-two points invents precision the set does not have.

Where the model stops

The family is analytic and every member is a single absorption band on a pale base. Real surfaces have several bands, and a sample with structure in two places pairs with two departures at once in a way no member of this family does.

The parameters are moved one at a time. Real observers differ in all six simultaneously, and the combination is not the quadrature sum used above unless the departures are orthogonal in stimulus space — which they are not, since the lens and the macular both absorb in the blue and their directions overlap substantially.

And the family is not weighted. Every one of the forty-two counts equally, which is a declaration that all forty-two are equally likely to be encountered, and that is certainly false for any real application.

One more caution belongs with the table and it is about the word spans. A factor of thirty-five sounds like an unstable measurement and it is not; every one of the forty-two values is computed exactly and reproducibly, and the spread is a property of the family rather than of the arithmetic. Measurement uncertainty and variation across a set are different quantities that a single error bar cannot distinguish, and reporting a span as though it were an uncertainty would be the same conflation the census had to unpick.

The generalisation

The habit is about what a bar chart claims.

A bar chart of k quantities makes an implicit claim that the quantities have single values. When each of them is really a distribution over some parameter the chart does not show, the ordering displayed is the ordering at one setting of that parameter, and the reader has no way to know how stable it is.

The repair is to draw the distribution rather than the value, which costs nothing and changes the reading completely. Here it turns “the lens is the largest departure” into “the lens has the highest median and the second-widest spread, and its lowest member is below the macular’s median”, which is a sentence nobody would mistake for a ranking.

The failure mode is to compute the chart on the sample that was to hand, publish the ordering, and have it repeated. A ladder that has not been checked over a set is a ladder of one example, and this round has now found four of them, including two of its own.

Where the ladder goes next

Both factors of the pairing have been varied and the ordering has come apart under each. What has not been varied is the arithmetic underneath, and one row of the light table is a zero that has nothing to do with eyes: the collection’s own grid removed an observer departure entirely, and it is the largest one in the table.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Individual variationMacular pigmentMeanMeasurement uncertaintyObserver metamerismSelf-screeningSpecificationStandard observerTest setWorst case