Matching and measuring

A deviation is not a difference

The previous round priced twenty departures in colour differences and treated each number as a property of the thing that departed. Every one of them is a product of two factors — how far the reading moved, which is linear and belongs to the departure, and what a unit of that movement is worth where it landed, which is not linear and belongs to the colour. Dimming one surface sixteen times scales the first by exactly sixteen and the second by eight.

Assumes A narrow primary buys a disagreement, The reference had to be built and Whose eyes.

The previous round audited what this collection’s own arithmetic assumes, and it reported its findings in colour differences. Six ways an observer can differ, four ways a wavelength grid can, four ways a measurement geometry can — each one a stated change, applied to a stated stimulus, and priced. The numbers were the round’s product and they were read as properties of the things that departed.

One surface dimmed sixteen times, and the two things that happen to it. A single surface, dimmed by successive halvings, with the lens age departure measured on it at every level. The tristimulus deviation falls by exactly the dimming factor — 16 times over the sweep, to the last bit, because the colour integral is linear in the stimulus. What a unit of that deviation is worth rises by 8.2 times over the same sweep. The colour difference the audit reports is the product of the two, and it falls by only 1.95.
Fig. 1 One surface, dimmed by successive halvings, with the same observer departure measured on it each time. The tristimulus deviation follows the dimming exactly and the colour difference does not, because between the two sits a price that is not a constant.

The claim

A colour difference is a product of two factors that were never separated: a deviation, which is linear and belongs to the departure, and a price, which is not and belongs to the colour.

  • The deviation is exact arithmetic. Halving a surface’s reflectance halves the tristimulus deviation an observer departure produces on it, to the last bit, sixteen times running.
  • The price is not. Over the same dimming it rises by a factor of 8.2, so the reported difference falls by 1.95 rather than by 16.
  • Over the surfaces the audit averaged across, the price of one departure spans a factor of between 5 and 10. Over a set with dark surfaces in it, between 36 and 81.
  • And the ranking is not one ranking. On the audit’s own forty-two surfaces the six departures come out in twenty-seven distinct orders, and four different departures are the largest on at least one of them.

What the previous round could rely on

The round that produced those numbers was able to be exact about a great deal, and the reason is worth restating because it is exactly the reason none of it survives here.

A tristimulus value is an integral: a spectrum multiplied by a curve and summed. That makes it linear in the stimulus and linear in the curve, and linear objects have properties that can be asserted rather than measured. Two changes to a stimulus add. A gain applied to the cones is exactly the same observer once the adaptation is done in the cones, which is what a von Kries hypothesis says written as arithmetic. A single wavelength is a stimulus every observer agrees about, to the floating-point floor. A neutral is observer-invariant for the same reason, and so is a stimulus judged against itself. None of those is approximately true; each is an identity, and the round tested each by handing it a case it had to refuse.

The file that computes those observers says so in its own words. The relative cone excitations, it records, are the quantity the pairing is linear in — and then, in the next sentence, that ΔE₀₀ is a function of them and is not a linear one.

That sentence is the whole of what follows. It was written as a caveat and it is a boundary. Everything to the left of the three numbers is an integral and can be audited by pairing two deviations; everything to the right of them is a compression, a chroma, a hue angle and a set of weighting functions, and none of that is linear in anything.

Six departures of the observer, each at a stated strength, under a 6500 K thermal radiator. Each bar is two observers differing in one argument, looking at the same sample under the same light, in ΔE₀₀. The strengths are the literature's: the working-age lens, two standard deviations of the reported macular and density spreads, the long-wavelength polymorphism, the CIE's own second observer, and a rod contribution of a tenth. They are within a factor of 2.0 of one another, which is the point: there is no single term to fix. Every one of them is above the ΔE of about one that a delivery tolerance is written in.
Fig. 2 The previous round’s six departures, ranked by what each costs. Every bar is a colour difference, and every colour difference is a product this essay takes apart.

The ladder above is the round’s summary and it is the object under examination. Read as a statement about eyes it says that a lens is the largest thing separating two observers and a cone optical density the smallest. Read as arithmetic it says something weaker: that on one sample, under one light, the product of deviation and price came out in that order.

The dimming

The cheapest way to separate the two factors is to change one of them and leave the other alone, and there is an operation that does exactly that.

Multiply a surface’s reflectance by a constant. Its shape is untouched — every ratio between one wavelength and another is what it was — so nothing about how two observers disagree about it has changed. What changes is how much light comes back, and because the integral is linear, the tristimulus deviation between the two observers’ readings is multiplied by the same constant.

The figure at the head of this essay is that experiment, run five times on one surface from the audit’s own family. The reflectance is halved, halved again, and twice more, and at each level the difference between a twenty-year-old lens and a seventy-year-old one is computed. The deviation goes 1.60 × 10⁻², 8.00 × 10⁻³, 4.00 × 10⁻³, 2.00 × 10⁻³, 1.00 × 10⁻³. That is a factor of sixteen, and the agreement with the dimming is to nine figures, which is what an identity looks like when a computer checks it.

The colour difference goes 0.663, 0.567, 0.483, 0.406, 0.340. That is a factor of 1.95.

The ratio between those two ratios is the price, and it moved by a factor of 8.2 while nothing about the eyes changed at all. The surface went from L* 82 to L* 22.9 and CIELAB’s compression is steeper at the bottom than at the top, so a movement of a given size in tristimulus values buys more lightness units down there.

Two columns that were one

Separating the factors is arithmetic rather than modelling: the deviation is the length of the tristimulus difference between the two observers’ adapted readings, and the price is the colour difference divided by that length. Both are already inside every number the audit published; neither was ever printed.

The other factor: how far each departure moves the reading. The same six departures, drawn by the size of the tristimulus deviation they produce rather than by what it costs. This is the linear half — a property of the two observers and of how much light the surface returns, and the quantity the previous round's identities are about. It spans a factor of 77 across 42 surfaces, and multiplying it by the price gives the number that was published.
Fig. 3 The first factor: how far each of the six departures moves the reading, over the forty-two surfaces the audit averaged across. This is the linear half, and it is a property of the two observers and of how much light the surface returns.

The deviations span a wide range and the range is legible. The lens and the macular pigment move the reading most, because both are filters in front of everything and a filter multiplies the whole spectrum. The cone optical density moves it least, because self-screening broadens a curve rather than raising it — a derivative rather than a displacement — and a broadening is a small perturbation to an integral.

Nothing in that column is surprising and nothing in it needs a metric. It is the answer to a question about light and pigment.

What one unit of deviation costs, over the surfaces it lands on. Each bar is one of the six observer departures, drawn from the cheapest surface in the set to the dearest, on a logarithmic axis. The quantity is colour differences per unit of tristimulus deviation — the price, which belongs to the colour and not to the eye. The narrowest spans a factor of 5 and the widest, field size, a factor of 10. The audit published one number for each of these, over 42 surfaces.
Fig. 4 The second factor: what a unit of deviation costs, over the same forty-two surfaces. The bar is the range and the point is the median. No departure has one price, and the widest spans a factor of ten.

The second column is the one that has no business being a range and is one anyway. It is measured in colour differences per unit of tristimulus deviation, and for a departure to have a price at all is already a statement: the quantity is not a property of the departure, since a departure is a pair of observers and this number changes when the surface does.

Across the audit’s own forty-two surfaces the field-size departure’s price runs from 36 to 355, a factor of ten. The narrowest range belongs to the cone optical density at a factor of five. So even inside the set the round actually used, the second factor is never better than a five-fold uncertainty on the first.

The other thing the separation buys is a way to ask whether a departure is large for what it is. A lens that has yellowed by fifty years moves the reading a long way and a cone optical density two standard deviations from the median moves it a short one, and the ratio between those two deviations is a statement about eyes that no metric can touch. It is 6.9 to 1 at the median of the audit’s own surfaces.

The colour differences the audit publishes put the same two at 2.40 and 0.72, a ratio of 3.3. Between 6.9 and 3.3 sits the price, which is higher for the density departure than for the lens on nearly every surface — 110 against 64 at the median — because the two move the reading in different directions and the metric charges differently for different directions.

So the audit’s headline ratio between its largest and smallest departure is half the underlying one, and the compression that halves it is not mentioned anywhere in the round.

The set the audit had

There is a reason those factors are as small as they are, and it is not a fact about colour.

Where the audit's surfaces sit, and where they do not. The forty-two surfaces every observer departure was averaged over, plotted by lightness and chroma, with the same family at four reflectance levels behind them. The audit's own set runs from L* 57.1 to 92.6: every one of them is a pale surface. Nothing chose that — the family was written to vary where a band sits and how wide it is, and the level was a constant because no question then being asked depended on it.
Fig. 5 Where the audit’s forty-two surfaces sit in lightness and chroma, with the same family at four reflectance levels behind them. Every one of the forty-two is a pale surface, and nothing decided that.

The family of surfaces the round averaged over is built as a constant reflectance of 0.82 with a Gaussian absorption band taken out of it. The band’s centre, its width and its depth all vary — seven centres, three widths, two depths, forty-two combinations — and the constant does not, because when that family was written the questions being asked were about where a feature sits and how wide it is, and the level was not one of the variables.

The consequence is that the darkest surface in the set is L* 57.1 and the median is L* 86.6. Nothing the audit measured was measured on a dark colour. That is not a mistake anybody made; it is a default that nobody had a reason to examine, which is the shape almost every finding in this collection’s self-audits has taken — and it is the same fault as a mean with an unexamined set under it.

Widening the family is one line: multiply the whole reflectance by a level, so each surface keeps its shape and changes only how much light it returns. Four levels give a hundred and sixty-eight surfaces reaching down to L* 18.3, and on that set the price of a departure spans a factor of between 36 and 81 rather than between 5 and 10.

The price is a function of how dark the surface is. Every surface in the set, with its lightness across the bottom and what a unit of the the macular pigment deviation costs on it up the side, logarithmic. The correlation is -0.77. A deviation that lands on a surface at L 20 is worth several times what the same deviation is worth at L 90, which is a statement about CIELAB's compression and not about anybody's eye.
Fig. 6 The price of the macular departure against the lightness of the surface it lands on, over the widened family, logarithmic. The correlation is −0.77, and on the audit’s own set it is −0.15 because there is not enough lightness in that set to see it.

The correlation between the log price and lightness runs from −0.77 to −0.86 across the six departures on the widened family. On the audit’s own forty-two it runs from −0.23 to +0.54, which is the same relationship measured through a window too narrow to contain it. A set spanning thirty-five lightness units cannot report a effect whose whole range is eighty.

So the audit’s numbers are not wrong and they are not general. They are correct statements about pale surfaces, and the price factor that would carry them to a dark one was never computed because the price was never named.

Which departure is the largest

The consequence that matters most is not a level but an ordering, because an ordering is what a reader takes away from an audit.

How many rankings six departures have, over one set of surfaces. The audit ranked its six departures once. Ranked separately on each of the 42 surfaces it was averaged over, there are 27 distinct orderings, and 4 different departures are the largest one on at least one surface — lens age on 20, pigment peaks on 12, macular on 8, cone density on 2. The eight commonest are drawn; the mean ordering, which is the one published, is lens age, then macular, then pigment peaks, then rods, then field size, then cone density.
Fig. 7 The six departures ranked separately on each of the forty-two surfaces. There are twenty-seven distinct orderings, and four different departures are the largest on at least one surface.

The published ranking is the lens, then the macular pigment, then the pigment peaks, then the rods, then the field size, then the cone optical density. It is the ranking of the means, and the means are honest.

Ranked surface by surface there are twenty-seven distinct orderings across forty-two surfaces, and the commonest of them occurs six times. The lens is the largest departure on twenty surfaces, the pigment peaks on twelve, the macular pigment on eight and the cone optical density on two. On the widened set the pattern is the same and firmer: thirty-seven orderings, and the same four departures taking the top place on 80, 46, 34 and 8 surfaces.

That is not noise around a stable answer. It is a different answer on a quarter of the surfaces, and the mechanism is not mysterious: the deviations point in different directions, the price is direction-dependent as well as position-dependent, and a surface that puts one departure’s deviation along a cheap direction and another’s along a dear one reorders them.

A reader who takes away the lens is the largest observer difference has taken away a statement that is false on half the surfaces they will meet. The correct statement is that the lens has the largest mean over a set of pale surfaces, and that which departure dominates is a joint property of the eye pair and the colour.

Carrying a number to another surface

The separation is worth something practical, and the practical thing is a conversion.

An audit number is a colour difference measured on one surface. A reader wants it on theirs. The deviation half transfers by a ratio of light returns, which is arithmetic anybody can do from a reflectance; the price half transfers by a ratio of prices, and that ratio is the thing this essay computes and the audit did not.

Worked on the dimming above: the lens departure on that surface at L* 82 is 0.663 colour differences. The same surface at L* 22.9 returns a sixteenth of the light, so the deviation is a sixteenth. The price at L* 22.9 is 8.2 times the price at L* 82. The product is 0.663 × (1/16) × 8.2, which is 0.340, and that is the measured value to three figures.

Two factors, each computed independently, multiplying to the number the formula reports. That is the whole claim demonstrated on one case, and it is a claim that could have failed: if the price were not well defined — if it depended on the direction of the deviation as well as its size — the product would not have come out. It does come out here because the dimming leaves the direction alone.

It does not come out in general, and the reason is the subject of the next rung. A departure that changes direction as well as size has a price that is not a single number, and the two-factor decomposition becomes a three-factor one with an angle in it.

The third factor, named and not measured

There is a factor this essay has held fixed throughout and it is not small.

The audit runs its departures under six lights, from a tungsten radiator to a three-laser projector, and reports the spread. Everything above is computed under one of them. The price depends on the light, because the light decides where the adapted reading sits, and a surface under a tungsten lamp and the same surface under a laser projector are two different colours in the space the price is measured in.

So the audit’s numbers are a product of three things and this essay has separated two. The third is worth naming precisely because naming it is what makes the omission a shortfall rather than an oversight: the price under each of the six lights, over the same surface set, is one loop over machinery that already exists, and it would say whether the ranking of lights the audit publishes is a ranking of lights or a ranking of where those lights put the samples.

The prediction, stated so that it can be wrong, is that the narrow-band lights will show a smaller price spread than the smooth ones, because a narrow-band light lands its samples in a smaller region of the space. If that is right, the audit’s finding that observers disagree most under narrow-band sources is understated by its own arithmetic.

What was computed, and how

The deviations are computed the way the audit computes its own numbers: each observer forms its own tristimulus values from the same spectral power distribution, normalises to its own white, and adapts to D65 through CAT16. Skipping the adaptation would report the difference between two observers’ whites as a difference between two samples, which is the largest thing in the calculation and is not a departure at all — every observer agrees about the white it adapted to, exactly.

The price is the ΔE₀₀ between those two adapted readings divided by the Euclidean length of the tristimulus difference. Nothing about that choice of length is canonical; a different norm would rescale every price by a bounded factor and leave every ratio in this essay where it is, because the essay’s quantities are all ratios of prices.

The dimming experiment uses one surface from the audit’s family, and the identity it demonstrates is checked in code rather than asserted: the ratio of successive deviations divided by the ratio of successive dimmings is required to be one to within a millionth, and the machinery refuses to draw the figure otherwise.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 8 This collection’s inventory of what it computes differences with. Every number in the audit passes through one of these, and every one of them is a nonlinear function of the three tristimulus values.

The formula used throughout is ΔE₀₀, because that is the one the audit used. It is the least linear of the menu — a lightness compression, a chroma-dependent weighting, a hue-dependent weighting and a rotation term — and it would be reasonable to suspect the whole effect of being an artefact of the formula rather than of the space.

It is not. Running the dimming experiment in ΔE*ab, which is a plain Euclidean distance in CIELAB, still gives a price that moves, because the cube root is already there. The prices change and the structure does not, which is the general statement: any difference formula that is not a Euclidean distance in tristimulus values has a varying price, and a formula that is one is the thing colour science spent a century demonstrating does not describe anybody.

What this does not say

It does not say the audit should have reported the price instead. A reader asking whether two observers will disagree about a printed sample wants the product, because the product is the disagreement.

It says the product should not be read as the first factor. A number that is a property of a pair of eyes and a number that is a property of a colour look identical on the page, and the audit’s tables are full of the second kind labelled as the first.

The practical form of that is a rule about transfer. An audit number measured on a pale surface transfers to a dark one by multiplying by the ratio of the two prices, and that ratio is between two and eight over the range of surfaces anybody delivers, which is more than the spread between the two standard observers. Nothing in the round said so, because nothing in the round had the second column.

Where the model stops

The surfaces are analytic and the widening is a scalar multiplication, which is the cleanest possible way to change lightness and not a realistic one. Real dark surfaces are dark because they absorb selectively, so a real set would move chroma as well and the two effects would not separate as neatly as they do here.

The price is defined against a Euclidean length in adapted tristimulus values, which is a coordinate-dependent choice. Every ratio survives it; no absolute price should be quoted outside this collection.

And the whole of this is one light. The audit runs its departures under six lights and the price depends on the light through the adapted reading, so there is a third factor here that has not been separated from the other two. Its size is not measured in this essay and is not small.

The generalisation

The habit is about a number that is a product of two things when only one of them is named.

An audit, a benchmark, a score — anything that reports one figure for a change applied to a system — is nearly always a composition: how far the change moved the system, and what the reporting instrument charges for movement there. The first belongs to the change and transfers; the second belongs to the operating point and does not. A table of such numbers reads as though every entry were the first kind.

The move is to divide before reporting. It costs one extra column, the column is already computable from what was measured, and it converts a table of results into a table of results plus the conversion factor a reader needs to apply them anywhere else.

The failure mode is subtler than reporting a wrong number, and it is what happened here. Every number was right and the ordering they were read for was a property of the sample set, which nobody had chosen and nobody could see.

Who found it, and when

That colour differences are not uniform is the oldest measured fact in this subject: MacAdam’s ellipses date from 1942 and every uniform space since has been an attempt to make them rounder. That the residual non-uniformity varies over the solid by a large factor is likewise standard, and this collection has minimised its anisotropy over every chromaticity diagram there is and found the floor.

What does not appear to be standard is applying it to an audit’s own output — treating a published sensitivity as a product and asking which factor a reader is entitled to carry away. The arithmetic is elementary and the practice is not common, in this collection or elsewhere.

Where the ladder goes next

Separating the two factors leaves a question the separation itself raises. If the price depends on direction as well as on position, then two departures measured on the same surface cannot be combined by their sizes alone — and the audit combined them by their sizes, because a size is all it recorded. What two departures cost together needs the angle between them, and the angle has never been computed.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 10 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AuditCIEDE2000CIELABColour differenceDeclared inputLightnessMetric axiomsObserver variabilitySensitivityTristimulus