Matching and measuring

Which mixture bows most depends on the ruler

The physical half-and-half mixture of two lights misses the midpoint of their two readings in any unit a person is shown, and the share it misses by is about the same in ΔE₀₀ and in the appearance model's own space — 12.8 and 14.6 per cent at the median. Which mixtures miss most is not. The rank correlation between the two units is 0.57, a display's blue and yellow is the worst pair in one and the mildest in the other, and the smaller number usually quoted for the appearance model is its power correction rather than a straighter line.

Assumes The mixture line bows, A model judged in another model's unit and The census in six units.

Grassmann’s second law makes an additive mixture exactly linear in tristimulus values, and the mixture line bows the moment it is read in any coordinate a person is shown. Measured in ΔE₀₀ the half-and-half mixture of two colours misses the midpoint of their two readings by a median of 5.5 colour differences, and a display’s green and blue miss by 21. That measurement ended by noting that other units give other numbers, and that the appearance model’s gives smaller ones. The second half of that sentence turns out to be a statement about an exponent.

Four mixtures of display primaries, bowing by different amounts in two units. The half-and-half mixture of each pair of sRGB primaries, measured from the midpoint of the two readings as a share of the pair's own separation. The upper bar is ΔE₀₀ and the lower is the appearance model's J′a′b′. In ΔE₀₀ green and blue bows most, at 25 per cent; in the model it is red and blue, at 20, and blue and yellow falls from 22 to 12. The right-hand column gives the model's distance with its power correction, which is smaller than the Euclidean one for every pair here.
Fig. 1 The half mixture of each pair of sRGB primaries, as a share of the pair’s own separation, in ΔE₀₀ and in the appearance model’s own space. The pair that bows most in one unit is not the pair that bows most in the other.

The same bow, a different ranking

How far a mixture bows, as a share of its pair’s separation, is nearly the same in the two units at the median. Which mixtures bow most is not.

  • Over 1,450 random pairs inside sRGB the median share is 12.8 per cent in ΔE₀₀ and 14.6 in the appearance model’s J′a′b′.
  • The rank correlation between the two is 0.57. Of the tenth of pairs that bow most in ΔE₀₀, 26 per cent are also in the model’s top tenth.
  • On a display’s primaries the ranking reverses. Green and blue bows most in ΔE₀₀, at 25 per cent; red and blue bows most in the model, at 20. Blue and yellow falls from 22 per cent to 12.
  • The model’s distance with its power correction is smaller than the Euclidean one for every pair of primaries, because the correction shrinks every difference above two and a half units. That is where a smaller number for the model comes from.

Two readings of one mixture

The physical event is the same in both units. Two lights are switched on together at complementary levels; their tristimulus values add; the mixture walks a straight line in tristimulus values from one light to the other. What differs is only the coordinate the walk is read in.

ΔE₀₀ reads it in CIELAB: a cube-root compression of each tristimulus ratio, differences taken with coefficients 116, 500 and 200, and then a difference formula that weights lightness, chroma and hue by where the pair sits. The appearance model reads it in J′a′b′: adapted cone signals, a hyperbolic compression, opponent signals, a lightness from the achromatic signal, a colourfulness from the chromatic ones, and a logarithmic compression of colourfulness. Both are compressions and both bow a straight line. They compress different things by different amounts.

Two primaries mixed, and the line a reader assumes they take. The additive mixture of two sRGB primaries, walked in twenty steps, plotted in the a and b of CIELAB. The filled points are where the light actually goes, which is exactly straight in tristimulus values because that is Grassmann's second law. The open points are the straight line between the two readings. They part company by 28.2 ΔE₀₀ at their furthest, at 55 per cent of the way along, and the half-and-half mixture misses the midpoint by 23.2.
Fig. 2 The mixture of a display’s blue primary and yellow — its red and green together — in CIELAB’s a* and b*, against the straight line between the two readings. In this unit it is the worst of the four primary pairs.

In ΔE₀₀ the blue and yellow half mixture misses its midpoint by 23.2 colour differences, across a separation of 103. In the model the same mixture misses by 12.4 units across 107. The separations are almost the same size and the bows differ by a factor of two, in opposite directions for different pairs: red and green bows 8.2 in ΔE₀₀ and 17.8 in the model.

The whole cube, pair by pair

Four pairs of primaries might be four exceptions. The general statement needs pairs from everywhere in the cube, and the same pairs in both units.

Each mixture's bow in one unit against the same mixture's bow in the other. 1450 random pairs of colours inside sRGB — the same pairs, drawn with the same seed, in both units. Across the bottom is the half mixture's miss as a share of the pair's separation in ΔE₀₀, up the side the same share in the appearance model's J′a′b′. The medians are 12.8 and 14.6 per cent, nearly the same, and the rank correlation between the two is only 0.57: of the tenth of pairs that bow most in ΔE₀₀, 26 per cent are also in the model's top tenth.
Fig. 3 Each of 1,450 pairs’ bow as a share of its separation in ΔE₀₀, across, against the same pair’s share in the model’s space, up. The cloud sits on the diagonal at its middle and spreads far from it everywhere else.

The 1,450 pairs are drawn with the same seed that produced the ΔE₀₀ census, so every point is one mixture read twice. The medians agree closely — 12.8 per cent in ΔE₀₀ and 14.6 in the model — and the ninetieth percentiles agree too, 23.5 and 25.4. A statement about how much mixtures bow in general transfers between the units to within a couple of points.

A statement about which mixture bows does not. The rank correlation between the two shares is 0.57, and of the 145 pairs that bow most in ΔE₀₀ only 38 are also among the 145 that bow most in the model. A designer who has learned from ΔE₀₀ which gradients need an intermediate colour stop has learned a list that is three quarters wrong in the other unit.

Better than chance, far from the same

A rank correlation of 0.57 is easy to read as agreement or as disagreement, and the overlap of the two worst tenths puts a scale on it.

If the two rankings were unrelated, the 145 pairs that bow most in ΔE₀₀ would share about fifteen members with the model’s worst 145 by chance alone — a tenth of them. They share 38: two and a half times what chance gives, and a quarter of what agreement would. The two units are measuring one phenomenon, and which cases sit at its extreme is mostly decided by the unit.

Put as the choice a designer faces, a list of the worst tenth drawn up in one unit holds 107 pairs the other unit does not place in its worst tenth, and leaves out 107 that it does. Neither list is wrong in its own unit. A tool that shows a designer one of them has chosen the unit on the designer’s behalf, and nothing in the list says so.

Why the medians agree at all

That the medians agree to two points is not evidence that the units agree about mixtures, and it is not a coincidence either.

The likely reading is this. Both units were built for one job — to make equal numbers mean roughly equal visible differences across ordinary colours — and both were fitted or tuned against judgements of pairs. A unit that is roughly uniform must bend a straight line in tristimulus values by roughly the amount that makes such a line look uneven, and a median over fourteen hundred pairs is a statement about that average amount. Where each unit puts its bending is a separate matter. CIELAB compresses each tristimulus ratio on its own before it forms opponent differences; the model compresses cone signals, forms its opponent signals from them, and then compresses colourfulness logarithmically. The average is fixed by the shared job, and the distribution over colours by the construction, which is exactly what differs.

How far the half-and-half mixture misses, over a thousand pairs. 1450 random pairs of colours inside sRGB, at least ten colour differences apart. Across the bottom is how far apart the pair is; up the side is how far the physical half-and-half mixture lands from the midpoint of the two readings. The median is 5.5 ΔE₀₀ and the worst is 30. As a share of the pair's own separation it is 15 per cent at the median and 49 at the worst.
Fig. 4 The ΔE₀₀ census the pairs above were drawn from: the half mixture’s miss against the pair’s separation. Its median share is the 12.8 per cent the model very nearly reproduces, with a different set of pairs in each part of the cloud.

Why the smaller number was the exponent

The claim that the appearance model gives smaller bows came from comparing the model’s distance with ΔE₀₀, and the model’s distance, as used throughout the appearance essays here, is not the Euclidean distance in J′a′b′. It is that distance raised to the power 0.63 and multiplied by 1.41 — a correction fitted so that the unit’s small and large differences agree better with how people scale them.

The power correction, against the distance it corrects. The model's distance with its power correction, 1.41 times the Euclidean distance to the power 0.63, against the Euclidean distance itself, both axes logarithmic, with the identity drawn faint. The two agree at exactly one size, 2.53 units. Below it the correction makes a difference larger, above it smaller — a difference of 20 becomes 9.3 — and on a logarithmic plot the correction is a straight line of slope 0.63, which is the whole of its effect on anything added up in it.
Fig. 5 The power correction against the Euclidean distance it corrects, both logarithmic. Above 2.53 units the corrected number is smaller than the distance, and a bow of twenty becomes nine.

The corrected and uncorrected distances are equal at exactly 2.53 units; below that the correction makes a difference larger, above it smaller. Every bow on a pair of display primaries is well above 2.53, so every one is reported smaller with the correction than without: red and green’s 17.8 becomes 8.7, green and blue’s 17.1 becomes 8.4. Over the census the median bow is 5.52 in ΔE₀₀, 7.58 in J′a′b′ and 5.05 with the correction.

So of the two numbers that could be quoted for the model, the Euclidean one is larger than ΔE₀₀ at the median and the corrected one is smaller. The smaller number measures the correction, not the space’s straightness. The model’s space bends an additive mixture by slightly more than CIELAB-based difference does, at the median, and by very different amounts on particular pairs. A distance raised to a power has no length, and a bow read along its whole path, rather than at its midpoint, is a length.

A smaller bow and a larger share

The correction has a second effect on a bow, and it points the other way.

A share divides the bow by the pair’s separation. If both are taken in the corrected unit the factor of 1.41 cancels, and the corrected share is the Euclidean share raised to the power 0.63. A power below one pulls every number below one towards one, so a corrected share is always larger than the Euclidean share it came from, and the smaller the share the larger the increase.

The model’s median Euclidean share of 14.6 per cent becomes 29.8 per cent in the corrected unit. So the model’s mixtures, reported with the correction, bow by smaller distances than ΔE₀₀’s and by more than twice ΔE₀₀’s share — two numbers from one calculation that point in opposite directions, either of which could be quoted as the model’s verdict on mixtures. The ranking survives the correction, because raising to a power preserves order. The size of the effect relative to its own pair does not.

What the room does to it

The appearance model has an argument ΔE₀₀ does not: the room. A mixture’s bow in the model can be asked how it changes when the light level and the surround change, which in ΔE₀₀ is not a question at all.

One mixture in five rooms: red and green. The half mixture of red and green sRGB primaries in the appearance model's space, from a dark room at 4 candelas a square metre to daylight at 10,000. The faint bar is the pair's separation and the filled bar inside it is the bow. Both grow by a half across the range, because the model's colourfulness rises with the light, and the share stays between 19.0 and 20.2 per cent. The room scales a mixture's bow and leaves its shape.
Fig. 6 The red and green half mixture in five rooms, from a dark room at 4 cd/m² to daylight at 10,000. The bow grows with the separation it sits inside, and its share barely moves.

For red and green, the bow grows from 14.1 units in a dark room at 4 candelas a square metre to 22.9 in daylight at 10,000, and the separation from 73.5 to 113.1, so the share stays between 19.0 and 20.2 per cent. Green and blue behaves the same way, 15.9 to 17.5 per cent. The model’s colourfulness rises with the luminance level — brighter looks more colourful — and it rises for the pair and for its mixture together.

One mixture in five rooms: blue and yellow. The half mixture of blue and yellow sRGB primaries in the appearance model's space, from a dark room at 4 candelas a square metre to daylight at 10,000. The faint bar is the pair's separation and the filled bar inside it is the bow. Both grow by a half across the range, because the model's colourfulness rises with the light, and the share stays between 10.3 and 13.3 per cent. The room scales a mixture's bow and leaves its shape.
Fig. 7 Blue and yellow in the same five rooms. Here the share falls as the light rises, from 13.3 per cent in the dark room to 10.3 in daylight — a pair whose bow is less a matter of colourfulness than the others’.

Blue and yellow is the exception. Its bow grows only from 11.3 to 12.9 units while its separation grows from 85 to 126, so its share falls from 13.3 per cent to 10.3. Blue and yellow differ mostly in lightness, and the model’s lightness is relative to the white and nearly immune to the room, while its colourfulness is not. A pair whose bow is carried by lightness keeps its bow as the room brightens and gains separation; a pair whose bow is carried by colourfulness gains both. The room reorders the pairs too, slightly, and ΔE₀₀ has no way to say so.

On a wider gamut

The mixture line bows found that moving to Display P3’s primaries barely changes the blue-and-yellow bow in ΔE₀₀ — 23.0 against 23.2 — and read that as evidence that the bending belongs to the metric rather than the display. The model agrees about that pair and not about the others.

On P3 the model’s blue and yellow bow is 11.8 against 12.4 on sRGB, as stable as ΔE₀₀’s. Its red and green bow rises, from 17.8 to 19.6, where ΔE₀₀’s falls by nearly half, from 8.2 to 4.6. A wider red and green bows less in one unit and more in the other. The conclusion that the bend belongs to the ruler survives; which ruler is doing the bending decides the direction a gamut change moves it.

A stop placed by one ruler

A gradient that bows is usually repaired with a stop: a third colour inserted where the gradient ought to pass halfway, so that the rendered gradient runs through it. The stop is computed in some unit, as the midpoint of the two ends’ readings, and turned back into a colour.

In the unit that placed it, the repaired gradient’s midpoint no longer misses. In the other unit it misses by the distance between the two units’ midpoints, and those are different colours, because a midpoint taken in one set of coordinates is not the midpoint in another — the midpoint is not half in any sense that survives the change. On red and green the ΔE₀₀ repair undoes a miss of 8.2 colour differences and the model’s repair a miss of 17.8 units. They are repairs of different sizes, to different colours.

A gradient repaired in one unit carries a residual bow in the other, and neither census predicts its size, because it is a distance between two midpoints rather than between a midpoint and a mixture. That residual is the rank correlation in the form a designer actually meets it. A design system that specifies gradients by their stops has specified them in the unit its tool computed the stops in, whether or not the specification names that unit.

What a designer can take from it

Three consequences, from the most general to the most particular.

The size of the effect transfers and the list of offenders does not. A rule of thumb like a mixture misses its midpoint by about a seventh of its separation holds in both units. A rule like blue-to-yellow gradients need extra stops holds in one and not the other.

A gradient tool that advertises perceptual evenness has chosen a unit. A gradient is a path, and interpolation in CIELAB, Oklab or the appearance model’s space gives different midpoints. The bow here is the complementary fact: whichever unit a tool uses to judge the physical mixture, a different unit will find a different set of mixtures unacceptable.

And a comparison between units has to state its exponent. A model judged in another model’s unit found that swapping units reorders conclusions; the census in six units found the same for adaptation. Here a single comparison — model against ΔE₀₀ — reversed its headline depending on whether the model’s distance carried its power correction.

What was computed, and how

The mixtures are formed in tristimulus values by linear interpolation between two sRGB colours, which is what two lights at complementary drive levels produce. Each unit reads the half mixture and the two endpoints, and the bow is the distance between the half mixture’s reading and the midpoint of the two endpoints’ readings. The share divides that by the distance between the endpoints in the same unit.

The appearance model runs at an adapting luminance of 100 candelas a square metre, a twenty per cent background and an average surround unless a room is stated, with D65 as the white. Its unit here is the Euclidean distance in CAM16-UCS’s J′a′b′ unless the power correction is named.

The census draws 1,500 pairs uniformly in the sRGB cube with a fixed seed, keeps the 1,450 at least ten ΔE₀₀ apart, and reads each in both units. The rank correlation is Spearman’s, over the two shares.

Where the measurement stops

Two units only. ΔE*ab, Oklab and CIELUV would each give another ranking, and the weighting is the disagreement suggests the chroma weighting would sort them.

The appearance model is being used as a ruler for pairs viewed side by side under one condition. The bow of a physical mixture is a matching question — two lights, one room — and the model is the right tool for the room’s part of it and no better than any other unit for the matching part.

And the pairs are drawn uniformly in the sRGB cube, which over-represents saturated colours relative to the colours anybody blends. A census over the colours in real gradients would move both medians and would probably move the rank correlation.

The habit

The habit is about a result that survives a change of unit in the aggregate and not in the particular.

Two rulers that agree about the average size of an effect are easily taken to agree about the effect. They may agree about its size and disagree about where it is, and the second disagreement is the one a practitioner acts on, because a practitioner acts on particular cases.

The move is to compare rankings, not only summaries. A rank correlation costs one sort.

The failure mode is to validate a unit against another on a mean and then use it to choose between cases. A median can agree to two points while the list of worst cases is three quarters different.

Who noticed it first

That different colour spaces give different interpolation paths and different colour differences is general knowledge, and CAM16-UCS was introduced precisely because it predicts difference data better than earlier spaces. The power correction to its distance comes from later work fitting both small and large colour-difference data sets with one formula.

That a physical mixture’s bow ranks pairs differently in the two units, and that the correction rather than the space is what makes the model’s bows look smaller, follows from computing both on the same pairs. The sources consulted here compare units on difference data rather than on mixtures.

Still open: what an observer sees in a mixture

Both units are fitted to judgements of pairs. A physical mixture is a third patch that an observer compares with two others, and whether its perceived position between them follows ΔE₀₀, the model, or neither is an experiment rather than a computation — a bisection experiment on physical mixtures of display primaries, which is cheap to run and would decide which ranking describes what a viewer sees.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Additive mixtureCIECAM16CIEDE2000Colour differenceDisplay gamutGradientPerceptual uniformityPrimariesSurroundTristimulus