Difference and uniformity

Three constants nobody quotes

CIEDE2000 is defined with three parametric factors in it, the CIE leaves them to the industry using the formula, and the two settings in ordinary use differ by a factor of two in one of them. Twenty-nine per cent of acceptance decisions change between the two — a larger disagreement than any between the formulae themselves.

Assumes A tolerance is a shape and Which of two is worse.

Every colour difference on this site is written ΔE00, and the notation is incomplete. The formula as the CIE published it is

ΔE00=(ΔLkLSL)2+(ΔCkCSC)2+(ΔHkHSH)2+RTΔCkCSCΔHkHSH\Delta E_{00} = \sqrt{\left(\frac{\Delta L'}{k_L S_L}\right)^2 + \left(\frac{\Delta C'}{k_C S_C}\right)^2 + \left(\frac{\Delta H'}{k_H S_H}\right)^2 + R_T \frac{\Delta C'}{k_C S_C}\frac{\Delta H'}{k_H S_H}}

and three of the symbols in it are not defined by the CIE at all. They are left to the industry using the formula, the two settings in ordinary use differ by a factor of two in one of them, and a specification that says “ΔE00 ≤ 1.0” has not said which.

The three constants a specification does not quoteΔE2000 is defined with three parametric factors in it — kL, kC and kH — which the CIE leaves to the industry using the formula rather than fixing. Graphic arts uses ones throughout; the textile standard weighs lightness at half, which is kL = 2. Applied to 4000 pairs at a tolerance of 1, the two settings accept 46 and 75 per cent, and 29 per cent of pairs change verdict — a larger disagreement than any between the formulae themselves. The reference conditions the ones assume are diffuse illumination at 1000 lux, a mid-grey surround, samples abutting, subtending over four degrees, differing by under five units, with no visible texture.graphic, kL 145.6%textile, kL 274.5%verdict changes28.9%pairs accepted, at a tolerance of ΔE00 14000 pairsΔE2000 with kL 1 against kL 2
Fig. 1 Four thousand pairs at a tolerance of one unit, judged twice. The graphic-arts setting accepts 45.6 per cent of them and the textile setting 74.5 — and 28.9 per cent of individual pairs change verdict between the two. Nothing about the samples, the formula or the observer differs between the two judgements.

The claim

Three of the eight quantities in a colour difference are set by the trade rather than by the standard, and they decide more acceptance decisions than the choice of formula does.

  • Twenty-nine per cent of pairs change verdict between the two settings in common use — graphic arts at ones throughout, textiles with lightness weighted at half — at a tolerance of one unit.
  • The effect does not shrink at any tolerance. From half a unit to five, the flip rate runs 29.4, 28.9, 28.5, 27.4, 26.5 per cent. It is not a boundary artefact and it does not go away by loosening the specification.
  • And it is larger than the differences this site has already made essays out of. Two colour-difference formulae disagree about which of two pairs is worse in 13.4 per cent of comparisons. Two settings of the same formula disagree about acceptance in 29.

What the factors are for

They are not a fudge, and the essay would be unfair if it left that impression.

The formula was fitted against judgements made under a stated set of conditions, and the conditions are unusually specific: diffuse illumination at about a thousand lux, a mid-grey surround, the two samples touching each other, subtending more than four degrees, differing by less than five units, with no visible texture.

Almost nothing in industry is like that. A piece of cloth has texture, is looked at draped rather than flat, and is judged at arm’s length. A car’s paint is seen curved, outdoors, in sunlight. A print is seen on a substrate whose surround is whatever else is on the page.

Departing from the reference conditions changes how much each of the three axes matters — and a viewing condition is an argument rather than a footnote — so the factors are the standard’s own mechanism for saying so. kL = 2 is the textile trade’s statement that a lightness difference on cloth is worth half what a lightness difference on a flat patch is worth — which is a claim about cloth, is defensible, and has been in the relevant standards for decades.

The complaint is not that they exist. It is that they are the only part of the specification that is never written down.

The same measurement at five tolerances. A quarter to a third of pairs change verdict at every tolerance from half a unit to five, so this is not an artefact of pairs sitting on a boundary. The pairs at each tolerance are drawn so that their differences scale with it, which is the only way to ask the question: a fixed sample would be measuring how far apart the samples happened to be rather than what the factors do.
Fig. 2 The same measurement at five tolerances, with the pairs drawn so their differences scale with the tolerance being tested. Between a quarter and a third of pairs change verdict everywhere from half a unit to five, so this is a property of the weighting rather than of pairs sitting on a boundary.

What was computed, and how

Four thousand pairs, generated from a seeded sampler over CIELAB with chroma bounded at 40, and then pulled together onto the scale the tolerance lives on: each pair is a random point and three per cent of the way toward a second random point, so the differences cluster around the tolerance rather than being spread over the whole space. A tolerance decision is never made between two random colours.

That pulling scales with the tolerance, which is the one detail that makes the sweep mean anything. The first version used a fixed interpolation, and the sweep it produced was 6.6 per cent at half a unit, 28.9 at one and 0.7 at two — a shape that looks like a resonance and is entirely an artefact of where the sample happened to sit. Scaling the interpolation with the tolerance makes each row a fair test of the same question, and the answer turns out to be flat.

Each pair is then judged twice, with the same formula and the same coordinates, differing only in kL. A pair counts as flipped when the two verdicts differ.

And the direction is asserted as well as the size. The setting that halves the weight on lightness accepts more — 74.5 per cent against 45.6 — which it must, since dividing one of three positive terms by two can only lower the total. An implementation that got that backwards would pass every structural check and fail this one.

Why it is invisible

Three reasons, and they compound.

The defaults are silent. Every implementation of ΔE2000 in every colour library defaults to ones, so a specification that says nothing gets the graphic-arts setting whether or not anybody meant it. A default that is also the most common intention is a default nobody has to think about, which is exactly the condition under which the exceptions get missed.

The number looks the same. A textile tolerance of 1.0 and a graphic-arts tolerance of 1.0 are printed identically, stored identically, and compared identically — which is the same failure as a tolerance quoted without its measurement condition. There is no unit, no annotation and no field. Two databases of acceptance results, one from each trade, can be merged without any error being raised and without any of the merged rows meaning what they say.

And the disagreement is invisible in aggregate. Both settings accept a large fraction of pairs and reject the rest. Comparing summary statistics — pass rates, mean differences — shows two systems that look broadly similar. It is only pair by pair that a third of the verdicts turn out to be different, and nobody compares pair by pair because the two trades do not share samples.

A fourth reason is structural rather than cultural. The factors sit in the denominators, so setting one of them to two is arithmetically identical to loosening the tolerance along that one axis — and a laboratory that wanted a looser lightness tolerance could get the same result by writing a different number after the inequality. The two routes are not equivalent in intent: one says this material is judged differently, the other says this contract is more forgiving. They are exactly equivalent in effect, they produce the same accepted set, and only one of them appears in the specification. So a reader of a textile result cannot tell, from the number, whether they are looking at a tight tolerance under a generous weighting or a loose tolerance under a strict one.

What it costs, worked through

Take a supplier submitting to ΔE00 ≤ 1.0 with a process whose output sits, on the graphic-arts setting, at a mean of 0.8 with the ordinary amount of scatter. Under that setting they pass comfortably.

Now suppose the customer’s laboratory is a textile house and uses kL = 2, as its trade’s standard tells it to. Every submission’s lightness term is halved before the root is taken, so every sample’s reported difference falls — and the supplier’s rejected tail, the samples whose failure was mostly a lightness error, moves inside the limit. The pass rate goes up and nothing about the goods has changed.

The reverse is the case that costs money. A supplier whose process is well controlled in lightness and poor in chroma gains nothing from kL = 2, because the term being discounted was not the one causing the failures. Two suppliers with the same mean difference and different error shapes are therefore affected in opposite directions by the same change of convention — and the shape of a process’s error is precisely the thing a mean difference does not record.

That is the practical sting. The factors do not shift a threshold up or down for everybody. They rotate the acceptance region, and a rotation helps whoever is already aligned with it.

Where the model stops

A flip rate is not a cost. Twenty-nine per cent of pairs changing verdict does not mean twenty-nine per cent of a supplier’s output changes hands: real submissions are not uniformly distributed near the tolerance, and a competent supplier aims well inside it. What the rate measures is the size of the ambiguity in the decision rule, which is the thing a contract is supposed to remove.

The two settings are quoted, not derived. kL = 2 for textiles is a convention, and this site has no data with which to check that cloth really does halve the weight on lightness. If the true factor for some material were 1.4 — a value in use for coatings and standardised nowhere — the flip rate against ones would be smaller and would not be zero.

And kC and kH are almost never moved, so this essay has effectively measured one factor of the three. That is honest as far as it goes and it is a floor: a trade that moved two of them would disagree with the default by more than a trade that moves one.

The sampler is uniform in the space and no real submission is. Pairs are drawn over CIELAB with chroma bounded at 40, which populates regions where nothing is ever manufactured as heavily as regions where everything is. A flip rate measured on a real production distribution would be different, would depend on the product, and would be a number about that product rather than about the formula — which is why the uniform sampler is the right instrument for the question asked here and the wrong one for any question a factory has.

The generalisation

The sentence worth carrying is: a specification is a formula, its arguments, and the conditions its data were taken under, and only the first of those is ever written down.

This site keeps finding the same shape. A colour temperature does not specify a spectrum. A white bin does not specify a chromaticity. A hex code does not specify a colour. And a ΔE00 does not specify a difference — it specifies a difference given three constants, a set of viewing conditions and an observer, none of which travels with the number.

The pattern is that a quantity gets adopted for its convenience, its arguments get defaulted, and after a decade the defaults become invisible enough that people forget the arguments were there. The repair in every case is the same and is not expensive: quote the arguments beside the number.

The surprising connection is with the observer population. That essay found a spread of about fifteen per cent in what an ordinary pair’s difference is worth to different people, and treated it as a serious limitation on what a tolerance can mean. The factors move the same quantity by a factor of two, deliberately, and nobody records it. The unavoidable uncertainty is being measured to two decimals while the avoidable one is left blank.

Who found it, and when

The factors were in CIE94, published in 1995, and they arrived with the formula rather than after it. The 1994 recommendation explicitly names the reference conditions and explicitly says the factors exist so that departures from them can be accounted for.

CIEDE2000 kept them unchanged in 2001. The 2000 formula’s improvements — the hue rotation term, the a* rescaling, the revised weighting functions — were fitted with all three factors at one, on data taken under the reference conditions, and the factors were carried through untouched.

kL = 2 for textiles predates both, and comes from the CMC(l:c) formula of 1984, whose whole notation is a pair of parametric factors: CMC(2:1) is a lightness factor of two and a chroma factor of one, written in the name of the formula. That is the honest way to do it, and the convention did not survive the move to CIE94.

The loss of that notation is the whole story of this essay. CMC(2:1) carries its parameters in its name and cannot be quoted without them. ΔE00 carries them nowhere and is quoted without them universally. A change in notation, made for good reasons, removed the only place the information had ever been written.

What the pictures cannot show

They cannot show a pair being judged. The flip rate is arithmetic on a formula; whether a person would accept either sample is a decision made by a person under conditions the figure does not describe. The factors exist precisely because those conditions vary, so a figure drawn under one set of conditions is the wrong instrument for arguing about them.

And they cannot show cloth. The textile setting is a claim about texture, drape and viewing distance, none of which is representable on this page — every patch here is flat, matte, abutting and seen at whatever distance the reader is sitting. The essay is therefore in the odd position of measuring the consequences of a factor whose justification it cannot draw.

The three constants a specification does not quote. ΔE2000 is defined with three parametric factors in it — kL, kC and kH — which the CIE leaves to the industry using the formula rather than fixing. Graphic arts uses ones throughout; the textile standard weighs lightness at half, which is kL = 2. Applied to 4000 pairs at a tolerance of 2, the two settings accept 44 and 73 per cent, and 29 per cent of pairs change verdict — a larger disagreement than any between the formulae themselves. The reference conditions the ones assume are diffuse illumination at 1000 lux, a mid-grey surround, samples abutting, subtending over four degrees, differing by under five units, with no visible texture.
Fig. 3 The same measurement at a tolerance of two units, which is where a great deal of packaging and textile work is done. The acceptance rates move and the flip rate does not — a supplier working to a looser tolerance is exposed to exactly the same ambiguity as one working to a tight one.

A tight specification is the third of the three settings worth drawing, and it is where the constants have the largest say over the verdict.

The three constants a specification does not quote. ΔE2000 is defined with three parametric factors in it — kL, kC and kH — which the CIE leaves to the industry using the formula rather than fixing. Graphic arts uses ones throughout; the textile standard weighs lightness at half, which is kL = 2. Applied to 4000 pairs at a tolerance of 0.5, the two settings accept 46 and 76 per cent, and 29 per cent of pairs change verdict — a larger disagreement than any between the formulae themselves. The reference conditions the ones assume are diffuse illumination at 1000 lux, a mid-grey surround, samples abutting, subtending over four degrees, differing by under five units, with no visible texture.
Fig. 4 And at half a unit, which is where a tight specification lives. The three constants matter most exactly where the tolerance is tightest, which is the opposite of what a reader would guess and the reason they are worth quoting.
Which of two reproductions is better depends on which statistic is asked. A specification for a proof or a print run is written against a set and has to reduce the set to one number. Here are two candidates: the one with the lower mean has the higher ninety-fifth percentile, so the mean prefers A and the tail prefers B. Across 2000 pairs of candidates generated the same way, the two statistics disagree about the winner 30 per cent of the time. Both numbers are honest; only one of them is what somebody notices.
Fig. 5 And the same three constants seen from the other side: two candidate reproductions of one set, where which of them is better depends on which statistic the constants are fed into. A grader who never quotes them has still chosen them.

The flip rate is not a third measurement

Three numbers are quoted from the four thousand pairs — 45.6 per cent accepted under the graphic-arts setting, 74.5 under the textile one, and 28.9 per cent of pairs changing verdict — and they read as three findings. 74.5 minus 45.6 is 28.9 exactly, and that is not a coincidence to be noted in passing: it is forced.

Halving kL divides one of three positive terms under the root, so every pair’s reported difference falls or stays the same. The textile setting therefore accepts every pair the graphic-arts setting accepts, and some more. The two accepted sets are nested, so the count of pairs that differ is the difference of the two counts, always, with no room for arithmetic to say anything else.

Two consequences follow, and the second matters more than the first.

The flip rate carries no information beyond the acceptance rates. Twenty-nine per cent of decisions change and the pass rate rises by twenty-nine points are the same sentence. Quoted together they suggest a decision rule that is unstable in both directions; what they describe is a threshold that moves one way.

And no pair can go from accepted to rejected. That is a genuine correction to the worked example above, which says two suppliers with the same mean difference and different error shapes are affected in opposite directions by the change of convention. They are not. The supplier whose errors are mostly in lightness gains a great deal and the supplier whose errors are mostly in chroma gains almost nothing, and neither loses anything at all. Nobody’s goods are rejected by kL = 2 that would have passed under ones.

The rotation language is right about the shape of the acceptance region and wrong about its consequence. The region does stretch along one axis rather than growing uniformly, so the benefit is unevenly distributed — but a stretch is not a rotation, and a strictly larger region cannot harm anybody. The ambiguity here is one-sided, and that makes it a different kind of contract problem: not a coin toss about who wins, but a discount whose size depends on which way a supplier’s process happens to fail.

Three units is the other tolerance in common use, and running the same four thousand pairs there says whether the flip rate is a feature of one particular threshold or of the whole range.

The three constants a specification does not quote. ΔE2000 is defined with three parametric factors in it — kL, kC and kH — which the CIE leaves to the industry using the formula rather than fixing. Graphic arts uses ones throughout; the textile standard weighs lightness at half, which is kL = 2. Applied to 4000 pairs at a tolerance of 3, the two settings accept 43 and 71 per cent, and 27 per cent of pairs change verdict — a larger disagreement than any between the formulae themselves. The reference conditions the ones assume are diffuse illumination at 1000 lux, a mid-grey surround, samples abutting, subtending over four degrees, differing by under five units, with no visible texture.
Fig. 6 The same measurement at a tolerance of three units. The graphic-arts setting accepts 43 per cent and the textile one 71, and 27 per cent of pairs change verdict — within a point of what the tolerance of two gave, so the disagreement is not a property of any one threshold.

The number a supplier would want

Stated as a share of all pairs, 28.9 per cent understates what is at stake for the party the specification exists to constrain. A supplier does not care what fraction of the whole space changes verdict; they care what fraction of their rejections would survive under the other convention.

Under the graphic-arts setting 54.4 per cent of the four thousand pairs are rejected. Of those, 53.1 per cent are accepted under the textile setting. More than half of every rejection is overturned by a constant that appears nowhere in the specification.

Read from the other end it is nearly as stark: of everything the textile setting accepts, 38.8 per cent would be rejected by a laboratory using the defaults. A shipment passing one house’s inspection has a better than one in three chance of failing the other’s, with the same instrument, the same formula and the same tolerance printed on the same page.

Those two figures are the same measurement as the 28.9 and they are the ones a contract argument would be conducted in.

At five units the tolerance is loose enough that most working specifications would call it slack, and it is worth knowing whether the disagreement finally goes away there.

The three constants a specification does not quote. ΔE2000 is defined with three parametric factors in it — kL, kC and kH — which the CIE leaves to the industry using the formula rather than fixing. Graphic arts uses ones throughout; the textile standard weighs lightness at half, which is kL = 2. Applied to 4000 pairs at a tolerance of 5, the two settings accept 42 and 68 per cent, and 26 per cent of pairs change verdict — a larger disagreement than any between the formulae themselves. The reference conditions the ones assume are diffuse illumination at 1000 lux, a mid-grey surround, samples abutting, subtending over four degrees, differing by under five units, with no visible texture.
Fig. 7 Five units, with 42 and 68 per cent accepted and 26 per cent of pairs flipping. The flip rate falls as the tolerance loosens, which is the direction the mechanism predicts and is still a quarter of the pairs at the loosest setting anybody uses.

Two disagreements that are not comparable

The essay’s closing comparison — 28.9 per cent against the 13.4 per cent by which two formulae disagree, “more than twice as much” — is arithmetically right and compares two quantities with different denominators and different kinds.

The 13.4 is a rate of rank swaps: how often two formulae order one pair of pairs oppositely, over comparisons of two pairs. The 28.9 is a rate of verdict changes: how often one pair falls on different sides of a threshold, over pairs. One is a statement about a ranking and needs four colours; the other is a statement about a threshold and needs two.

They also differ in symmetry, which is the more important difference. A rank swap can go either way — sometimes one formula puts a pair higher, sometimes the other — so the 13.4 per cent is a genuine two-sided ambiguity. The 28.9 per cent, as the section above establishes, is entirely one-sided.

None of that rescues the defaults, and the essay’s point survives intact: an unstated constant is moving more decisions than the argument about which formula to use. But twice as much is a comparison of a rate against a different rate, and the honest version is that the two disagreements are large in different ways and neither is a scaled copy of the other.

The other end of the range is the one that matters to anybody holding a tight specification, and a quarter of a unit is about as tight as a real process is asked to be.

The three constants a specification does not quote. ΔE2000 is defined with three parametric factors in it — kL, kC and kH — which the CIE leaves to the industry using the formula rather than fixing. Graphic arts uses ones throughout; the textile standard weighs lightness at half, which is kL = 2. Applied to 4000 pairs at a tolerance of 0.25, the two settings accept 46 and 76 per cent, and 30 per cent of pairs change verdict — a larger disagreement than any between the formulae themselves. The reference conditions the ones assume are diffuse illumination at 1000 lux, a mid-grey surround, samples abutting, subtending over four degrees, differing by under five units, with no visible texture.
Fig. 8 A quarter of a unit, where the two settings accept 46 and 76 per cent and 30 per cent of pairs change verdict. That is the highest flip rate of the five thresholds drawn here, so the constants nobody quotes matter most exactly where the tolerance is tightest.

The sweep is flatter than it looks

One last reading of the five-tolerance sweep. The flip rates run 29.4, 28.9, 28.5, 27.4 and 26.5 per cent from half a unit to five — described above as flat, and the description is fair: the total fall is 2.9 points, or ten per cent of the value, across a tenfold change in the tolerance.

That flatness is the essay’s strongest structural evidence and it is worth saying why. If the effect were pairs sitting near a boundary, the flip rate would fall as the boundary moved into a sparser part of the sample, and it would fall fast. Ten per cent of decay across a factor of ten in tolerance is very nearly no decay at all, and it says the disagreement is a property of the weighting rather than of where the threshold happens to sit — which is exactly what the scaled sampler was built to be able to test.

Where the ladder goes next

The nearest unfinished piece is the automotive setting. A value near 1.4 is in use for coatings, is not standardised anywhere this site can find, and sits between the two measured here — so the flip rate against both would be smaller than 29 per cent and the three-way disagreement would be worse than any pairwise one. Three trades, three settings, one notation.

The second is the interaction with the observer. The factors are a correction for viewing conditions and the population is a spread over eyes, and the two act on the same number without either knowing about the other. Whether a pair that is fragile to one is fragile to the other is computable from the machinery this phase built and has not been asked.

And the third is the one that would settle the whole argument: whether the factors are the right shape. Weighting an axis by a constant assumes that departing from the reference conditions scales that axis uniformly, everywhere in the space. Nothing in the standard argues for that, and the S terms next door are position-dependent for exactly the reason a constant would be suspicious.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AcceptabilityChromaCIEDE2000ΔELightnessMeasurement errorPerceptual uniformityQuality controlSpecificationStandard observerToleranceViewing condition