What it takes to deliver it

A tolerance is a boundary through pairs

Twenty-four pairs of surfaces built to sit exactly on a ΔE2000 tolerance of one read from 0.69 to 1.93 in the other five units. A contract quoting "one unit" without naming the formula does not become slightly wrong in another; it accepts and rejects a different set of deliveries, and in one case rejects every single one.

Assumes A tolerance is a region, A choice with no magnitude and A delivery tolerance is three tolerances.

A tolerance is where a colour-difference formula stops describing and starts deciding. A delivery either passes or it does not, and the line between the two is drawn by whichever formula the contract happens to name.

Pairs built to sit exactly on a ΔE2000 tolerance, read in every other unit. Twenty-four pairs of surfaces, each constructed by walking one member along a fixed direction until the difference is exactly 1.0 ΔE2000 under D65. The bar spans what those same pairs read in each unit, after calibration, with the tick at the mean. ΔE2000's own row is a point at 1.0 by construction. Every other unit spreads them: ΔE*uv reads them from 0.69 to 1.75, so a contract written at "one unit" accepts and rejects a different set of deliveries depending on which unit it means. CAM16-UCS rejects all twenty-four: it reads the closest of them at 1.43.
Fig. 1 Twenty-four pairs of surfaces, each constructed by walking one member along a fixed direction until the difference under D65 is exactly 1.0 ΔE2000, then read in every other unit after calibration. ΔE2000’s own row is a point at 1.0 by construction. The tick on each bar is the mean.

The claim

A tolerance is a boundary drawn through the space of pairs, and the six formulae draw six different boundaries enclosing different sets.

  • The pairs are built to sit exactly on the line. Each is walked until the difference is 1.0 ΔE2000 to twelve figures, so any spread afterwards is entirely the formula’s doing.
  • They spread from 0.69 to 1.93 across the five alternatives, after each has been calibrated onto ΔE2000’s scale.
  • CIELUV is the widest at 0.69 to 1.75 and accepts 54 per cent of them; ΔE*94 is the narrowest at 0.91 to 1.42 and accepts 33.
  • CAM16-UCS accepts none of them. Its closest reading of a pair on the ΔE2000 boundary is 1.43.
  • Nothing here is a rescaling. Every unit is calibrated first, so the spread is disagreement about which pairs are close, not about how big a unit is.

What a tolerance actually is

A specification says ΔE ≤ 1.0, and what it means is easy to state wrongly.

It does not name a distance in a physical sense. It names a region: for every reference colour, the set of colours a delivery may take. This collection has already established that the region is not a sphere — the formula’s weighting functions make it an ellipsoid, tilted, and its size and shape depend on where the reference colour is.

So a tolerance is a boundary drawn through the space of pairs. Every pair is either inside or outside. The number 1.0 is a label on the boundary and not a measurement of anything, and the boundary’s shape is entirely the formula’s.

That reframing is what makes the question here answerable. Instead of asking how much a number changes when the formula changes — which is bookkeeping, and is removed by calibration — ask which pairs change sides.

How the pairs were built

Twenty-four surfaces from the collection’s own test set, each with a perturbation direction: a small smooth wobble applied to the reflectance, the same shape for every surface, so that walking along it changes the colour without changing what kind of object it is.

For each, the walk is bisected until the difference under D65 lands on exactly 1.0 ΔE2000 — forty bisections, so the residual is below floating-point resolution. That gives twenty-four pairs that are, by construction, precisely on the boundary in the published unit.

Then read them in the other five. Since every unit has already been multiplied by the single factor that best carries it onto ΔE2000 over a common sample, a unit that agreed about anything but scale would read all twenty-four at 1.0.

unit lowest reading mean highest of the 24, how many pass
ΔE*94 0.91 1.07 1.42 33%
Oklab 0.87 1.10 1.48 38%
ΔE*ab 0.77 1.08 1.63 58%
ΔE*uv 0.69 1.08 1.75 54%
CAM16-UCS 1.43 1.58 1.93 0%

The two ways a boundary can differ

Reading down the last two columns shows that the boundaries differ in two independent ways, and separating them matters for what a specification writer should do.

The four matching units make that separation cleanly, and the table shows it more sharply than the prose does. Their means are 1.07, 1.10, 1.08 and 1.08 — a spread of 2.8 per cent across four formulae that were calibrated independently. Their ranges are 0.51, 0.61, 0.86 and 1.06 — a factor of 2.08.

So the four agree about size to within three per cent and disagree about shape by a factor of two. There is no size disagreement among them worth correcting; the whole of what separates them is which pairs they put where.

And that produces an inversion in the acceptance column worth stopping on. Since every pair sits exactly on ΔE2000’s boundary, ΔE2000 accepts all twenty-four, and each alternative’s acceptance rate is its agreement rate. Ordering the four by how tightly their readings cluster:

unit spread of readings agrees with ΔE2000 about
ΔE*94 0.51 33%
Oklab 0.61 38%
ΔE*ab 0.86 58%
ΔE*uv 1.06 54%

The formula whose readings agree best with ΔE2000 agrees worst about the decision, and the rank correlation between spread and agreement is +0.80.

The mechanism is arithmetic rather than perceptual and is worth having explicitly, because it makes the acceptance column easy to misread. All four sit with a mean a little above the line — 1.07 to 1.10, so the typical reading of a pair on the boundary is just outside it. A formula whose readings cluster tightly around that mean puts nearly everything just outside and rejects nearly everything. A formula whose readings scatter widely around the same mean puts half of them below the line by accident.

So the acceptance rate is measuring leniency and reading as agreement. ΔE*uv’s 54 per cent is not evidence that CIELUV is a closer match to ΔE2000’s boundary than ΔE*94’s 33; it is evidence that CIELUV’s boundary is a different shape and that the difference happens to fall on both sides of the line rather than consistently on one. The two formulae are wrong about these pairs in the same average amount and in different patterns, and the pattern that scatters scores better on a yes-or-no test.

That has a practical edge for anybody porting a threshold. A high agreement rate between two formulae on marginal pairs is not a reason to trust the port, because it can be produced by a formula that disagrees more. The quantity to look at is the spread of readings, and on that measure the ordering is the reverse: ΔE*94 is the formula whose boundary is nearest ΔE2000’s shape, and it is the one the acceptance column ranks last.

A boundary can be the wrong size. CAM16-UCS’s readings are all above 1.0 by between 43 and 93 per cent, so its boundary at 1.0 is strictly tighter than ΔE2000’s at 1.0 and admits none of the same marginal pairs. That is a calibration disagreement that survived calibration — and it survives because the calibration was fitted over pairs spanning a fraction of a unit to about ten, while these pairs are all at exactly one, which is the band where CAM16-UCS departs from the published unit most.

The repair for a wrong size is easy: change the number. A contract wanting CAM16-UCS to accept what ΔE2000 accepts should specify 1.58, not 1.0, and the two would then agree about roughly half the marginal pairs.

A boundary can be the wrong shape. CIELUV reads the same twenty-four pairs from 0.69 to 1.75 — a factor of 2.5 between the pair it thinks closest and the pair it thinks furthest, all of which are identical in the published unit. No change of threshold fixes that. Whatever number is chosen, some pairs ΔE2000 accepts will be rejected and some it rejects will be accepted.

The two are separable and the second is the one that costs money. A shape disagreement means two laboratories can both be measuring correctly, both be using a defensible formula, and disagree about individual batches indefinitely.

The three constants a specification does not quoteΔE2000 is defined with three parametric factors in it — kL, kC and kH — which the CIE leaves to the industry using the formula rather than fixing. Graphic arts uses ones throughout; the textile standard weighs lightness at half, which is kL = 2. Applied to 4000 pairs at a tolerance of 1, the two settings accept 46 and 75 per cent, and 29 per cent of pairs change verdict — a larger disagreement than any between the formulae themselves. The reference conditions the ones assume are diffuse illumination at 1000 lux, a mid-grey surround, samples abutting, subtending over four degrees, differing by under five units, with no visible texture.graphic, kL 145.6%textile, kL 274.5%verdict changes28.9%pairs accepted, at a tolerance of ΔE00 14000 pairsΔE2000 with kL 1 against kL 2
Fig. 2 A tolerance read as a probability rather than a boundary, which is this collection’s earlier treatment of the same object. The spread of readings above is what turns a sharp boundary into a probabilistic one when the formula is not agreed.
Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu.
Fig. 3 Where on the scale the units disagree with the published one. A tolerance at 1.0 sits in the leftmost band, which is where the proportional disagreement is largest — 16 to 91 per cent depending on the unit.

Which pairs move, and why

The pairs the alternatives read highest are the saturated ones, and this is the chroma weighting arriving in a place where it decides a commercial outcome.

A pair at ΔE2000 = 1.0 between two saturated surfaces has, in raw CIELAB terms, a considerably larger difference than a pair at 1.0 between two pale ones — because CIEDE2000 divided the saturated pair’s chroma difference by 1 + 0.045 C on the way. Read the saturated pair in ΔE*ab, which does no dividing, and it comes back high. Read the pale pair and it comes back near 1.0.

So the disagreement across the twenty-four pairs is almost entirely ordered by the chroma of the reference surface. The pair CIELUV reads at 1.75 is one of the most saturated in the set; the pair it reads at 0.69 is one of the palest.

That has a practical shape. A specification written in ΔE*ab and tested in ΔE2000 will be systematically lenient on saturated colours and strict on pale ones relative to the specification’s intent, or the reverse depending on the direction of substitution. A brand colour is usually saturated; a substrate or a neutral is usually pale. The two ends of a print job are on opposite sides of this effect, which is one reason a delivery tolerance has to be more than one number.

The appearance unit rejects everything, and that is not a bug

Zero of twenty-four is a striking number and deserves care, because it invites a wrong reading.

It does not mean CAM16-UCS thinks these pairs are obviously different. It means its boundary at the label 1.0 is tighter than ΔE2000’s at the label 1.0, after a calibration fitted over a wide range of separations. The label is arbitrary in both.

What it does mean is that the two units cannot share a threshold, and a specification that switches formula without re-deriving the number will change its acceptance rate to nearly zero or nearly one. Given that appearance-based tolerancing is proposed periodically and that CAM16-UCS is the natural unit for it, the number is worth having on record.

The deeper reason is the exponent. CAM16-UCS raises its distance to the power 0.63, which stretches the near end of the scale relative to the far end, and a global scale factor cannot undo a power law. Any unit with a non-unit exponent will fail to share a threshold with any unit without one, for arbitrarily good reasons on both sides.

What an overlap of a half actually costs

The acceptance percentages are the practical output, and they need translating into what a supplier and a customer would experience.

Suppose a run of deliveries whose errors are distributed around the tolerance — which is the normal state of a controlled process, since a process comfortably inside the tolerance is over-engineered and one comfortably outside is not shipping. Then the marginal batches are the ones near the boundary, and the acceptance figures above say what fraction of those two formulae agree about.

ΔE*ab and ΔE2000 agree about 58 per cent of the marginal batches — the best of the four, for the reason above, which is not that its boundary is the closest in shape. So on a run where half the batches are marginal, about a fifth of all batches would be decided differently by the two formulae. That is not a rounding error in a commercial relationship; it is a supplier and a customer disagreeing about one job in five, each with a defensible measurement.

And the disagreement is not random. It is ordered by chroma, as above, so it presents as a pattern: the saturated colours fail more often under one formula and the neutrals under the other. A pattern is worse than noise here, because it looks like a process problem and invites a process change that will not fix it. The same shape of misattribution turns up whenever an instrument’s convention is mistaken for the thing being measured.

The one comfort is that the disagreement is stable and computable. Two parties with the same set of pairs can derive the number that makes their two formulae agree about their own product, which is a half-hour of arithmetic and is what the third rule below asks for.

What follows for a specification

Three rules, in increasing order of how often they are broken.

Name the formula. ΔE ≤ 1.0 names no formula and there are at least six. The practice in printing is to write ΔE00 ≤ 1.0, which is right; the practice in a great deal of else is not.

Name the parametric factors too. CIEDE2000 and ΔE*94 both take kL, kC, kH, and the textile industry’s values differ from the graphic arts’. This collection has measured what those decide and it is not small.

And do not port a threshold between formulae by rescaling it. The mean readings above are 1.07, 1.08, 1.08, 1.10 and 1.58, so a naive port using the mean would be nearly right for four of the five and would still get between a third and a half of the marginal pairs wrong, because the mean is not the boundary. Porting a threshold properly means choosing the number that reproduces the acceptance rate on a representative set of pairs, which requires the set — and a specification that does not name its set is making the mistake this collection spends a round on.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 4 The six units by how far each is from a rescaling of the published one. A tolerance can only be ported between two units whose scatter is small, and the smallest here is 15 per cent — which is why the third rule above exists.
What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 5 Six published quantities across the menu. The tolerance question is the same question with a boundary attached, which is what turns a spread of readings into a change of decision.

A single one of the measured ellipses is where the six units stop agreeing even about which of them is roundest, and it is the finest grain this whole comparison has to offer.

One of MacAdam's ellipses as each unit sees it, at x 0.280, y 0.385. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 1.46; the least round is ΔE*ab at 2.24. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 6 One of MacAdam’s ellipses drawn as its perimeter distance in each of the six units, with each outline scaled to its own mean radius so the six compare as shapes. A unit in which one step meant one thing in every direction would have drawn a circle.

The number is not the only thing a contract has to name

A tolerance in a real specification carries more than a formula, and it is worth listing what else is in the same position — chosen once, rarely stated, and worth as much as this.

The illuminant. A pair at one unit under D50 is not at one unit under a shop’s fluorescent tubes, and the collection has measured the spread: a tolerance is not invariant under a change of light and adaptation does not make it one.

The observer. Two standard observers are about 2.65 ΔE2000 apart over a set of ordinary surfaces, which is well over twice a typical tolerance.

The measurement geometry. An instrument reporting 45°/0° and one reporting d/8° disagree about a glossy sample by more than the tolerance, which is a difference in what was measured rather than in how it was scored.

And the set the tolerance is asserted over. A specification saying every patch must be inside one unit and one saying the mean must be are different specifications, and a mean is not a worst case by a factor of about two.

Five choices, each worth as much as the formula, each usually unstated. The formula is the one this essay measures because it is the one an audit of units can reach; the honest summary is that it is one of five and not the largest.

Where the model stops

Twenty-four pairs is a small set and they are constructed. Their perturbation direction is one shape applied to every surface, so the pairs differ in colour but not in the kind of difference between them — and a real batch differs from its standard in whatever way the process happened to drift, which is often a lightness shift or an ink-density change rather than a smooth spectral wobble. A set of realistic drift directions would give a wider spread, not a narrower one.

They are also all at exactly one unit. A specification at 2.0 or at 0.5 would sit in a different band of the curve where the formulae part, and the proportional spread there is different — smaller at 2.0, larger at 0.5.

And the whole exercise holds the illuminant and the observer fixed. A tolerance measured under one light is not the same tolerance under another, and that effect is comparable in size to this one. The two compound, and neither is usually in the contract.

The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.
Fig. 7 The adaptation census in six units, for scale. A census row is a level and moves by a factor of two; a tolerance is a boundary and moves the set of things it admits, which is the harder change to correct for.

Who found it, and when

That different formulae accept different batches is the oldest practical complaint in industrial colour control, and it is why every specification worth anything names its formula. The CIE’s own technical reports on tolerancing say so at length.

What is less often done is to hold the pairs fixed and vary the formula, which isolates the effect from everything else that varies between two laboratories. The usual experiment varies both — two labs, two instruments, two formulae, two sets of samples — and reports an overall disagreement that cannot be attributed. Constructing pairs to sit exactly on one boundary and then reading them elsewhere costs nothing and attributes the whole of the result, and it is available to anybody with the formulae in one program.

Where the ladder goes next

A camera profile is fitted by a linear least-squares solve in tristimulus space, which is an objective, and it is on nobody’s menu. Refitting the same 3×3 to minimise each of the six units instead turns a linear solve into a nine-parameter search — and moves the matrix, which means the camera renders different pixels rather than merely reporting a different score.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 9 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AcceptanceAppearance modelCalibrationChromaCIEDE2000Colour differenceDeliveryJust-noticeable differenceSpecificationTolerance