A tolerance is a boundary through pairs
Assumes A tolerance is a region, A choice with no magnitude and A delivery tolerance is three tolerances.
A tolerance is where a colour-difference formula stops describing and starts deciding. A delivery either passes or it does not, and the line between the two is drawn by whichever formula the contract happens to name.
The claim
A tolerance is a boundary drawn through the space of pairs, and the six formulae draw six different boundaries enclosing different sets.
- The pairs are built to sit exactly on the line. Each is walked until the difference is 1.0 ΔE2000 to twelve figures, so any spread afterwards is entirely the formula’s doing.
- They spread from 0.69 to 1.93 across the five alternatives, after each has been calibrated onto ΔE2000’s scale.
- CIELUV is the widest at 0.69 to 1.75 and accepts 54 per cent of them; ΔE*94 is the narrowest at 0.91 to 1.42 and accepts 33.
- CAM16-UCS accepts none of them. Its closest reading of a pair on the ΔE2000 boundary is 1.43.
- Nothing here is a rescaling. Every unit is calibrated first, so the spread is disagreement about which pairs are close, not about how big a unit is.
What a tolerance actually is
A specification says ΔE ≤ 1.0, and what it means is easy to state wrongly.
It does not name a distance in a physical sense. It names a region: for every reference colour, the set of colours a delivery may take. This collection has already established that the region is not a sphere — the formula’s weighting functions make it an ellipsoid, tilted, and its size and shape depend on where the reference colour is.
So a tolerance is a boundary drawn through the space of pairs. Every pair is either inside or outside. The number 1.0 is a label on the boundary and not a measurement of anything, and the boundary’s shape is entirely the formula’s.
That reframing is what makes the question here answerable. Instead of asking how much a number changes when the formula changes — which is bookkeeping, and is removed by calibration — ask which pairs change sides.
How the pairs were built
Twenty-four surfaces from the collection’s own test set, each with a perturbation direction: a small smooth wobble applied to the reflectance, the same shape for every surface, so that walking along it changes the colour without changing what kind of object it is.
For each, the walk is bisected until the difference under D65 lands on exactly 1.0 ΔE2000 — forty bisections, so the residual is below floating-point resolution. That gives twenty-four pairs that are, by construction, precisely on the boundary in the published unit.
Then read them in the other five. Since every unit has already been multiplied by the single factor that best carries it onto ΔE2000 over a common sample, a unit that agreed about anything but scale would read all twenty-four at 1.0.
| unit | lowest reading | mean | highest | of the 24, how many pass |
|---|---|---|---|---|
| ΔE*94 | 0.91 | 1.07 | 1.42 | 33% |
| Oklab | 0.87 | 1.10 | 1.48 | 38% |
| ΔE*ab | 0.77 | 1.08 | 1.63 | 58% |
| ΔE*uv | 0.69 | 1.08 | 1.75 | 54% |
| CAM16-UCS | 1.43 | 1.58 | 1.93 | 0% |
The two ways a boundary can differ
Reading down the last two columns shows that the boundaries differ in two independent ways, and separating them matters for what a specification writer should do.
The four matching units make that separation cleanly, and the table shows it more sharply than the prose does. Their means are 1.07, 1.10, 1.08 and 1.08 — a spread of 2.8 per cent across four formulae that were calibrated independently. Their ranges are 0.51, 0.61, 0.86 and 1.06 — a factor of 2.08.
So the four agree about size to within three per cent and disagree about shape by a factor of two. There is no size disagreement among them worth correcting; the whole of what separates them is which pairs they put where.
And that produces an inversion in the acceptance column worth stopping on. Since every pair sits exactly on ΔE2000’s boundary, ΔE2000 accepts all twenty-four, and each alternative’s acceptance rate is its agreement rate. Ordering the four by how tightly their readings cluster:
| unit | spread of readings | agrees with ΔE2000 about |
|---|---|---|
| ΔE*94 | 0.51 | 33% |
| Oklab | 0.61 | 38% |
| ΔE*ab | 0.86 | 58% |
| ΔE*uv | 1.06 | 54% |
The formula whose readings agree best with ΔE2000 agrees worst about the decision, and the rank correlation between spread and agreement is +0.80.
The mechanism is arithmetic rather than perceptual and is worth having explicitly, because it makes the acceptance column easy to misread. All four sit with a mean a little above the line — 1.07 to 1.10, so the typical reading of a pair on the boundary is just outside it. A formula whose readings cluster tightly around that mean puts nearly everything just outside and rejects nearly everything. A formula whose readings scatter widely around the same mean puts half of them below the line by accident.
So the acceptance rate is measuring leniency and reading as agreement. ΔE*uv’s 54 per cent is not evidence that CIELUV is a closer match to ΔE2000’s boundary than ΔE*94’s 33; it is evidence that CIELUV’s boundary is a different shape and that the difference happens to fall on both sides of the line rather than consistently on one. The two formulae are wrong about these pairs in the same average amount and in different patterns, and the pattern that scatters scores better on a yes-or-no test.
That has a practical edge for anybody porting a threshold. A high agreement rate between two formulae on marginal pairs is not a reason to trust the port, because it can be produced by a formula that disagrees more. The quantity to look at is the spread of readings, and on that measure the ordering is the reverse: ΔE*94 is the formula whose boundary is nearest ΔE2000’s shape, and it is the one the acceptance column ranks last.
A boundary can be the wrong size. CAM16-UCS’s readings are all above 1.0 by between 43 and 93 per cent, so its boundary at 1.0 is strictly tighter than ΔE2000’s at 1.0 and admits none of the same marginal pairs. That is a calibration disagreement that survived calibration — and it survives because the calibration was fitted over pairs spanning a fraction of a unit to about ten, while these pairs are all at exactly one, which is the band where CAM16-UCS departs from the published unit most.
The repair for a wrong size is easy: change the number. A contract wanting CAM16-UCS to accept what ΔE2000 accepts should specify 1.58, not 1.0, and the two would then agree about roughly half the marginal pairs.
A boundary can be the wrong shape. CIELUV reads the same twenty-four pairs from 0.69 to 1.75 — a factor of 2.5 between the pair it thinks closest and the pair it thinks furthest, all of which are identical in the published unit. No change of threshold fixes that. Whatever number is chosen, some pairs ΔE2000 accepts will be rejected and some it rejects will be accepted.
The two are separable and the second is the one that costs money. A shape disagreement means two laboratories can both be measuring correctly, both be using a defensible formula, and disagree about individual batches indefinitely.
Which pairs move, and why
The pairs the alternatives read highest are the saturated ones, and this is the chroma weighting arriving in a place where it decides a commercial outcome.
A pair at ΔE2000 = 1.0 between two saturated surfaces has, in raw CIELAB terms, a considerably larger difference than a pair at 1.0 between two pale ones — because CIEDE2000 divided the saturated pair’s chroma difference by 1 + 0.045 C on the way. Read the saturated pair in ΔE*ab, which does no dividing, and it comes back high. Read the pale pair and it comes back near 1.0.
So the disagreement across the twenty-four pairs is almost entirely ordered by the chroma of the reference surface. The pair CIELUV reads at 1.75 is one of the most saturated in the set; the pair it reads at 0.69 is one of the palest.
That has a practical shape. A specification written in ΔE*ab and tested in ΔE2000 will be systematically lenient on saturated colours and strict on pale ones relative to the specification’s intent, or the reverse depending on the direction of substitution. A brand colour is usually saturated; a substrate or a neutral is usually pale. The two ends of a print job are on opposite sides of this effect, which is one reason a delivery tolerance has to be more than one number.
The appearance unit rejects everything, and that is not a bug
Zero of twenty-four is a striking number and deserves care, because it invites a wrong reading.
It does not mean CAM16-UCS thinks these pairs are obviously different. It means its boundary at the label 1.0 is tighter than ΔE2000’s at the label 1.0, after a calibration fitted over a wide range of separations. The label is arbitrary in both.
What it does mean is that the two units cannot share a threshold, and a specification that switches formula without re-deriving the number will change its acceptance rate to nearly zero or nearly one. Given that appearance-based tolerancing is proposed periodically and that CAM16-UCS is the natural unit for it, the number is worth having on record.
The deeper reason is the exponent. CAM16-UCS raises its distance to the power 0.63, which stretches the near end of the scale relative to the far end, and a global scale factor cannot undo a power law. Any unit with a non-unit exponent will fail to share a threshold with any unit without one, for arbitrarily good reasons on both sides.
What an overlap of a half actually costs
The acceptance percentages are the practical output, and they need translating into what a supplier and a customer would experience.
Suppose a run of deliveries whose errors are distributed around the tolerance — which is the normal state of a controlled process, since a process comfortably inside the tolerance is over-engineered and one comfortably outside is not shipping. Then the marginal batches are the ones near the boundary, and the acceptance figures above say what fraction of those two formulae agree about.
ΔE*ab and ΔE2000 agree about 58 per cent of the marginal batches — the best of the four, for the reason above, which is not that its boundary is the closest in shape. So on a run where half the batches are marginal, about a fifth of all batches would be decided differently by the two formulae. That is not a rounding error in a commercial relationship; it is a supplier and a customer disagreeing about one job in five, each with a defensible measurement.
And the disagreement is not random. It is ordered by chroma, as above, so it presents as a pattern: the saturated colours fail more often under one formula and the neutrals under the other. A pattern is worse than noise here, because it looks like a process problem and invites a process change that will not fix it. The same shape of misattribution turns up whenever an instrument’s convention is mistaken for the thing being measured.
The one comfort is that the disagreement is stable and computable. Two parties with the same set of pairs can derive the number that makes their two formulae agree about their own product, which is a half-hour of arithmetic and is what the third rule below asks for.
What follows for a specification
Three rules, in increasing order of how often they are broken.
Name the formula. ΔE ≤ 1.0 names no formula and there are at least six. The practice in printing is to write ΔE00 ≤ 1.0, which is right; the practice in a great deal of else is not.
Name the parametric factors too. CIEDE2000 and ΔE*94 both take kL, kC, kH, and the textile industry’s values differ from the graphic arts’. This collection has measured what those decide and it is not small.
And do not port a threshold between formulae by rescaling it. The mean readings above are 1.07, 1.08, 1.08, 1.10 and 1.58, so a naive port using the mean would be nearly right for four of the five and would still get between a third and a half of the marginal pairs wrong, because the mean is not the boundary. Porting a threshold properly means choosing the number that reproduces the acceptance rate on a representative set of pairs, which requires the set — and a specification that does not name its set is making the mistake this collection spends a round on.
A single one of the measured ellipses is where the six units stop agreeing even about which of them is roundest, and it is the finest grain this whole comparison has to offer.
The number is not the only thing a contract has to name
A tolerance in a real specification carries more than a formula, and it is worth listing what else is in the same position — chosen once, rarely stated, and worth as much as this.
The illuminant. A pair at one unit under D50 is not at one unit under a shop’s fluorescent tubes, and the collection has measured the spread: a tolerance is not invariant under a change of light and adaptation does not make it one.
The observer. Two standard observers are about 2.65 ΔE2000 apart over a set of ordinary surfaces, which is well over twice a typical tolerance.
The measurement geometry. An instrument reporting 45°/0° and one reporting d/8° disagree about a glossy sample by more than the tolerance, which is a difference in what was measured rather than in how it was scored.
And the set the tolerance is asserted over. A specification saying every patch must be inside one unit and one saying the mean must be are different specifications, and a mean is not a worst case by a factor of about two.
Five choices, each worth as much as the formula, each usually unstated. The formula is the one this essay measures because it is the one an audit of units can reach; the honest summary is that it is one of five and not the largest.
Where the model stops
Twenty-four pairs is a small set and they are constructed. Their perturbation direction is one shape applied to every surface, so the pairs differ in colour but not in the kind of difference between them — and a real batch differs from its standard in whatever way the process happened to drift, which is often a lightness shift or an ink-density change rather than a smooth spectral wobble. A set of realistic drift directions would give a wider spread, not a narrower one.
They are also all at exactly one unit. A specification at 2.0 or at 0.5 would sit in a different band of the curve where the formulae part, and the proportional spread there is different — smaller at 2.0, larger at 0.5.
And the whole exercise holds the illuminant and the observer fixed. A tolerance measured under one light is not the same tolerance under another, and that effect is comparable in size to this one. The two compound, and neither is usually in the contract.
Who found it, and when
That different formulae accept different batches is the oldest practical complaint in industrial colour control, and it is why every specification worth anything names its formula. The CIE’s own technical reports on tolerancing say so at length.
What is less often done is to hold the pairs fixed and vary the formula, which isolates the effect from everything else that varies between two laboratories. The usual experiment varies both — two labs, two instruments, two formulae, two sets of samples — and reports an overall disagreement that cannot be attributed. Constructing pairs to sit exactly on one boundary and then reading them elsewhere costs nothing and attributes the whole of the result, and it is available to anybody with the formulae in one program.
Where the ladder goes next
A camera profile is fitted by a linear least-squares solve in tristimulus space, which is an objective, and it is on nobody’s menu. Refitting the same 3×3 to minimise each of the six units instead turns a linear solve into a nine-parameter search — and moves the matrix, which means the camera renders different pixels rather than merely reporting a different score.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A tolerance has no light level appearance model · chroma · ciede2000 · colour difference · specification · tolerance
- A unit rests on a space that was ranked appearance model · calibration · chroma · ciede2000 · colour difference · just-noticeable difference
- A difference has no rate chroma · ciede2000 · colour difference · just-noticeable difference · tolerance
- The weighting is the disagreement appearance model · calibration · chroma · ciede2000 · colour difference
- A budget drawn through one hue colour difference · delivery · specification · tolerance
- A dial through a discrete menu calibration · chroma · ciede2000 · colour difference
What links here
The 8 essays that link to this one and share the most of its objects, of 9 that link here.
The objects this essay names
Each one links to every other essay that touches it.
AcceptanceAppearance modelCalibrationChromaCIEDE2000Colour differenceDeliveryJust-noticeable differenceSpecificationTolerance