Difference and uniformity

A tolerance is a shape

A total colour difference below one, or every component within one — the two sound like the same requirement stated twice. They are different shapes, they disagree about most of what either accepts, and which one a supplier is held to is worth money.

Assumes How far apart are two colours and A threshold is not a unit.

18 min read 10 figures Computed, not quotedSay which colour

A colour specification in a supply contract is a number and a rule. The number is usually 1, and the rule is usually one of two: a total colour difference below the number, or each component within it.

They sound like the same requirement. They are different shapes in the same space, and they disagree about a startling proportion of everything either of them accepts.

A box tolerance and a ΔE tolerance around the same colourA slice through CIELAB at L* 50, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 2.1 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 74% — accepted by one specification and rejected by the other, on the same measurement.gold outline: ΔE2000 = 1blue box: ±1 in each componentelongation 2.13disagree on 74%of everything either acceptsΔa*at L* 50, a* 55, b* 25 — the shape changes elsewheresame measurement, two verdictsCIE 1931 2° observer
Fig. 1 A slice through CIELAB with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is elongated in one direction and the box is square. Of every sample either rule accepts, the two disagree about most of them.

Where in the space the contour is traced changes its shape as much as its size, and one centre is one case.

A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 55, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 2.2 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 58% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 2 The same two rules around a blue reference rather than an orange one. The ellipsoid has turned and changed its aspect ratio; the box has not, because a box is a statement about coordinates and not about colours.
A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 20, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 1.5 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 50% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 3 And around a dark near-neutral, where the two rules disagree most. A specification that names a box is naming a different region at every point it is applied, and the difference is largest exactly where a printer’s blacks live.

Two more centres and a wider tolerance say that neither the shape nor the mismatch between the two rules is a property of any one place in the space.

A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 70, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 1.8 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 78% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 4 A light yellow-green reference, which is where a press has most room and a specification is most relaxed. The ellipsoid has turned again and the box has not, because a box has nothing to turn.
A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 40, with the ΔE2000 = 2 contour traced point by point and a ±2 component box drawn over it. The contour is 3.1 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 62% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 5 And both rules loosened together at a dark blue centre. Doubling the number doubles the box in every direction and the ellipsoid in none of them equally, so the two rules disagree more the looser they get.

Why the ellipsoid is not a sphere

A component box is axis-aligned and the same size everywhere. A ΔE2000 limit is not.

The entire content of a modern difference formula is that perceptual difference is anisotropic and non-uniform. The formula weights lightness, chroma and hue differences differently; the weights change with position; and there is a rotation term coupling chroma and hue in the blues. So the set of colours within ΔE 1 of a reference is an ellipsoid whose size, elongation and orientation all depend on where the reference is.

That is not a defect being worked around. It is the measurement the formula encodes, and a formula that produced spheres would be claiming the space was uniform, which it is not.

The consequence for specification is immediate. A single box cannot match a family of differently-shaped ellipsoids. Any box either contains the ellipsoid, accepting samples the ΔE rule rejects, or is contained by it, rejecting samples the ΔE rule accepts, or crosses it and does both. There is no box that does neither.

How much they disagree

Set the box to the same numerical value as the ΔE limit — which is what a specification writer who thought the two were interchangeable would do — and measure.

Around a mid red, the two rules disagree about 74% of the samples either accepts. Near the neutral axis, 55%. Out in the saturated yellows, 82%.

“Disagree about” means: accepted by one rule and rejected by the other, on the same measurement, with no ambiguity about the measurement itself. A supplier and a customer applying these two rules to identical instrument readings will reach opposite conclusions most of the time.

The disagreement is also not symmetric. In some regions the box is the stricter rule and in others the ΔE limit is, and which one favours the supplier changes with the product’s colour. That is worth knowing before signing.

What a tolerance is actually for

The deeper problem is that neither rule is measuring the thing anybody cares about.

A tolerance in a contract exists to answer: will the customer accept this? That is a question about acceptability, and acceptability is not perceptibility. A difference can be plainly visible and entirely acceptable — nobody rejects a jumper because the two sleeves differ by a ΔE of 2 — and a difference can be barely visible and completely unacceptable, which is what happens when two parts of the same moulding meet at a seam.

Acceptability depends on the product, the arrangement of the parts, the customer’s expectations and the market, and none of those is in any formula. What the formula supplies is a defensible number, and the industry’s practice is to calibrate the number against experience: a threshold is set at whatever value stopped producing complaints, per product line, which is an empirical fix for a formula that answers the wrong question.

This is why the parametric factors exist. Textile practice uses kL = 2, halving the weight on lightness differences, because in fabric a lightness difference matters less than a chroma difference of the same size. That factor is not a correction toward a truer perceptual metric. It is a correction toward acceptability in one industry, and it was found by seeing what worked.

The number arrives with an error bar nobody writes down

A tolerance is applied to a measurement, and the measurement has its own disagreement.

Two instruments of different bandpass, looking at a sample with real spectral structure, differ by 2.34 ΔE. Under a discharge lamp rather than daylight, the same pair of instruments differ by 6.30 on a sample they agree about to within 0.88 under daylight.

So a specification of ΔE ≤ 1 on an interference pigment is a specification whose measurement uncertainty exceeds its tolerance by a factor of two. The pass-or-fail decision it produces is not a measurement of the sample; it is a measurement of which instrument was used.

The industry’s answer is procedural: both parties must use the same instrument model, and often the same physical instrument. That works, and it is an admission — a tolerance that only means something when the apparatus is fixed is a tolerance on an apparatus rather than on a colour.

What a good specification says

Everything above points at what a specification has to carry to mean anything, and the list is longer than one number.

Which formula. ΔE76, ΔE94 and ΔE2000 disagree by more than a unit on the same pairs, so a bare “ΔE ≤ 1” is not a requirement until the formula is named.

Which parametric factors. Left unstated they default to 1, which is the original experiment’s viewing arrangement rather than the product’s.

Which illuminant, and which observer. A pair of samples can match under one illuminant and not another, and a real lamp is not a standard illuminant — that is metamerism, and it is common in production, where a substitute pigment matches under the specified light and fails under the shop’s.

Which instrument geometry and bandpass, since the measurement uncertainty can exceed the tolerance.

And what the parts look like in use — adjacent or separated, which changes the acceptable difference by a factor of several in either direction.

A specification carrying all of that is a paragraph rather than a number. Most specifications are a number.

Why boxes persist anyway

Given all of this, the component box looks indefensible. It is not, quite, and the reasons are worth stating fairly.

It is diagnostic. A ΔE tells a supplier that something is wrong. Component limits tell them what: too light, too red, too blue. On a production line, where the response to a failure is an adjustment to a mix, knowing the direction is worth more than knowing the magnitude.

It is comprehensible without training. “Within half a unit of a*” is checkable by anybody with the instrument’s output in front of them. A ΔE2000 requires a computation with a dozen terms, several of them piecewise, and an implementation trap in the hue-angle arithmetic that a great deal of software has fallen into.

It is stable. A box does not change when somebody updates the formula. The move from ΔE76 to ΔE2000 changed a lot of pass-or-fail decisions on unchanged products, which is a genuine cost of using a better metric.

The best practice, where anybody can afford it, is both: component limits for diagnosis and a total limit for acceptance, with the two chosen deliberately rather than by setting them to the same number.

The tolerance that is not a number at all

The most reliable colour specification in industrial use is not a number, and the exception is worth stating because it is what everything above is an attempt to replace.

A physical master is a sample of the agreed colour, held by both parties, and the specification is: match this. Assessment is visual, under agreed lighting, by trained assessors, with the parts arranged as they will be in the product.

That method has obvious defects. Masters fade, they have to be physically shipped, they cannot be emailed, and assessor agreement is imperfect and hard to audit. Every one of those is a reason instruments and formulae exist.

It also has one property nothing numerical has managed to reproduce: it measures the right quantity. A visual assessment against a master under the product’s own conditions is a direct measurement of whether the difference is acceptable, which is the question the contract is actually about, rather than a proxy for it.

The industries with the most at stake still keep masters and use the numbers alongside them. That is not conservatism, and it is worth reading as evidence: eighty years of measurement has produced excellent tools for detecting and quantifying differences, and has not produced a replacement for asking somebody.

What was computed here

The ΔE contour is traced rather than assumed. Around each centre, bisection along a hundred and eighty directions finds the exact locus where ΔE2000 reaches the limit — fifty steps per direction, which is what it takes to measure an elongation reliably.

Fitting an ellipse would be wrong, and not only approximate. ΔE2000 has piecewise terms, so its unit contour is genuinely not an ellipse, and fitting one would smooth over the structure the figure exists to show.

The disagreement rates are computed over a low-discrepancy sequence rather than a random sample. That is a requirement rather than a preference: a build drawing random numbers would print a different percentage in every caption on every run, and a number that changes between builds cannot be quoted in prose. The additive-recurrence sequence covers the cube far more evenly than sampling would at the same count and gives the same answer on every machine.

The population is spread over a cube several times the tolerance, so the reported rates describe the region where the disagreement exists rather than being diluted by distant samples that both rules obviously reject.

One assertion runs in the gate, and its parameters are chosen to make it hard rather than easy. The box is set to the same numerical value as the ΔE limit — the comparison a specification writer who believed they were equivalent would produce — because a box far larger or smaller would disagree about almost everything and prove nothing. The assertion then requires a substantial disagreement at every one of the five points, so a result driven by one unusual region would fail.

Half the disagreement is a ball not being a box

The three disagreement rates are offered as evidence that a box cannot follow a family of differently shaped ellipsoids, and about half of each of them would be there if every ellipsoid were a perfect sphere.

A unit ball sits entirely inside a cube of half-width one, so the two accept different things by the volume between them: 4.19 against 8, which is 47.6 per cent of the union. That is the disagreement between a box and a round region of matched nominal size, with no anisotropy anywhere and no non-uniformity in the space.

Against the measured rates:

region measured from the shape alone from the anisotropy
near the neutral axis 55 % 47.6 % 7 points
a mid red 74 % 47.6 % 26 points
the saturated yellows 82 % 47.6 % 34 points

Near the neutral axis almost the whole disagreement is geometry, and the formula’s anisotropy adds seven points. Out in the saturated yellows the anisotropy adds thirty-four and is the larger term.

That is exactly what the formula’s own weightings predict. S_C = 1 + 0.045C and S_H rise with chroma, so the ellipsoid’s elongation against the lightness axis runs 1.4 to 1 at a chroma of 10, 2.3 at 30 and 3.7 at 60. Near the neutral axis the contour is nearly round and the box is arguing with a ball; in the saturated yellows it is nearly four times longer one way than another and the box is arguing with a cigar.

So the essay’s finding survives and its emphasis moves. The two rules disagree about most of what either accepts, and near the neutral axis they would disagree about nearly half of it if the space were perfectly uniform — because an axis-aligned box and any round region of matched size have less than half their volume in common, which is a fact about six corners rather than about human vision.

The disagreement is a function of where in the space the tolerance is centred, and a near-neutral light colour is the case a paper or a paint specification is usually written about.

A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 85, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 1.8 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 52% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 6 A slice at L* 85 near the neutral axis. The contour is 1.8 times longer one way than the other against a square box, and 52 per cent of everything either rule accepts is accepted by one and rejected by the other.

Which corners the box is accepting

The 47.6 per cent has a location, and it is the one a specification writer would care about.

A cube’s excess over its inscribed ball is entirely in the corners — the eight regions where two or three components are near their limits at once. A sample at (1, 0, 0) is on both surfaces; a sample at (0.8, 0.8, 0.8) has a total difference of 1.39 and passes the box comfortably.

So the box’s leniency is concentrated on samples that are off in every direction at once, which is the failure mode a customer is most likely to call unacceptable: not too light, not too red, but a bit of everything. And the box’s strictness is concentrated on samples off in one direction only, which is the failure a production line can correct with one adjustment.

That inverts the usual defence of component limits. The box is diagnostic where it is strict and blind where it is lenient — it tells a supplier what to change on exactly the samples the ΔE rule would have rejected, and passes without comment the ones where nothing single is wrong.

Tightening the ΔE tolerance while leaving the component box where it was is what happens when a specification is revised in one of its two halves, and the arithmetic is unkind.

A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 50, with the ΔE2000 = 0.5 contour traced point by point and a ±1 component box drawn over it. The contour is 2.1 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 69% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 7 The same centre with the contour at ΔE2000 0.5 and the box still at ±1. The shapes now disagree about 69 per cent of what either accepts, so halving one of the two numbers makes the pair of rules less compatible rather than more.

Six times, not two

The measurement-uncertainty argument quotes one of its two figures.

A specification of ΔE ≤ 1 on an interference pigment is a specification whose measurement uncertainty exceeds its tolerance by a factor of two uses the daylight number, 2.34. The same essay’s other figure — the same pair of instruments on the same sample under a discharge lamp — is 6.30, which is a factor of six.

And the discharge lamp is not the exotic case. It is the light in most shops, most warehouses and a good many inspection benches, which is the arrangement a supplier and a customer are most likely to be measuring in when they disagree.

So the sentence is right and understated by three times, and the corrected version says something stronger about the procedural remedy. Fixing the instrument model removes the whole of the 6.30, because the disagreement is between two instruments rather than in either — which is why the practice of naming the instrument is not fastidiousness but the largest single thing a specification can do, worth more than the choice of formula, more than the choice of shape, and more than the tolerance itself.

A mid-lightness blue is the last quadrant the sweep has not visited, and it is where the contour is roundest of all.

A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 60, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 1.5 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 72% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 8 A slice at L* 60 in the blue-green quadrant. The contour is only 1.5 times longer one way than the other — the roundest in this essay — and the two rules still disagree about a large share of what either accepts.

Where the model stops

The comparison is in CIELAB, at fixed lightness, on a two-dimensional slice. A real tolerance is three-dimensional and the third dimension is where a substantial part of the disagreement lives — lightness differences are weighted differently from chromatic ones by every formula and by every industry’s parametric factors.

More fundamentally, this essay compares two rules against each other and neither against a person. What settles which rule is right is acceptability data: samples at known differences, and judgements about whether a customer would accept them. That data exists, is proprietary, is industry-specific, and is not on this site.

So what is established here is that the two rules are materially different, which is enough to make the choice matter. Which of them better predicts acceptance for a given product is a question this essay cannot answer, and the honest form of the advice is: find out for the product, rather than assume the number transfers.

Loosening the ΔE tolerance while leaving the box is the other half of the revision, and it moves the disagreement the other way.

A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 50, with the ΔE2000 = 2 contour traced point by point and a ±1 component box drawn over it. The contour is 2.2 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 96% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 9 The same centre with the contour at ΔE2000 2 and the box still at ±1. The contour is 2.2 times longer one way than the other, so changing either number alone changes which rule is the binding one.

What the pictures cannot show

The figures draw tolerance regions as geometry, and a tolerance is a decision about a business relationship.

Nothing here can show what a ΔE 1 difference looks like, and that is not an oversight. Two swatches a stated ΔE apart could be drawn — but whether the reader sees a difference depends on their display, their room, and whether the swatches share an edge, and those three factors move the answer by more than the difference being illustrated. A figure captioned “these differ by ΔE 1” would be inviting a judgement the page cannot support.

The one demonstration that would work is physical: two painted panels, one seam, under a stated light. That is what colour tolerance work actually looks like, and it is the reason the field still runs on physical standards despite eighty years of measurement.

A dark saturated green is where the two rules are furthest apart in this collection’s own sweep, and it is worth drawing under the wider-field observer.

A box tolerance and a ΔE tolerance around the same colour. A slice through CIELAB at L* 35, with the ΔE2000 = 1 contour traced point by point and a ±1 component box drawn over it. The contour is 1.6 times longer in one direction than the other, and the box is square. Of every sample either rule accepts, the two disagree about 73% — accepted by one specification and rejected by the other, on the same measurement.
Fig. 10 A slice at L* 35 in the green quadrant, scored with the CIE 1964 observer. The contour is the roundest of the three at 1.6, and the disagreement is the largest at 73 per cent — so a rounder contour does not mean a better-behaved specification.

What passes and what is accepted

There is a final gap worth naming, because it is where the whole apparatus meets the thing it was built for.

A tolerance produces a pass or a fail. A customer produces an acceptance or a rejection. These are different events, and the correlation between them is the only measure of whether a tolerance is any good.

Nothing in this essay measures that correlation, and nothing can from a computer. Establishing it requires samples at known differences, real assessors, and the product in the arrangement it will be used in — the parts adjacent or separated, at the size they will be, under the light of the place they will be seen.

The industries that do this well collect the data. Automotive colour work maintains physical master panels and assessor panels, and the numerical tolerance is calibrated against them rather than derived. Textile practice does the same with its parametric factors.

The industries that do it badly copy a number out of a standard, apply it to a product nobody assessed, and discover the mismatch through complaints. That is the failure this essay is really about — not choosing the wrong shape, but treating a shape as a substitute for having asked anybody.

Who found it, and when

Colour tolerancing became a formal discipline when supply chains lengthened enough that the customer and the supplier were not in the same building. The textile and automotive industries drove it, both because their products come in colours and because both assemble parts made in different places that have to match.

The CMC formula, published by the Colour Measurement Committee of the Society of Dyers and Colourists in 1984, was the first widely adopted formula built explicitly around acceptability rather than perceptibility, and it is still specified in textile work. Its ellipsoids were fitted directly to accept-or-reject judgements from industrial assessors, which is a different and more honest target than perceptual difference.

ΔE2000 arrived in 2001 with better perceptual uniformity and a more complicated formula, and its adoption was slower than its merits would suggest — because changing the formula changes which existing products pass, and a metric that reclassifies a warehouse is expensive regardless of being better.

That is the pattern this essay ends on. The tolerance in a contract is not chosen to be correct. It is chosen to be defensible, stable, and compatible with what was agreed last time, and every technical improvement has to argue against all three.

Where this goes next

The formulae the numbers come from are how far apart are two colours. The reason a threshold and a tolerance are different quantities is a threshold is not a unit. And the measurement uncertainty a tolerance sits on top of is what the instrument reports.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 36 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AcceptabilityCIELABΔEEllipsoidLightnessQuality controlSpecificationThresholdTolerance