Difference and uniformity

One unit in another room

Twenty-three pairs built at exactly ΔE00 1.000 under D65, re-measured under every change of light this site models with the observer adapted to each, come out anywhere between 0.64 and 1.57. A tolerance is written as a property of a pair and it is a property of a pair and a room.

Assumes A tolerance needs a second number and What no adaptation can remove.

A colour tolerance is a number in a contract. It says that two things are close enough if the difference between them is under some limit, and the limit is quoted with an illuminant and an observer beside it because everybody involved knows those matter. A shape rather than a number would be better still, and is a separate argument.

What the illuminant is understood to do is decide the numbers the two samples produce. What it also does, and what nothing in the document accounts for, is decide the size of the difference between them — because the light multiplies both samples, and the difference between two products is not the product of the difference.

A tolerance of one unit, re-measured under every light in the census. Every one of the 23 pairs behind this figure is at exactly ΔE00 1.000 under D65 by construction. Each bar is what those same pairs measure under another light, after the observer has adapted to it: the line is the median and the bar spans the pairs. A tolerance is written as a property of a pair and it is not one — the light multiplies both members, and the difference between two products is not the product of the difference. The widest row is a lens at twenty against a lens at seventy, spanning 0.81 to 1.57.
Fig. 1 Twenty-three pairs, every one of them at exactly ΔE00 1.000 under D65 by construction, re-measured under each of the census’s changes of light with the observer adapted to the new one. The line is the median and the bar spans the pairs. The specification’s number is the vertical rule.

The claim

A difference is not invariant under a change of light, and adaptation does not make it one.

  • Pairs at exactly one unit under D65 run from 0.64 to 1.57 across the census, after the observer has adapted — a factor of 2.4 between the extremes.
  • The medians move too, from 0.79 under two bounces off a green wall to 1.07 under tungsten, so it is not merely a spread around a stable centre.
  • The widest single row is fifty years of lens yellowing, 0.81 to 1.57, which is a difference between two observers rather than between two rooms and cannot be removed by controlling the light at all.
  • The most stable row is a change to a blackbody at the same colour temperature — 0.98 to 1.02 — which is the smallest change in the census and the one closest to being a pure gain.
  • And a tolerance of one unit is therefore not a quantity a specification can enforce without saying more than any specification says.

Why a difference moves at all

The arithmetic is short. Two surfaces ρ_A and ρ_B under a light E give tristimulus values A(E ρ_A) and A(E ρ_B), and the difference between them is A(E (ρ_A − ρ_B)) — the same integral applied to the difference of the reflectances.

Under a different light it is A(E′ (ρ_A − ρ_B)), which is the same difference-of-reflectances weighted differently. So the tristimulus difference between a pair changes with the light exactly as any single surface’s tristimulus values would: linearly, and by a matrix that is the census’s T.

Adaptation applies the same gain to both members, so it applies that gain to the difference too — and a gain scales a difference. What it cannot do is make the scaling one, because the gain is the ratio of the whites and the pair’s difference does not live along the white.

What is held fixed, and how

Everything here rests on the pairs being identical in size to begin with, so the construction matters.

Each pair starts from one surface in the site’s three-dimensional family. A second surface is made by walking along a fixed direction in that family until the CIEDE2000 difference under D65 lands on exactly 1.000, found by bisection to a millionth of a unit. Twenty-three pairs survive that construction; those that would need a step taking them outside the physical range are dropped.

So every pair enters the comparison at the same number. Whatever spread appears afterwards is entirely the light’s doing, and the machinery refuses to run at all if fewer than six pairs reach the target — a sample of three would be an anecdote about three pairs.

The rooms, in order

The most stable row is D65 to a blackbody at the same colour temperature: 0.98 to 1.02, a spread of four hundredths. That change is the smallest in the census, ΔE00 2.24 unadapted, and it is very nearly a gain — so the pairs are scaled by very nearly the same factor and hold their sizes.

The daylight rows are next, running about 0.92 to 1.16 across all three. Tungsten is wider, 0.93 to 1.25, and moves its median up: a pair at one unit under D65 is at 1.07 under tungsten on average and up to 1.25 for the worst of them.

Two bounces off a green wall is the row that moves the median furthest down, to 0.79 — the same corner that makes a metameric match fail makes an ordinary difference shrink, because the sharpened illuminant weights the part of the spectrum most of these pairs differ in less heavily.

The same wall, applied once and applied twice. A room lit by light that has bounced off its own walls is a change of illumination like any other, and a corner is the same change applied twice. Squaring a reflectance sharpens it, a sharper change of light is further from being a gain, and the residual an adapted observer is left with therefore grows faster than the change does: the second bounce is 1.33 times the change and 1.96 times the residual. This is the adaptation half of what a corner does to a metameric match.
Fig. 2 The two surface rows of the census: one bounce off a green wall, and two off the same wall. The second bounce is what moves this essay’s median furthest down, because each bounce weights the middle of the spectrum — where most of these pairs differ — less heavily than the one before it.
Every change of light this site models, and how much of it a gain removes. Each row is a change of illumination. The pale bar is how far it moves an ordinary surface for an observer who does not adapt; the solid bar at its left end is what is left after the observer has applied the one gain adaptation gives them, which is the ratio of the two whites in the CAT16 basis and is not fitted to anything. Sorted by the fraction left rather than by the size of the change, because the two orderings are different: the largest change here is removed almost entirely and the worst row is a change less than a third its size.
Fig. 3 The census the rows come from, for reference. The tolerance figure is this table asked a different question — not how far a surface moved, but how far two surfaces moved apart.

Which direction the errors go

A spread is not the whole story if the movement has a direction, and here it has one that is worth stating on its own because it inverts the intuition a quality manager would bring.

Eight of the fourteen rows have medians below one. That is, for most changes of light, an ordinary pair gets smaller, and a specification enforced under the wrong lamp is more likely to pass a bad pair than to fail a good one.

The reason is that most of these pairs differ across the middle of the spectrum, where a broad, warm light weights less heavily and a sharpened one weights unevenly. A change that concentrates the light — a corner, a phosphor’s bands, a warm lamp — puts less of it where these pairs differ, and the difference shrinks.

That direction is the awkward one for anybody enforcing a limit. A test that is systematically lenient in the field and correct in the laboratory produces exactly the pattern that gets blamed on the supplier: everything passes at the point of inspection and complaints arrive later, from people looking at the goods somewhere else.

The exception is tungsten, and it is the exception in the useful direction: it moves the median up, to 1.07. So the two commonest judging conditions in the world — a warm domestic lamp and a shop’s fluorescent tube — sit on opposite sides of the specification’s number, and the same pair can fail in one and pass in the other with nothing whatever changing about the goods.

The row that is not a room

The widest row in the whole comparison is not a change of light in any useful sense. Fifty years of lens yellowing takes the same pairs from 0.81 to 1.57 — a spread of three quarters of a unit on a tolerance of one.

That row belongs to a person rather than to a room, and it cannot be controlled by specifying the illuminant, the geometry or the instrument. Two inspectors of different ages, in the same booth, looking at the same pair, are measuring different quantities, and the spread between them is larger than the spread produced by any lamp here.

This is the population argument arriving through a different door. That essay measured how many people would accept a pair the instrument passed; this one measures how much the size of the difference itself moves between observers, which is upstream of any acceptance decision.

The share of itself each change leaves behind, and the smallest is inside the eye. The residual as a fraction of the change rather than as a colour difference, which sorts the census differently. At the top is the macular pigment — the filter in front of the central few degrees of one's own retina — leaving 2.4 per cent of itself. It is a fixed transmittance multiplying the light and the white together, which is as close to a pure gain as anything here gets, and it is why nobody notices they have one.
Fig. 4 The census drawn as the share of itself each change leaves behind rather than as a colour difference, which sorts it differently. The two ocular rows are filters inside the observer: a gain cannot remove them, and no viewing booth can standardise them away.

What a specification would have to say

A specification that wanted to be enforceable would have to name three things it does not currently name.

The light the pair will be judged under, not merely measured under. Those are usually different, and the document names the second. Measurement is done under a stated illuminant computed from a spectrophotometer’s reading; judgement happens under whatever is in the room.

How far the two samples differ spectrally, not merely colorimetrically. The spectral index exists precisely for this: a pair whose spectra agree transfers exactly, with a fitted constant term of 1.014 that nothing required, and a pair whose spectra differ transfers badly. The rows in this essay’s figure with the widest bars are the ones containing spectrally distant pairs.

And the observer’s age, or an acceptance of the range it implies. A tolerance of one unit with a 0.76-unit observer spread inside it is a tolerance that cannot distinguish a pass from a fail.

One tolerance decision, about pairs that agree less and less about the spectrumEvery point is a pair of samples that the reference observer reports as exactly ΔE00 1.0 apart — the same number, the same decision, the same line in the same specification. Along the axis is how far apart their two reflectances are. Up the side is the 95th percentile of what two hundred other eyes report. It runs from 1.13 to 2.32. The document records the horizontal line and not the axis it is plotted against.1.131.271.532.32what the specification says: ΔE00 1.00102001295th percentile ΔE00 over the populationspectral index — weighted rms difference, per centΔE00 1.0 held fixed200 observers, spectral index against the 95th percentile
Fig. 5 A family of pairs held at one difference and separated by how far apart their spectra are. That separation predicts how a pair travels between lights, which is the missing field in every tolerance document this site has looked at.
Whether the second number changes any answer. Each row is a tolerance. The bars are how often the two readings of the same document agree with what the population actually reports — the difference alone, and the difference with the index beside it. At most tolerances the two agree, because the pairs are comfortably inside or comfortably outside. At ΔE00 1.5 they disagree about 7 pairs of 21, and the two-number reading is right about every one of them. A field that changed no decisions would be a column of numbers.
Fig. 6 And what adding it changes about the decisions. The pairs a colorimetric tolerance passes and a spectral one fails are exactly the ones this essay’s widest bars are made of.

What happens at other tolerances

Everything above is built at exactly one unit, which is the commonest number in a specification and is not the only one. The effect does not scale the way an error budget would.

Rebuilt at two units, the pairs travel proportionally further in absolute terms and by a smaller fraction of themselves. That is CIEDE2000’s doing rather than the light’s: the formula’s weighting functions are calibrated at the threshold, and a two-unit pair spans more of the space than two one-unit pairs laid end to end. So a wider tolerance is slightly more robust to a change of light, measured as a percentage, and slightly less robust measured in units.

Rebuilt at half a unit, the fractional spread grows. A pair at 0.5 under D65 can be at 0.35 or 0.6 elsewhere, and half a unit is inside the range where the formula’s own disagreements with itself — the region where it is not smooth — are comparable with the effect being measured.

The practical reading is that a tight tolerance is not a tight tolerance. Halving a limit from one unit to half a unit does not halve what gets through, because a growing share of what the number is measuring is the light, the observer and the formula rather than the pair. Somewhere below half a unit the specification stops being about the samples at all.

That is a familiar conclusion in this trade arrived at from an unfamiliar direction. It is usually stated as instrument repeatability: a spectrophotometer’s own noise is a few hundredths, so a tolerance near a tenth is measuring the instrument. This essay adds a second floor, larger than that one, and it is not in any instrument’s specification because it is not a property of the instrument.

A third of the range is the lamps and two thirds is not

The headline factor of 2.4 is quoted as what a change of light does, and the rows the essay names say which changes contribute it.

row range spread factor
a blackbody at the same colour temperature 0.98–1.02 0.04 1.04
the daylight rows 0.92–1.16 0.24 1.26
tungsten 0.93–1.25 0.32 1.34
fifty years of lens yellowing 0.81–1.57 0.76 1.94

Every lamp in the census sits between 0.92 and 1.25 — a factor of 1.36. The census’s full range of 0.64 to 1.57 is a factor of 2.45, and the two ends belong to the two rows that are not lamps: the top to the observer’s lens, and the bottom to a corner. On a logarithmic scale the lamps supply 34 per cent of the range and the observer and the geometry supply the other 66.

The same census, sorted by where the change of light came from. Each row is a change of illumination. The pale bar is how far it moves an ordinary surface for an observer who does not adapt; the solid bar at its left end is what is left after the observer has applied the one gain adaptation gives them, which is the ratio of the two whites in the CAT16 basis and is not fitted to anything. Sorted by where the change came from. The two kinds of light that existed before electricity sit at the top and leave the smallest share of themselves behind; the discharge lamps are worse, and the worst of them is d65 to a triphosphor tube at 33 per cent.
Fig. 7 The same rows grouped by where the change of light came from. The lamps are one group and they are the narrow one: both ends of this essay’s 0.64-to-1.57 range belong to rows that are not lamps at all — a corner at the bottom and an observer’s own lens at the top.

That reorders the essay’s own conclusion in a useful way. A tolerance is a property of a pair and a room is right, and the room is the smaller half of it. A specification that named the judging lamp perfectly — which is the fix a viewing-condition standard offers, and the only one any of them offers — would have bought a third of the problem.

The lens row alone is 2.4 times the widest lamp row and nineteen times the narrowest, which puts a number on the spread between them is larger than the spread produced by any lamp here. Two inspectors thirty years apart disagree about the size of a difference by more than the gap between a viewing booth and a tungsten lamp.

The leniency is in the size, not in the count

Eight of the fourteen rows have medians below one is offered as the evidence that a specification enforced under the wrong lamp is systematically lenient, and eight of fourteen is 57 per cent — a bare majority, and not much of an argument on its own. A fair coin gives eight or more out of fourteen about four times in ten.

The two medians the essay names make a much stronger case. The green-wall row sits 0.21 below one and the tungsten row 0.07 above — so the downward excursion is three times the upward one, on the two rows chosen as the extremes of the median’s own range.

The asymmetry is in the magnitude rather than in the count, and that is the right form of the argument anyway: a test whose errors are eight-to-six in one direction but equal in size is unbiased on average, and one whose errors are seven-to-seven but three times larger downward is not. The second describes this census and the first does not.

It also sharpens the practical warning. A pair mismeasured in the lenient direction is out by up to a fifth of the tolerance; one mismeasured in the strict direction is out by less than a tenth. So the false passes are not merely more numerous, they are worse, and the complaints-arrive-later pattern the essay describes follows from the sizes rather than from the tally.

The half-unit figures point the other way

Rebuilt at half a unit, the fractional spread grows is the section’s conclusion and the two numbers offered for it say the opposite.

A pair at 0.5 landing between 0.35 and 0.6 is a fractional range of 0.70 to 1.20, a factor of 1.71. The one-unit census runs 0.64 to 1.57, a factor of 2.45. As quoted, the fractional spread shrinks by a third rather than growing.

There is a reading that rescues it, and it is worth stating because it is probably the intended one: the half-unit figures look like a lamp row rather than the whole census, and against the lamps’ own 1.36 at one unit a fractional range of 1.71 at half a unit is indeed a growth — by about a quarter. On that reading the claim holds and the comparison in the text is between a subset and a whole.

Which matters for the section’s conclusion rather than for its direction. A tight tolerance is not a tight tolerance survives either way, because both readings have the fractional spread at half a unit above the fractional spread of the lamps at one. What does not survive is any statement about how much worse a half-unit tolerance is, since the two numbers being compared are not measured over the same set of rows.

Who found it, and when

That a pair’s difference changes with the illuminant is the definition of illuminant metamerism, and it is the oldest observation in industrial colour. The special index of metamerism was standardised to quantify it in the 1970s, and it works by computing the difference under a second illuminant and quoting it.

What that machinery does not do is ask the question in an adapted observer’s terms, and it does not run over a set of pairs held equal to begin with. The special index is computed for a pair that matches under the first illuminant — a metameric pair, ΔE zero — so it measures the growth from nothing. This essay measures the change in a difference that was not zero, which is the case every tolerance in a contract is actually about.

The distinction matters because the two behave differently. A metameric pair’s difference can only grow. An ordinary pair’s can grow or shrink, and in this census it shrinks more often than it grows: eight of the fourteen rows have medians below one.

The practical consequence of that asymmetry is that the special index is not a bound on this essay’s quantity. A pair with a low index — spectra close, metamerism slight — can still move by a tenth of a unit here, because the index is computed for a pair that started at zero and this one did not. Two pairs with the same colorimetric difference and the same special index can travel differently, and nothing in the standardised machinery separates them.

The measurement that does separate them is the weighted spectral distance, and it was built on this site for a different question — how much of the population would accept a pair — before it turned out to answer this one as well.

What was computed, and how

The pairs are constructed by bisection to ΔE00 1.000 ± 10⁻⁶ under D65 with the 2° observer, on the same three-dimensional reflectance family the census uses. Directions are alternated between the two modulation basis functions so the set is not all one kind of pair.

For each row, both members are multiplied by the new light, the adapted observer’s gain is applied — the ratio of the whites in the CAT16 basis, not fitted — and the difference is recomputed in CIELAB against the original white. Reporting against the original white is the choice that makes this a question about an adapted observer: they are judging the new stimulus with the machinery they walked in with.

The gate requires the whole set to span at least 0.3 units and requires at least one pair to move by 0.4, so a future change that quietly made differences invariant would stop the build rather than silently strengthening the claim.

Where it stops

The pairs are constructed rather than sampled from any real collection of colorants. They span the site’s smooth three-dimensional family, so they under-represent the spectrally spiky pairs — a fluorescent-pigment against a conventional one, say — where the effect would be much larger. The numbers here are a floor.

The comparison uses CIEDE2000 throughout, and the formula has opinions that vary across the space. Some of the spread within a row is the formula’s own non-uniformity rather than the light’s, and the two are not separated here.

And the observer is assumed to adapt completely to the new light. Incomplete adaptation would move the medians further and would not narrow the bars.

One further caution, and it is the one most likely to be misread. The spread reported here is a spread across pairs, not an uncertainty about any one of them. A given pair’s behaviour is entirely determined: it has a spectral difference, and that difference decides how it travels. What the figure says is that a specification naming only a colorimetric limit has not said which pair it is talking about, so it has bought the whole spread. A specification that named the spectra as well would have bought a single number, and that is the fix the spectral index essay already priced.

Where the ladder goes next

If a tolerance is a property of a pair and a room, then it is also a property of a pair and an instrument, because the instrument’s geometry decides what reflectance it reports at all. The two standard geometries disagree by ΔE00 8.35 on a dark sample, which is eight times the tolerance this essay is about, and the difference between them is not multiplicative and cannot be adapted away.

The other direction is the population. The lens row here says the observer contributes more spread than the lamp does, and a gamut boundary turns out to have the same problem: the quantities this subject reports as single numbers are, one after another, distributions with the observer inside them.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AcceptabilityAdaptationAssertionChromatic adaptationCIEDE2000ΔEIlluminantMeasurement uncertaintyMetamerismQuality controlSpecificationTolerance