Difference and uniformity

Two departures that partly cancel

A glossy translucent sample has two of this round's four departures at once, and the expectation was that they would compound. They do the opposite. An aperture takes light away that went into the material and came back too far out; an interface returns light that never went in at all — so measuring either alone overstates what both together do, on every material tested.

Assumes Either factor being zero, An aperture is a filter and A room is not a sphere.

Two errors in the same measurement are usually assumed to add, or at worst to add in quadrature. These two subtract, and the subtraction is larger than the smaller of the two.

An aperture and a gloss lobe, apart and together. Six materials, each measured through a four-millimetre radius and each given a gloss lobe, alone and at the same time. The pale bar is what the two cost added together as if they were independent; the dark one is what they cost when both are present. Every material comes out below the sum, by between 0.8 and 3.6 ΔE₀₀. The two departures partly cancel: the aperture removes light that went into the material and came back out too far away, and the interface returns light that never went in at all. Measuring either one alone therefore overstates what both together do, which is the opposite of the way interacting errors are usually assumed to behave.
Fig. 1 Six materials measured through an aperture, given a gloss lobe, and both at once. The pale bar is the two costs added as if they were independent; the dark one is what they cost together.

What the two departures are departures of, and how far an audit of this kind can be pushed, are the two things the interaction table sits inside.

The six arguments a surface's response has, and the one this model keeps. A surface's response to light is a function of six arguments: the wavelength, direction and place the light arrives with, and the wavelength, direction and place it leaves with. The model every colour here is computed from keeps one number per wavelength, which means it takes the diagonal of the first pair, integrates the second away, and assumes the third pair equal. Each departure drawn here restores one of them. The fourth departure is not on the diagram: the wavelength grid is the range of the index that was kept rather than an index that was dropped, which is why it is the cheapest of the four to fix and was still not fixed.
Fig. 2 The decision both belong to. A surface’s response has six arguments and the model keeps one number per wavelength, so the two that cancel are two of the five it dropped — which is why they are not independent in the first place.
Which of the collection's published quantities a departure can be pushed through. The six quantities the previous round recomputed under six different colour-difference units, and whether the same treatment works for a departure. Two do: the adaptation census and the metameric pair both take reflectances and a light, which is what a departure acts on. Four do not, and the reasons are different in each case rather than a single obstacle. A unit is a function applied to the answers, so it can be swapped at the end of any computation; a departure changes the object at the start, so it has to be accepted by every stage in between. That is the practical difference between auditing a convention and auditing a structure.
Fig. 3 And the boundary of what this instrument can be pointed at. Only two of the collection’s published quantities can be pushed through, which is why the cancellation is measured on single readings rather than on the census.

The claim

An aperture and a gloss lobe on the same sample cost less together than they do apart, on every material tested.

  • On a pigmented plastic, the aperture costs 1.96 ΔE₀₀ alone and the lobe 1.87 alone; together they cost 1.54, not 3.83.
  • The cross term is negative everywhere, from −0.80 on coated paper to −3.56 on skin, and it is between 118 and 286 per cent of the smaller of the two effects.
  • The mechanism is a subtraction of light against an addition of it. The aperture removes body light; the interface adds light that never entered the body.
  • The expectation was the opposite. Two departures on one sample were predicted to compound, and the arithmetic refused it.
  • And it breaks an error budget in the safe direction, which is the least useful direction for a budget to be wrong in.

The two departures on one sample

A glossy translucent sample is not an exotic object. It is a pigmented plastic sheet, a varnished print, a polished stone, a piece of skin — most of what a laboratory is asked to measure that is not paper.

Such a sample has two things going on at once. Light meeting it partly reflects at the interface without ever entering the material, and partly enters, scatters, and comes back out somewhere else. The first part is the specular term the geometry decides; the second is the body term the aperture decides.

The two are usually treated as independent because they are physically separate — one happens at a boundary, the other happens underneath it. That separateness is exactly what makes the interaction worth computing rather than assuming.

Why they cancel

The aperture’s error is negative: a finite hole throws away the light that came back outside it, so the reading is low.

The interface’s error, in a geometry that collects the specular component, is positive: light is added that carries no information about the pigment.

And the two are not simply two errors on the same number — they act on different parts of it. The aperture reduces the body term and does nothing to the interface term, because light reflected at the boundary never travelled sideways at all: it comes back exactly where it arrived.

So a narrow aperture on a glossy sample reads a smaller body plus an unchanged pedestal, which is closer to the correct total than either error alone would suggest. The pedestal partly fills the hole the aperture made.

That is a real cancellation rather than a coincidence of magnitudes, and it has a sign that can be predicted from the mechanism before any arithmetic: whenever one departure removes light that took a long path and another adds light that took none, they will oppose.

The numbers

Six materials, each measured through a four-millimetre radius, each given a lobe of roughness 0.05, and each computed both ways.

material aperture alone gloss alone both cross term
coated paper 0.53 0.96 0.69 −0.80
wall emulsion 1.16 1.12 0.94 −1.33
pigmented plastic 1.96 1.87 1.54 −2.29
pale marble 12.65 1.33 11.13 −2.86
skin 5.83 3.02 5.29 −3.56
candle wax 17.55 1.12 15.46 −3.21

Every cross term is negative. On coated paper the two together cost less than the gloss does alone. On marble the cancellation is 2.86 ΔE₀₀, which is more than twice the whole gloss effect.

The relative size is the striking part: the cross term is between 118 and 286 per cent of the smaller of the two main effects. An interaction that large is not a correction to an additive model; it is a statement that the additive model has the wrong shape.

The second departure improves the reading on every material

The table says something stronger than partly cancel and says it on all six rows rather than on one.

On coated paper the two together cost less than the gloss does alone is noted as the extreme case. It is not the extreme case; it is the general one. On every material the pair costs less than the larger of the two alone:

material larger alone both improvement
coated paper 0.96 0.69 28 %
pigmented plastic 1.96 1.54 21 %
wall emulsion 1.16 0.94 19 %
pale marble 12.65 11.13 12 %
candle wax 17.55 15.46 12 %
skin 5.83 5.29 9 %

Adding the second departure makes the measurement more accurate, by between nine and twenty-eight per cent, without exception. That is what a cross term larger than the smaller effect means, stated in the form somebody would act on: it is not that the two are less than their sum, it is that either one alone is worse than both.

The prediction that follows is testable and slightly absurd. Varnishing a translucent sample would improve an aperture measurement of it — the pedestal fills the hole — and on a coated paper it would improve it by more than a quarter. The essay’s own mechanism says exactly that and the table’s third column is the evidence for it, and neither is stated as advice anyone would follow, which is the mark of a result rather than a rationalisation.

The cancellation is a fixed geometry, until the metric bends

Between 118 and 286 per cent of the smaller effect is a 2.4-fold range, which reads as a loose result. Written as a geometry it is much tighter on the rows where it can be written at all.

Treat the two departures as vectors in a colour-difference space and solve for the angle between them from the three measured magnitudes:

material angle between the two departures
coated paper 44.5°
pigmented plastic 47.3°
wall emulsion 48.7°
skin 64.6°
pale marble no angle exists
candle wax no angle exists

Three of the six sit within four degrees of each other, at about 47°, which is a far more specific statement about the mechanism than a percentage range: the aperture’s error and the gloss pedestal point in directions a little under half a right angle apart, on three materials with nothing in common but being moderately translucent.

The last two rows are the informative failure. On marble and wax the implied cosine comes out at 1.13 and 1.79 — greater than one, so no Euclidean arrangement of two vectors produces those three magnitudes. The cancellation there is larger than any geometry allows, which means it is not geometry: it is CIEDE2000 compressing. Both materials have an aperture effect above twelve units, where the formula’s chroma and hue denominators are large, so the combined difference is discounted more than its two parts are.

That splits the six rows into two claims rather than one. On the four materials whose two effects are comparable, the cancellation is a real opposition of directions and the angle is nearly constant. On the two where the aperture dominates by a factor of ten, part of the reported cross term is the difference formula rather than the physics, and the essay’s cross term is 214 and 287 per cent of the smaller effect on those rows is measuring the metric.

Which does not weaken the finding and does relocate it. The sign is mechanical everywhere — one departure adds light and the other removes it, so they must oppose — and the size is mechanical only where the two are within a factor of about three of each other. Quoting a percentage range across all six mixes the two.

Why an additive budget is wrong in the worst way

An error budget adds contributions, and everybody who writes one knows it is approximate. The usual defence is that adding is conservative: an additive budget over-estimates the total, so a process inside its budget is safely inside.

Here that defence holds — the additive figure is larger than the truth — and it is worth seeing why that is not comforting.

A conservative budget that is conservative by a factor of two does not tell a laboratory it has margin; it tells it nothing. A process budgeted at 3.83 ΔE₀₀ that actually delivers 1.54 has 2.29 of headroom nobody knows about, and the decisions taken on the strength of the budget — a tighter tolerance rejected as unreachable, an instrument bought that was not needed — are wrong in the expensive direction.

And the sign is not guaranteed. It is negative for these two departures because one adds light and the other removes it. A pair where both remove light would compound, and this round has such a pair available: a narrow aperture on a fluorescent sample removes body light while the fluorescence adds light in the same place, but a narrow aperture plus an ultraviolet-free lamp removes both. The sign of an interaction is a property of the pair and has to be computed pair by pair.

Four departures from the model equation, each at an ordinary strength. What each of the four assumptions inside a colour integral costs, in ΔE₀₀, on a stated sample under a stated light. The wavelength index is a coated printing paper measured with and without the ultraviolet of D50; the range is the same paper integrated from 300 nanometres and from 380; the place index is a pigmented plastic through a four-millimetre radius; the direction index is an eggshell paint beside a window. The spread is a factor of 7.0. This is a ranking of four examples rather than of four departures — each of them can be made larger by choosing a more extreme sample, and the marble in the same collection of materials reaches 12.7 on the index that comes third here.
Fig. 4 The four departures alone, at ordinary strengths. Adding any two of these bars is an operation the interaction table says is not available.

What a positive interaction would look like

The sign is not a law and it is worth constructing the opposite case, because a rule with no counterexample is usually a rule nobody has tested.

Take an aperture and an ultraviolet-free lamp on a brightened translucent sheet. The aperture removes body light; the missing ultraviolet removes fluorescent light. Both subtract, and both subtract from the same part of the reading — the part that came out of the material — so the two compound rather than cancel.

Take an aperture and a wider field instead, and the two are not independent at all: widening the illumination is one of the two ways to remove the aperture’s error, so their interaction is as large as either of them and the pair is better described as one departure with two remedies.

Three pairs, three different structures: cancelling, compounding, and redundant. The only way to know which is to write down what each departure acts on, and that is a two-line argument rather than a computation — which is fortunate, since there are six pairs among four departures and this round has computed one.

What was computed, and how

Each material’s bulk reflectance is computed from its coefficients, and its aperture reading through a four-millimetre radius from the same kernel.

The gloss term is added as a pedestal, and the pedestal is not declared but computed: a Trowbridge–Reitz lobe of roughness 0.05 on a dielectric of index 1.5, integrated over the hemisphere for a detector eight degrees off the normal. Because the lobe contains no body reflectance, the integral is the same number at every wavelength, which makes the pedestal spectrally flat — and that flatness is a property of the model rather than an assumption, since Fresnel’s term at a fixed index barely varies across the visible.

Four colours are then computed for each material: no departure, aperture only, gloss only, and both. The three differences from the first are the three numbers in the table, and the cross term is the third minus the sum of the first two.

The assertion behind the figure requires the cross term to be negative on every material and larger than half the smaller main effect. A result that held on four materials and reversed on two would be an accident of the examples; requiring all six makes the claim about the mechanism.

Where the model stops

Three limits, and the second is the one that would change the sign if it were wrong.

The gloss is added and not coupled. Real light reflecting at the interface also fails to enter the material, so a sample with a stronger interface has a slightly weaker body. That coupling is the Saunderson relation, it is second-order in the four per cent involved, and it would make the cancellation slightly larger rather than smaller.

The geometry collects the specular component. In a geometry that excludes it — 45°/0°, or a sphere with a gloss trap — the interface term is nearly absent, both departures act on the body alone, and the cancellation disappears. So the sign of the interaction depends on the measurement geometry, which is a condition a specification is supposed to name and often does not.

And the lobe is narrow. At a roughness of 0.35 the pedestal is smaller and the cancellation with it, so the table’s numbers are for a fairly glossy surface. The sign holds across the roughness ladder; the magnitude does not.

The generalisation

The rule worth taking away is about what an interaction term means, and it is a rule about model structure rather than about arithmetic.

Two errors add when they act on the same quantity in the same way. They interact when they act on different parts of it — and the sign of the interaction is decided by which parts. Here one departure scales the body and the other adds to the total, so the composition is not a sum but a sum of a scaled thing and an unscaled thing, and the cross term is the difference between those two operations.

The diagnostic is cheap and available before any computation: write down what each error acts on. If the two act on the same variable, adding is defensible. If one acts on a component and the other on the whole, adding is not, and the sign of the interaction is the sign of the difference between the component and the whole.

This collection has met the same structure elsewhere. A partial correction is worth exactly its fraction because the residual it acts on is a norm along a ray, so the operations compose linearly. Two tolerances do not meet in a tolerance because they are regions in different spaces, so intersecting them is not an operation on either. Both are the same question — what does this act on — asked of a different pair.

pale marble at eleven apertures, and at none. The same slab of pale marble, under D65, through the CIE 1931 2° observer, measured through apertures from one millimetre to forty and then with no aperture at all. Each patch is the colour that measurement returns; the number under it is how far that colour is from the model's own, in ΔE₀₀. The lightness falls as the aperture narrows, which is expected, and the chroma falls with it, which is less so — the bands that were reflecting most lose the most, because they are the bands whose light travels furthest before it comes back. Below 5.3 millimetres the hue is on the other side of neutral from the sample's own.
Fig. 5 The aperture’s half of the interaction, on the material where the cancellation is largest in absolute terms.
Eight conditions under which the model equation is exact, and how exact each one is. Each of the four departures vanishes if either of its two factors is empty, which is eight conditions. The axis is logarithmic in the residual that is left when the condition is imposed. Three of the eight are identities: the fluorophore's loading is zero so the emitted term is an empty sum, and a Lambertian surface or a uniform field makes the pairing's second argument identically zero. The other five are limits — a Gaussian excitation band has no edge, an opaque sample still has a kernel a few microns wide, a four-metre aperture is still finite, and the observer is small rather than absent at 380 nanometres. Each limit is drawn with the sequence its residual falls along as the condition is pushed, because a small number is not evidence of a limit and a falling sequence is.
Fig. 6 The conditions each of the two vanishes under. Neither is satisfied by a glossy translucent sample, which is why the pair has to be computed rather than assumed away.

So much for one sample at a time. The same two departures pushed through the collection’s own census — where every reading is a mean over a hundred and twenty-five surfaces — behave quite differently.

The collection's adaptation census, with its surfaces departed. Each row is one of the fourteen changes of light in this site's adaptation census, and the bar is what a von Kries gain leaves behind. The open marks are the published numbers; the filled ones are the same computation with every one of the hundred and twenty-five test surfaces replaced by what an instrument with an aperture, or a room with a direction in it, actually reports. Nothing moves by more than 9 per cent. A departure that does not depend on the light is very largely absorbed by the observer's own gain, because it changes the reflectance and the gain is applied afterwards. The fourth departure is not on this chart and cannot be: a fluorescent sample has a different curve under every light, so there is no set of reflectances to hand the census at all.
Fig. 7 What the two together do to the collection’s own census, which is much less than either does to a single reading.

What it means for the round’s own ladder

The interaction has one immediate consequence for how the four departures should be read together, and it is a caution rather than a result.

The round’s ladder puts four bars on one axis: 6.98, 6.70, 1.96 and 1.00. It is tempting to read a total off it, and the total is not available. Any two of those bars belong to a pair whose interaction has to be computed separately, and one of the six pairs is now known to cancel by more than the smaller of its members.

The honest statement about a sample carrying several departures is therefore: each bar bounds what that departure could contribute alone, and the sum bounds nothing. A sample with all four present might be worse than any one of them and is certainly not the sum.

That is an unsatisfying answer and it is the correct one. The alternative — quoting a total and marking it approximate — would be a number readers would use, and this collection’s standing position is that a number without the computation that produced it is worse than no number — which is why a mean is not a worst case is an essay here rather than a footnote. Where a computation is missing, the right output is the gap.

The same caution applies in the other direction. Somebody reading only the interaction table might conclude that departures generally cancel, which would be a worse error than adding them: three of the six pairs among these four have not been computed, one of them is expected to compound, and a rule inferred from one measured pair is a rule with a sample size of one.

Who found it, and when

Interactions between measurement errors are standard in metrology and are usually handled by declaring the contributions independent and adding in quadrature, which is the right thing to do when the contributions are noise. These are not noise: they are systematic and they have signs, so quadrature is the wrong operation twice over.

The specific pair here has been noticed obliquely by the industries that meet it. Plastics laboratories know that a glossy translucent sample measured with the specular included behaves differently from a matte one, and the guidance is to measure with the specular excluded — which, in the language of this round, is advice to remove one factor of the interaction rather than to compute it.

What does not appear to be written down is the size or the sign. Both follow from a two-line argument about which part of the reading each departure acts on, and neither requires the arithmetic here to see; the arithmetic supplies the numbers.

What the cancellation is worth as a caution

The result has a use beyond its own subject, and it is a caution about how systematic errors get combined in practice.

Two departures with known magnitudes and unknown signs cannot be combined at all. Adding them assumes both act the same way; adding in quadrature assumes they are independent, which for systematic terms is a stronger assumption than adding and is usually made because it gives a smaller number. Neither is available here: the two act on different parts of the reading, and the sign of their interaction follows from which parts rather than from their sizes.

The cheap diagnostic is worth restating because it takes one sentence and settles the question: write down what each error acts on. Same variable, same direction: add. Different components of one total: compute. Nothing else is reliable, and the temptation to combine anyway is strongest exactly when the pair has not been computed.

Where the ladder goes next

Six pairs are available among four departures and only one has been computed. The one most likely to be interesting is fluorescence with an aperture, because a fluorescent sample re-emits light inside the material, so the emitted light has its own kernel and its own diffusion length — probably longer than the reflected light’s, since the emission band is blue and blue scatters more.

If so, a narrow aperture would cut the fluorescent component harder than the reflected one, and a brightened translucent sheet would read less bright and less blue through a small hole. That is a two-departure prediction with an obvious test, and it is the sort of thing the round’s own instrument cannot reach without the two modules being made to talk to each other.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

ApertureBilinearityFalsificationMarginalisationMeasurement errorPredictionSpecularSubsurface scatteringToleranceTrade-off