Matching and measuring

Five nanometres is a choice

Every integral here is taken in five-nanometre steps, and the interval has never had to be defended. Coarsening it to twenty costs daylight two hundredths of a colour difference and a fluorescent tube six and a half — and which way of coarsening is used decides a further factor of six.

Assumes What the instrument reports and A spectrum is not a colour.

Every number on this site is an integral, and every integral is a sum over eighty-one wavelengths from 380 to 780 nanometres in steps of five. That grid has been in the spectral machinery since the first commit and has never been asked to justify itself.

It should be. A wavelength grid is not an implementation detail of an integral — it is a claim about what the light has in it, and there is a source in every building for which the claim is false.

What a coarse wavelength grid costs, by source. Colour error against grid size, for four sources through one reflectance. Daylight survives every grid tested: 0.67 ΔE00 even at 40 nm. A source with lines in it does not — the narrowband source reaches 16.2. The grid is not a property of the arithmetic; it is a claim about what the light has in it.
Fig. 1 Colour error against grid size, for four sources through one reflectance, measured as ΔE00 against the answer on the full 5 nm grid. Daylight survives every grid tested — 0.67 even at 40 nm. A source with lines in it does not: the narrowband source reaches 16.2, which is not an error anybody would attribute to a choice of summation interval.

The claim

How finely a spectrum is sampled is part of the answer, and the size of the error is set by the source rather than by the arithmetic. Two consequences, both measured below:

  • On a smooth spectrum a coarse grid is nearly free. Coarsening daylight from 5 nm to 20 costs 0.17 ΔE00 and to 40 costs 0.67.
  • On a structured one it is not. The same coarsening costs a fluorescent tube 6.56 and a narrowband source 3.05, rising to 8.03 and 16.24 at 40 nm.

And a third, which is the practical half: there are two ways to coarsen and they are not the same operation. Sampling every nth wavelength is what a narrow slit at coarse steps does; averaging each band is what a wide slit does. On a tube at 20 nm they differ by a factor of six and a half.

Two ways to lose four fifths of the data

Sampling takes the value at 380, 400, 420 … and discards everything between. If a mercury line sits at 546 nm and the grid lands on 540 and 560, the line is simply not in the data — and the resulting colour is the colour of a lamp that does not have that line.

Averaging integrates each interval, so the line’s power is retained, redistributed across 20 nm of the spectrum. The colour is then the colour of a lamp whose line has been smeared, which is wrong in a much gentler way: the total is preserved and only the placement is lost.

That is the difference between an instrument with a narrow slit stepping coarsely and one whose bandpass matches its step size, and it is the reason the standards for measuring light specify both.

A fluorescent source at 5 nm and at 20 nm, sampled. The fine curve is this site's own grid and the stepped one is the same source read at 20 nm. Sampling takes the value at every fourth wavelength and discards the rest, which is what a narrow slit at coarse steps does. Through a reflectance the difference comes out at ΔE00 6.56 — a colour error caused by nothing but the choice of where to look.
Fig. 2 A triphosphor tube at 5 nm and the same tube sampled at 20. The mercury lines are the argument: at this interval the grid lands where it lands, and a line that falls between two grid points contributes nothing at all to any integral taken on it.
A fluorescent source at 5 nm and at 20 nm, averaged. The fine curve is this site's own grid and the stepped one is the same source read at 20 nm. Averaging integrates each band, which is what a wide slit does. Through a reflectance the difference comes out at ΔE00 1.00 — a colour error caused by nothing but the choice of where to look.
Fig. 3 The same tube averaged rather than sampled at the same interval. Every line’s power survives, spread across the band it fell in. The curve looks less like the original and the colour computed from it is far closer — which is the whole finding in one pair of pictures.

The measurements

Four sources through one broad red reflectance, integrated against the 1931 2° observer, each compared with its own answer on the 5 nm grid:

source 10 nm 20 nm 40 nm
D65 — sampled 0.09 0.17 0.67
D65 — averaged 0.00 0.03 0.13
a fluorescent tube — sampled 2.44 6.56 8.03
a fluorescent tube — averaged 0.07 1.00 9.85
a white LED — sampled 0.39 0.94 1.55
a white LED — averaged 0.04 0.23 1.19
a narrowband source — sampled 0.79 3.05 16.24
a narrowband source — averaged 0.26 1.21 5.46

Three things in that table are worth stating separately.

Averaging wins wherever there is structure, and by a large factor: 0.07 against 2.44 for the tube at 10 nm, a factor of thirty-five.

Daylight is nearly free at any interval tested, which is the control that makes the rest meaningful — a coarsening that damaged everything equally would be a broken coarsening rather than a finding about spectra.

And at 40 nm the ordering flips for the tube: averaged 9.85 against sampled 8.03. At that interval the bands are wider than the structure they are meant to preserve, so averaging is no longer smoothing a line, it is inventing a broad emission where a narrow one was — and the sampled version’s error becomes a lottery that happens to have paid off. Neither number is respectable; the crossover is the signal that the grid has stopped being a measurement at all.

Sampling against averaging, on the same coarse grids. The solid lines are sampled and the faint ones averaged, for four sources at three grid sizes. Averaging wins wherever the source has structure — a fluorescent tube costs 6.56 sampled against 1.00 averaged at 20 nm — and the two agree on daylight, which has nothing narrower than the grid in it either way.
Fig. 4 Sampled against averaged for all four sources. The dashed curves are averaged. The gap between the pairs is the difference between two instrument designs, measured in the unit the answer is quoted in.

Why the standards say what they say

The colorimetric standards handle this by publishing weighting tables: the colour-matching functions and a standard illuminant pre-multiplied and pre-integrated over each band, at 10 nm and at 20, in two versions — one for instruments that sample and one for instruments whose bandpass matches the interval.

Those tables look like a convenience for people without computers. They are not; they are the correction this essay measures. A table built by integrating over the band and a table built by sampling it differ by exactly the amounts above, and using the wrong one on a structured source is an error of several ΔE00 in a procedure everyone treats as exact arithmetic.

The measured half of the same problem is already on this site: a spectrophotometer sees a sample through a slit of finite width and writes down the truth convolved with its own bandpass, and nothing in the file says which values were measured and which were smoothed.

A 20 nm bandpass, and what it does to the sampleThe true reflectance, the same reflectance as a 20 nm instrument reporting every 20 nm returns it, and the slit function responsible, drawn at its own scale around 550 nm. The reconstruction is what every calculation downstream will use, and it differs from the truth by ΔE 1.09 under D65. Nothing in the file the instrument writes marks which values were measured and which were interpolated.20 nm slit400450500550600650700wavelength / nmgrey: the truthgold: what is reported21 samplesΔE 1.09under D65the interpolated values are not markedCIE 1931 2° observer
Fig. 5 The instrument’s half of the argument. A finite slit convolves the sample with its own response before the arithmetic starts, so a coarse grid and a wide slit are two ways of losing the same information — and an instrument that does both loses it twice.

What a coarse grid does to a match

The sharpest demonstration is not a lamp at all. It is a metameric pair, and it shows that a grid can manufacture a colour difference out of nothing.

Take two reflectances built to match exactly under D65 — a base and the same base plus a metameric black, which integrates to zero against all three colour-matching functions. On the full grid they differ by 2.6 × 10⁻¹³ ΔE00, which is the arithmetic’s floor. Now read the same two samples on a coarser grid:

grid sampled averaged
10 nm 3.59 0.24
20 nm 9.52 1.00
40 nm 18.40 4.02

At 20 nm, sampling reports two identical colours as nine and a half units apart. Nothing is wrong with the samples, the illuminant, the observer or the formula; the grid alone did it.

And the three rows have an order in them, which turns the comparison from a demonstration into a rule.

Doubling the grid multiplies the averaged error by 4.17 and then by 4.02. That is h2h^2 to within four per cent, twice — the second-order convergence a box rule has, arriving on the one test in this essay whose exact answer is known to be zero.

Doubling it multiplies the sampled error by 2.65 and then by 1.93. That is h1h^1, erratically.

10 → 20 nm 20 → 40 nm implied order
sampled ×2.65 ×1.93 ≈ 1
averaged ×4.17 ×4.02 ≈ 2

So the two coarsenings are not two grades of the same operation; they converge at different rates, and everything else in this essay follows from that one exponent. The advantage of averaging is the ratio of the two, which goes as 1/h1/h — measured at 15.0, 9.5 and 4.6 across the three grids, against a predicted halving at each step. A finer grid does not merely reduce both errors; it widens the gap between the two instrument designs.

That has a practical edge the table of four sources does not show as clearly, because its numbers are noisier. A laboratory choosing between the two designs should choose by how fine its grid is, and the finer the grid the more the choice matters. At 40 nm the two are within a factor of five and both are useless; at 5 nm the sampled instrument is behind by a factor that has grown eightfold, on a measurement everybody now regards as exact.

And the noise in the sampled column is the phase effect, showing up as an inconsistent exponent. Across the four sources the averaged errors rise by 4.3 to 5.2 per doubling — a tight band around the theoretical 4 — while the sampled ones rise by anywhere from 1.2 to 5.3. A convergence order that will not hold still is the signature of an error term that depends on where the grid landed rather than on how coarse it is, which is exactly the limitation this essay records at the end and declines to compute.

So the two halves fit together. Averaging has a genuine convergence order, so its error can be predicted from the interval alone and a table of it means something. Sampling does not, so its error is a property of one alignment and quoting it as a number is quoting one draw from a distribution. The reason the standards publish two sets of weighting tables is not that one instrument is worse; it is that only one of the two has an error that a table can describe.

The reason is exact rather than statistical. A metameric black is, by construction, a spectrum that oscillates — it has to, since it must integrate to zero against three smooth non-negative functions and still be non-zero. Its oscillation is precisely the structure a coarse grid cannot represent, so the one kind of sample guaranteed to defeat a coarse grid is the kind quality control most needs to measure.

What this site’s own grid can and cannot carry

Five nanometres is fine for everything on this site, and it is fine because of what is on this site rather than because it is fine in general.

The sources here are constructed: mercury lines are added at the wavelengths mercury emits at, with a width chosen to be resolvable on the grid. So the grid is not measuring those lamps — it is defining them, and their line widths are an instrumental width rather than a physical one. That is a real limitation and it was recorded when the machinery was built: any line narrower than 5 nm cannot exist here at all.

The consequence for the numbers above is worth being exact about. The narrowband source’s 16.24 ΔE00 at 40 nm is a fact about a source whose lines are 5 nm wide by construction. A real laser line, or a real low-pressure sodium lamp, is far narrower, and the error on a coarse grid would be larger still — so the table is a floor rather than an estimate, in the same way the fluorescence numbers here are a floor because the ultraviolet is outside the range.

What was computed, and how

The coarsening replaces each block of the fine grid with a single value — the block’s first sample, or its mean — and holds it across the block, so that the coarsened spectrum is defined on the same 81 wavelengths as the original. That keeps the downstream integral identical in every respect but the data, which is the only way the comparison is a comparison of grids rather than of code paths.

The colour is computed the same way for every row: the coarsened source multiplied by the reflectance, integrated against the 1931 2° functions, converted to CIELAB against the source’s own white point — so a coarsening that shifts the white point does not show up twice — and differenced with CIEDE2000.

The reflectance is a smooth broad bump centred at 600 nm on a pedestal, chosen because it has no structure of its own: a structured reflectance would add a second sampling problem and confuse which one the table is about.

And the guards. Two assertions, in opposite directions. One requires daylight at 20 nm to cost less than half a unit, so a broken coarsening that damaged everything would be caught. The other requires the narrowband source at 20 nm to cost more than a unit, so a coarsening that quietly did nothing would be caught too. A third requires averaging to beat sampling on at least two of the three structured sources, which is what makes the instrument-design half of the argument a measurement rather than an assertion.

A narrowband source averaged rather than sampled is the combination that separates the grid from what the instrument does with it.

A narrowband source at 5 nm and at 20 nm, averaged. The fine curve is this site's own grid and the stepped one is the same source read at 20 nm. Averaging integrates each band, which is what a wide slit does. Through a reflectance the difference comes out at ΔE00 1.21 — a colour error caused by nothing but the choice of where to look.
Fig. 6 The fine curve is this site’s own grid and the stepped one the same source read at 20 nm by averaging each band, which is what a wide slit does. Through a reflectance the difference comes out at ΔE00 1.21.

Where the model stops

The reflectance is smooth. Real samples are not: an interference pigment has structure narrower than any practical grid, and this site has already measured what that does to an instrument. A structured reflectance under a structured source is the case where the two sampling problems multiply, and nothing here computes it.

There is no ultraviolet and no infrared. The range is 380–780 nm throughout, so a coarsening’s effect on a fluorescing sample cannot be seen here at all.

Only one observer. The 1931 2° functions are themselves smooth, so the grid interacts almost entirely with the source. Under a sharper kernel the numbers would be larger, and no such kernel is standardised.

And the phase of the grid is arbitrary. Sampling at 380, 400, 420 lands differently from sampling at 385, 405, 425, and for a line source the difference can be the whole error. The table reports one phase; a full treatment would report the distribution over phases, which is a more honest object and a much harder one to draw.

A forty-nanometre step is coarser than any standard allows and is what makes the sampling failure unmistakable.

A fluorescent source at 5 nm and at 40 nm, sampled. The fine curve is this site's own grid and the stepped one is the same source read at 40 nm. Sampling takes the value at every fourth wavelength and discards the rest, which is what a narrow slit at coarse steps does. Through a reflectance the difference comes out at ΔE00 8.03 — a colour error caused by nothing but the choice of where to look.
Fig. 7 The same fluorescent source read at every fortieth wavelength, discarding the rest, which is what a narrow slit at coarse steps does. Whether a mercury line is caught or missed is decided by where the grid happens to fall.

The generalisation

The pattern is one this collection keeps arriving at from different subjects: a discretisation is a modelling assumption wearing the clothes of an implementation detail.

The tell is always the same. The grid is chosen once, early, by whoever wrote the first integral; it is invisible in every result; and it is exactly right until an input arrives whose structure is finer than the grid — at which point the arithmetic does not fail, it returns a confident wrong number.

Three neighbours of that shape sit close by. A profile’s lattice is a grid in colour rather than in wavelength, and its error is set by where its nodes are rather than by how many. A halftone screen is a grid in space, and it works only because the eye’s own grid is coarser. And the reader’s display is a grid too — a set of code values, whose spacing decides whether a gradient is a gradient.

One more neighbour is worth naming because it is the same arithmetic in the eye rather than in the instrument: a cone’s sensitivity is an integral too, and the smoothness of the colour-matching functions is what makes a 5 nm grid adequate for the observer while being inadequate for the source. If the kernel were as spiky as the light, no practical grid would work at all.

A narrowband source at 5 nm and at 40 nm, sampled. The fine curve is this site's own grid and the stepped one is the same source read at 40 nm. Sampling takes the value at every fourth wavelength and discards the rest, which is what a narrow slit at coarse steps does. Through a reflectance the difference comes out at ΔE00 16.24 — a colour error caused by nothing but the choice of where to look.
Fig. 8 The worst case in the table, drawn. Three narrow bands read at 40 nm intervals: two of the three land between grid points and are simply absent from the data, which is how a colour error of sixteen units is produced by an arithmetic in which nothing went wrong.

There is a second rule in the convergence orders, and it transfers further than the first. An error with a stable exponent can be extrapolated away and an error without one cannot. A box rule at 20 nm and at 10 nm gives two numbers whose difference is four thirds of the finer one’s error, so the limit is reachable from two coarse measurements — the same Richardson step used here to extrapolate a voxel count to a volume. A point-sampled measurement offers nothing of the kind, because the thing that varies between its two grids is not only their spacing.

That is worth knowing before deciding to measure something more finely. The question to ask of a discretisation is not how small the error is but whether it has an order, because an error of known order at a coarse setting is worth more than a smaller error of unknown order at a fine one.

The rule that generalises: state the grid beside the answer. A colour computed from a spectrum is not fully specified by an observer, an illuminant and a space; it also needs the interval it was summed on and how the interval was formed, and the two together can be worth more than the difference between two colour spaces.

Who found it, and when

The problem is as old as instrumental colorimetry. The CIE published weighting tables for 10 nm and 20 nm intervals precisely because early instruments reported at those intervals and the naive sum was known to be wrong for discharge lamps.

ASTM’s tables, which are what most industrial software actually uses, exist in two forms for the two instrument designs, and the standard says in as many words that using the wrong one is an error of the size measured here. The distinction between them is the difference between an instrument with a bandpass matched to its interval and one without — the same distinction this essay’s two coarsenings model.

The deeper statement belongs to sampling theory rather than to colorimetry, and it is Nyquist’s: a sampled function is only recoverable if it was band-limited below half the sampling rate. A daylight spectrum very nearly is, which is why the whole business works. A mercury line is emphatically not, which is why it does not.

Where this goes next

The obvious extension is the one this site cannot afford: widening the range. The 380–780 nm window is load-bearing in two places already — it truncates an optical brightener’s excitation band and it sets the narrowest representable line — and widening it touches the observer functions, every illuminant, every metamer construction and every figure downstream. That is a phase’s worth of work and should be decided rather than drifted into.

The nearer piece of work is the phase of the grid. Every number in this essay is for one alignment, and the variation over alignments is the honest error bar on a coarse-grid measurement of a line source. Computing it needs nothing new — the machinery already coarsens at any offset — and it would turn a table of point estimates into a table of ranges, which is what the quantity actually is.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 18 that link here.

The objects this essay names

Each one links to every other essay that touches it.

BandpassColorimetryΔEIntegrationMeasurement errorQuality controlSamplingSpecificationSpectrophotometryWavelength grid