A mean is not a difference
Assumes Which of two is worse and A tolerance is a shape.
Nothing in industry is ever a pair of patches. A proof is checked against a chart of a few hundred; a profile is validated on a test set; a print run is measured on a control strip repeated across a sheet. In every case the measurement produces a set of colour differences and the specification wants one number.
Choosing which number is a decision about what matters, it is made in every standard, it is made differently in different standards, and it decides the answer about a third of the time.
Which of the two the parametric factors favour is the same question asked of one number rather than of a whole set, and it has the same answer.
The size of the tolerance is the other free argument in the same comparison, and the ambiguity it produces does not shrink as the tolerance grows.
A tight tolerance is the case a reader would expect to be the safest of the three, and it turns out to be the one where the two judgements disagree with each other most.
The claim
Reducing a set of differences to one number is a modelling decision, and the two decisions in common use rank reproductions differently about a third of the time.
- Twenty-nine and a half per cent of candidate pairs are ranked one way by the mean and the other way by the ninety-fifth percentile.
- Neither statistic is wrong. A reproduction that is slightly off everywhere and a reproduction that is exact nearly everywhere and badly wrong in a corner are genuinely different objects, and there is no fact of the matter about which is better without saying what the reproduction is for.
- And the choice is usually made by whoever wrote the specification’s template, rather than by anybody thinking about the product — which is how a standard ends up preferring the reproduction its users would reject.
The two shapes of error
Real reproduction error has a characteristic distribution and it is not symmetric. A profile, a proof or a press is accurate over most of its range and inaccurate in specific places: near the gamut boundary, in the deep shadows, on the saturated primaries, wherever the mapping had to compromise.
So the error set looks like a bulk plus a tail. Most samples sit at some small value that reflects the general fidelity of the process; a minority sit much higher and reflect the places the process gave up.
Two candidates can differ in either part independently. One can have a better bulk and a worse tail — a process that is slightly sloppy everywhere but never catastrophic. The other can have a tighter bulk and a longer tail — a process that is nearly exact where it works and fails hard at the edges.
The mean is dominated by the bulk because the bulk is most of the samples. A high percentile is dominated by the tail because that is what a percentile is for. The two statistics are therefore reading two different halves of the same distribution, and when the halves disagree, the statistics do.
What was computed, and how
Two thousand pairs of candidates. Each candidate is a set of two hundred errors built from three drawn parameters: a bulk level, the share of samples that fall in the tail, and the tail’s size. Every value is then multiplied by a uniform factor to give within-set scatter, so no candidate is a step function.
The parameter ranges are chosen so that both candidates in a pair are plausible — a bulk between about 0.1 and 1.8 units, a tail affecting two to fifteen per cent of samples at one to nine units. Those are the shapes real reproduction error takes, and the point of the exercise is how often two ordinary candidates invert rather than whether a contrived pair can be made to.
Each candidate is then summarised twice, by its mean and by its ninety-fifth percentile, and the pair counted as inverted when the two statistics prefer different candidates.
The generator is seeded, so the 29.5 per cent is reproducible rather than sampled — and the number is stable across the seeds tried, which matters because a rate computed from a random construction is only worth quoting if it does not move when the randomness does.
And an empty set throws. setStatistics refuses to summarise nothing, which is not pedantry: a measurement run that produced no readings and a measurement run whose readings were all zero are different situations, and a mean of NaN silently propagating through a report is the kind of thing that gets discovered a year later.
What the standards actually say
Most specifications quote more than one statistic, which sounds like it should dissolve the problem and does not.
A typical print standard names a mean, a maximum and sometimes a ninety-fifth percentile, with a different limit on each — and a limit is a statement about a probability rather than about a sheet. That is better than one number, and it converts the ranking question into a conjunction: a candidate passes if it satisfies all of them. Conjunctions do not rank. Two candidates that both pass are equally acceptable to the specification, and two that both fail are equally unacceptable, so the moment anybody has to choose between candidates — which supplier, which profile, which set of press conditions — the ranking question comes straight back and is answered by whichever number the person happens to look at first.
A profile validation report is worse, because it usually leads with a mean. The mean is the number that goes in the summary, gets compared between profiles, and decides purchases; the maximum is in a table further down.
And an image difference is worse again, because the “set” is every pixel. A mean over an image is dominated by the sky, the wall and the background — which are the parts of the picture nobody was looking at. A high percentile is dominated by the few hundred pixels in the face, which is the only part anybody would have noticed. That is not a subtle statistical point; it is the whole reason image difference metrics exist and why none of them is a mean — the same reasoning that makes a mean and a worst case two different specifications.
The three-number table does not rank, and here is why that matters
The objection to all of this is that nobody uses one number: a serious specification quotes a mean, a maximum and often a percentile, each with its own limit. That is a real improvement and it solves a different problem.
A conjunction of limits partitions candidates into acceptable and not. It is exactly the right instrument for the question “may this sheet leave the factory”, which is a question about one object. It is no instrument at all for “which of these two profiles should the studio standardise on”, “which of these three suppliers gets the contract” or “has the new press condition improved on the old one” — and those are the questions that determine what gets built.
The moment a comparison is needed, somebody picks a number out of the table. Which number they pick is not written anywhere, is usually the first one, and the first one is almost always the mean, because that is how tables are ordered. So a document that carefully avoids reducing quality to a single statistic ends up reducing it to a single statistic whenever it is used for a decision, and the statistic is chosen by typography.
Where the model stops
The candidates are generated, not measured. Their shape is chosen to resemble real reproduction error and no real reproduction was involved. A rate computed against a corpus of measured profiles would be the number worth quoting industrially, and it would need a corpus this site has not got.
The percentile is one choice among many. Ninety-fifth is conventional; ninetieth and ninety-ninth give different rates, and the maximum — which several standards use, and which is a search rather than a statistic — gives a rate dominated by a single sample and is therefore the least stable of all the options. Nothing here argues that ninety-fifth is right, only that it is different from the mean.
And the errors are treated as exchangeable. A real set has structure: the two samples nearest the gamut boundary are both wrong for the same reason, so the tail is correlated rather than a random minority. Correlated tails make inversions more likely rather than less, so 29.5 per cent is a floor.
Nothing here weights the samples. Every element of the set counts once, which is what a chart-based validation does and what no viewer does. A test chart’s samples are laid out to cover the space evenly, so a colour that occupies one per cent of the chart and eighty per cent of the image contributes one per cent of the mean. That is a deliberate property of a chart — it is trying to characterise a process, not an image — and it becomes a defect the moment the resulting number is used to predict what somebody will think of a picture.
And the two statistics are not the only pair that disagree. Mean against maximum inverts more often than mean against ninety-fifth, because a maximum is one sample; median against mean inverts on any set with a skew, which is every reproduction error set there is. The essay measures one pair because it is the pair the standards actually use, and the general statement is that any two summaries of a skewed distribution can be made to disagree — what makes this one worth a measurement is that both summaries are in the same table, on the same page, describing the same run.
The rate for the other pairs, which the essay names and does not measure
The closing caveats say that mean against maximum inverts more often than mean against the ninety-fifth percentile, and that median against mean inverts on any skewed set. Both are stated as reasoning and neither is given a number, and the generator described above supplies them for the cost of running it four more times.
Rebuilding the construction exactly as specified — two hundred errors per candidate, a bulk between 0.1 and 1.8, a tail affecting two to fifteen per cent of samples at one to nine units, each value multiplied by a uniform scatter factor — reproduces the published figure at 29.7 per cent against 29.5, which is close enough to treat the reconstruction as the same experiment. Against that baseline:
| statistic compared with the mean | pairs it ranks the other way |
|---|---|
| median | 15.1% |
| 90th percentile | 19.9% |
| 95th percentile | 29.7% |
| 99th percentile | 35.7% |
| maximum | 37.0% |
The maximum inverts most, as the essay predicts, and the margin over the ninety-fifth is seven points rather than the large gap the reasoning suggests — a maximum is one sample, but on two hundred samples with a tail affecting up to thirty of them, the maximum and the ninety-fifth percentile are both reading the same tail and mostly agree with each other.
The median is the mildest of the four, not the ubiquitous disagreement the caveat implies. It inverts on one pair in seven, less than half as often as the ninety-fifth percentile does. That a skewed distribution has a median different from its mean is certainly true of every set here; that the difference is large enough to reverse a ranking is true one time in seven, and those are different claims.
One and a half units sits between the two tolerances most often quoted, and the flip rate there says whether the effect interpolates or jumps.
Which percentile is nearly as large a choice as whether to use one
The more consequential reading of that table is the middle three rows, because they are all the same kind of statistic.
Moving from the ninetieth percentile to the ninety-ninth takes the disagreement with the mean from 19.9 per cent to 35.7 — nearly doubling it, without changing the kind of summary at all. The choice of percentile is worth sixteen points of inversion rate, against the twenty-nine points that separate the mean from the middle of that range.
So a specification that names a percentile without naming which has not narrowed things nearly as much as it looks. The essay’s argument is that the mean-or-percentile decision is made by whoever wrote the template; the same is true one level down, and the second decision carries more than half the weight of the first. A document quoting “the 90th percentile” and one quoting “the 99th” are about as far apart as a document quoting a percentile and one quoting a mean.
Four units is where a specification has stopped being tight, and the rate there is the check on whether the whole effect is an artefact of tightness.
And the rate is stable against the one parameter the essay does not state
The construction is given in full except for how wide the within-set scatter factor is, which matters because a wider scatter blurs the bulk into the tail and could plausibly decide the answer.
It does not. Sweeping the scatter from ±10 per cent to ±70 per cent moves the ninety-fifth-percentile rate from 32.2 to 28.0 — four points across a sevenfold change in the one unstated parameter, with the published 29.5 sitting inside that band throughout.
That is worth recording beside the note about seeds. A rate computed from a random construction is worth quoting if it survives the randomness; it is worth relying on if it also survives the construction’s own free parameters, and this one does. The direction is mildly informative too: more scatter means slightly fewer inversions, because blurring the two parts of the distribution together makes the two statistics read more nearly the same thing.
The generalisation
The sentence worth carrying is: a summary statistic is a claim about what the reader will look at, and colour reproduction is judged by its worst region.
That is the asymmetry the mean cannot express. A viewer looking at a print does not integrate its accuracy; they find the part that is wrong. Everything about human vision that this site has measured points the same way — the eye is a difference detector, an edge carries the visible error in a gradient, a screen is judged by its loudest harmonic — and every one of those is a maximum rather than an average.
So there is an argument, stronger than a preference, for the tail being the right statistic in colour reproduction: it is the one that matches what the detector on the other end does. A mean is the statistic of a system that integrates, and nothing about a person looking at a printed page integrates.
The surprising connection is with dither. Dithering deliberately increases the mean error in order to decrease the visible one, and the phase before this one found that reading the result at a point rather than as components changes the answer by a factor of twenty. Both findings are the same statement: the number that predicts what somebody sees is not the number that summarises what the machine did.
Who found it, and when
The distinction is older than colour science and belongs to statistics, where the difference between a measure of central tendency and a measure of the tail has been understood since the nineteenth century. Nothing here is a discovery about statistics.
What is specific to colour is which one got adopted, and when. Early process control quoted a mean because a mean is what a small number of hand-measured patches supports: with six readings, a percentile is not a statistic. The convention outlived the constraint — a modern instrument reads a chart of hundreds in a minute — and by then it was in the templates.
Image difference metrics went the other way from the start, because their authors were explicitly modelling a viewer. S-CIELAB in 1997 filtered the image through the eye’s spatial response before taking any difference at all, which is the same instinct as taking a percentile and better executed: rather than choosing which part of the distribution matters, it changes the distribution so that the parts nobody can see are not in it.
And the standards’ three-number tables are the compromise that resulted, adopted through the 2000s, which is a genuine improvement over any single number and does not answer the question this essay is about — because a conjunction of limits still cannot rank.
What the pictures cannot show
They cannot show the picture. The whole argument is that a set of differences is judged by where in the image the large ones fall, and a bar chart of statistics has no image in it. The right figure would be two reproductions side by side, and this site cannot show a reproduction: every patch on this page is already a display, and the errors under discussion are of the size that a display’s own calibration would swamp.
And they cannot show a distribution the reader can check. The candidates are generated, so the histogram of any one of them is a fact about a random draw. What is stable across draws is the rate of inversion, which is a number about the procedure rather than about any picture that could be printed.
Where the ladder goes next
The nearest unfinished piece is the correlated tail. Real reproduction errors cluster in colour space, and clustering makes inversion more likely — so a measurement using a spatially structured error model rather than an exchangeable one would put a floor under the 29.5 per cent and say by how much.
The second is the weighted statistic. Between a mean and a percentile there is a family of options nobody uses: an error weighted by how much of the image sits at that colour, which is what an image difference metric approximates and what a chart-based validation could do directly if the chart carried occupancy weights. The arithmetic is trivial and the data is in every image.
And the third is the one that connects this to the rest of the phase: every statistic here is computed for one observer. A set of differences read by a population of observers has a spread per sample, and reducing a set of distributions to one number is a harder problem than reducing a set of numbers — with two axes to choose a statistic along instead of one.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- The booth is a luminaire acceptability · colour management · δe · measurement error · quality control · specification · tolerance
- Three constants nobody quotes acceptability · ciede2000 · δe · measurement error · quality control · specification · tolerance
- A difference is not a distance ciede2000 · δe · gamut mapping · quality control · specification · tolerance
- A tolerance needs a second number acceptability · ciede2000 · δe · quality control · specification · tolerance
- One unit in another room acceptability · ciede2000 · δe · quality control · specification · tolerance
- A brand colour for a population acceptability · colour management · quality control · specification · tolerance
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AcceptabilityCIEDE2000Colour managementΔEGamut mappingThe ICC profileImage differenceInterpolationMeasurement errorQuality controlSpecificationTolerance