What it takes to deliver it

A proof cannot be tuned for readers who disagree

A soft proof matched exactly for the standard observer is five colour differences wrong for one reader in twenty. Giving up that exactness to tune the display's three drives for a population instead moves the ninety-fifth percentile reader by 4 to 14 per cent even when the tuning is done on the very readers it is scored against — because what readers see is mostly each other's disagreement, and three drives act on every reader at once.

Assumes A soft proof is exact for one reader, Four primaries have a choice and The objective nobody chose.

A soft proof is a display driven so that its light and a printed patch produce the same three numbers for the standard observer. A soft proof is exact for one reader handed that match to two hundred observers drawn from a realistic population and found it several colour differences wrong for a large share of them. The obvious repair is to stop matching for the one observer nobody is, and to choose the display’s drives for the population instead.

A soft proof exact for one observer, as two hundred others see it. Each display is driven to match each of thirty printed patches exactly for the reference observer, so for that observer screen and print are the same colour to fourteen decimal places. The bars are what two hundred observers drawn from the population make of the same pairs: the median observer's difference, median over the patches, and the ninety-fifth percentile observer's: 1.8 and 4.8 on the wide-gamut LCD, 2.0 and 5.4 on the OLED, 3.1 and 7.5 on the laser projector. The narrower a display's primaries, the larger both become.
Fig. 1 The starting point: screen and print matched exactly for the standard observer on three displays, as two hundred other observers see them. The upper bar is the median observer and the lower the ninety-fifth percentile, median over the printed patches.

Tuned for readers, three drives move the tail by a few per cent

A three-primary soft proof re-tuned for a population serves that population only slightly better than an exact match for the standard observer, and the reason is structural: most of what readers see is their disagreement with each other, which no choice of three drives can reduce.

  • The part of the mismatch every reader shares is a median of 5.7 per cent of the squared mismatch on a wide-gamut LCD, 8.6 per cent on an OLED panel and 16.4 per cent on a laser projector. The rest is scatter between readers.
  • What they share is almost entirely their lenses being older than the standard observer’s. With every lens held at the standard observer’s age, the shared part falls to between 0.04 and 0.13 per cent.
  • Tuned on the very two hundred readers it is scored on — which no workflow could do — the ninety-fifth percentile reader moves from 5.12 to 4.86 on the OLED panel, from 4.57 to 4.39 on the LCD and from 7.54 to 6.50 on the projector.
  • Least squares over a population improves the median reader on most patches and makes the ninety-fifth percentile reader worse on most: on the OLED panel from 5.12 to 5.72.
  • Every tuned proof stops being exact for the standard observer, by a median of 0.67 to 1.88 colour differences, and the standard observer is how the proof is measured and approved.

Three drives act on every reader at once

A three-primary display has exactly three things to adjust for each patch: how hard to drive each emitter. Matching the standard observer spends all three, because that match is three equations in three unknowns with one answer. Giving the match up frees the three drives to do something else, and the question is what they can do.

What they cannot do is give different readers different light. Every reader sees the same spectrum leaving the screen, and changing a drive changes it for everyone at once. A drive change can move the whole population’s mismatch; it cannot pull the readers’ mismatches towards each other, except by accident of how differently each reader’s cones weight the change.

So the mismatch has to be taken apart first. Each reader’s mismatch with a patch is an offset in CIELAB: their reading of the screen minus their reading of the print. Over two hundred readers the offsets have a mean, which is the part they share, and a spread about it, which is the part they disagree about, and the squared offsets split exactly into the two, as a variance and a squared bias add. A tuning free to shift the display’s colour anywhere could at best remove the shared part.

The offset readers share, and the scatter they do not

What a soft proof's mismatch is made of, patch by patch, on an OLED panel. For each printed patch matched exactly for the reference observer on an OLED panel, the mismatch two hundred readers see, taken in CIELAB and split into the offset they share — the mean of their mismatches, which three drive levels could remove — and the root-mean-square scatter of each reader about that mean, which no drive levels can. Bars add the two lengths; the shares are of the squared mismatch, where the two add exactly. The shared offset is a median of 8.6 per cent of the squared mismatch and at most 33.4 per cent; the rest is readers disagreeing with each other.
Fig. 2 Every printed patch on the OLED panel, matched exactly for the standard observer, with the mismatch two hundred readers see taken apart. The dark end of each bar is the offset every reader shares; the long pale part is each reader’s scatter about it.

On the OLED panel the shared offset is under a tenth of the squared mismatch on the median patch, and it is never as much as half on any patch. On every one of the thirty patches the scatter between readers is larger than the offset they have in common. The largest shared part is on solid yellow, at a third of the squared mismatch: an offset of 1.17 against a scatter of 1.65. On unprinted paper, the most fragile patch in the set, it is a quarter — an offset of 3.52 against a scatter of 6.04. On solid cyan with forty per cent yellow it is two tenths of one per cent: an offset of 0.39 against a scatter of 7.86, which is a population whose members see the patch wrong in every direction and average to almost nothing.

That last case is the one to hold onto. Its ninety-fifth percentile reader is 7.5 colour differences from the print, and the mean of all two hundred readers’ offsets is 0.39. Any adjustment that moved the display’s colour enough to help the reader who sees it too green would hurt the reader who sees it too blue by the same amount. The population as a whole is almost exactly matched, and almost nobody in it is.

The part readers share is the age of their lenses

Why the shared part is small can be said before anything is computed, and saying it also says where to look for what is left.

If every reader’s departure from the standard observer changed their reading of a patch in proportion to the departure, and departures one way were as common as departures the other, the readers’ offsets would average to exactly zero and a tuning would have nothing to remove. For three of the four ways readers differ here that is nearly how the population is built: macular pigment, cone optical density and the three cone peaks are drawn symmetrically about the standard observer’s own values. The fourth is not. The standard observer’s lens is a thirty-two-year-old’s, which is the default age of the CIE’s physiological observer, and the readers are drawn from a working-age population between twenty and seventy, whose mean age here is 45. Every lens yellows in the same direction as it ages, so, as the observer has no age put it, the standard lens is not an average over readers but one moment in each reader’s life. On average these readers look through a yellower lens than the observer the proof was matched for, and that is an offset with a direction.

Holding variates still says whether that is the whole of it.

The offset readers share, with and without their lenses' ages. For each display, the offset readers share as a share of their squared mismatch, median over printed patches, in four cases: every reader as drawn, aged twenty to seventy; every lens held at the reference observer's age of 32; only the lens varied; and every reader as drawn but the proof matched exactly for an observer of the readers' mean age, 45. On a wide-gamut LCD the four shares are 6%, 0.11%, 42%, 0.35%, and the 95th percentile reader of the first and last is at 4.57 and 4.69. On an OLED panel the four shares are 9%, 0.04%, 42%, 0.32%, and the 95th percentile reader of the first and last is at 5.12 and 5.26. On a laser projector the four shares are 16%, 0.13%, 43%, 0.42%, and the 95th percentile reader of the first and last is at 7.54 and 6.90.
Fig. 3 The shared part of the readers’ mismatch, median over patches, with every reader as drawn; with every lens held at the standard observer’s age; with only the lens varied; and with every reader as drawn but the proof matched exactly for an observer of the readers’ own mean age.

With every lens held at the standard observer’s age and everything else varied as before, the shared part falls from 8.6 per cent to 0.04 per cent on the OLED panel, from 5.7 to 0.11 on the LCD, and from 16.4 to 0.13 on the projector. The shared offset itself falls from 1.79 to 0.12 on the OLED panel. With only the lens varied, the readers’ mismatch is smaller, and 42 per cent of what remains is shared — an offset of 1.69 on the OLED panel, 0.99 on the LCD and 2.99 on the projector, against 1.79, 1.04 and 3.03 with every variate free. The lens alone carries more than nine tenths of the shared offset on every display, and the other variates cancel to about a thousandth of the mismatch, which is what the symmetry argument predicts of them.

So the part a tuning could take is not a residue of curvature spread across the eye. It has one cause, and the cause is a choice: which age the observer the proof is matched for is taken to be.

What a soft proof's mismatch is made of, patch by patch, on a laser projector. For each printed patch matched exactly for the reference observer on a laser projector, the mismatch two hundred readers see, taken in CIELAB and split into the offset they share — the mean of their mismatches, which three drive levels could remove — and the root-mean-square scatter of each reader about that mean, which no drive levels can. Bars add the two lengths; the shares are of the squared mismatch, where the two add exactly. The shared offset is a median of 16.4 per cent of the squared mismatch and at most 26.6 per cent; the rest is readers disagreeing with each other.
Fig. 4 The same decomposition on a laser projector, whose emitters are two nanometres wide. The shared offsets are larger, largest on the solid magentas, and still smaller than the scatter on every patch.

On the laser projector the shared part rises to a median of 16.4 per cent, and the patches where it is largest are the magenta builds, with the black tint and paper close behind: solid magenta at 26.6 per cent, with an offset of 5.33 against a scatter of 8.87. A magenta is built on a display from its blue and red emitters, and an older lens takes its toll in the blue. The lens’s absorption falls steeply through the blue wavelengths; the print’s blue light is spread across that slope, while a two-nanometre emitter puts all of the display’s blue at one point on it, so a yellower lens takes a different fraction of the two. How much a lens costs depends on where a light puts its energy, which is what the lens is worst under tungsten found across lamps, and a laser’s blue is the most concentrated light in the set. Even so, on every one of the thirty patches the scatter is still the larger part.

Those Lab units are not the colour difference a reader is scored in. An offset of 5.33 on a saturated magenta is a chroma difference, and ΔE₀₀ counts chroma differences in saturated colours at well under their Lab length; solid magenta’s ninety-fifth percentile reader is at 5.34 in ΔE₀₀, with a Lab scatter of 8.87 behind it. The decomposition is done in CIELAB because that is the space in which a mean and a spread add exactly, and it says what share of the mismatch has a direction all readers agree on. What that share is worth to the ninety-fifth percentile reader has to be measured directly.

Four ways to choose the drives

A soft proof exact for one observer, and three proofs tuned for readers. For each display, the median over printed patches of the 95th percentile reader's mismatch between screen and print, for four ways of choosing the display's three drive levels: exact for the reference observer; least squares over a population of a hundred; tuned on that population's 95th percentile; and tuned on the two hundred readers it is scored on, which no workflow could do. On a wide-gamut LCD the four give 4.57, 4.99, 4.74, 4.39, and the three tuned proofs cost the reference observer 0.70, 0.82, 0.69. On an OLED panel the four give 5.12, 5.72, 4.99, 4.86, and the three tuned proofs cost the reference observer 1.35, 0.95, 0.67. On a laser projector the four give 7.54, 7.00, 6.81, 6.50, and the three tuned proofs cost the reference observer 1.88, 1.23, 1.14.
Fig. 5 The ninety-fifth percentile reader, median over patches, for four ways of choosing a display’s three drives: exact for the standard observer; least squares over a population of a hundred; tuned on that population’s ninety-fifth percentile; and tuned on the two hundred readers it is scored on. Each bar after the first also states what the tuning costs the standard observer.

Four proofs are built for every patch on every display, and all four are scored on the same two hundred readers, drawn with a different seed from anyone the tuning saw.

Exact is the proof as it is made now: three equations solved for the standard observer.

Least squares chooses the drives that minimise the summed squared mismatch over a training population of a hundred readers. It has a closed form, and it is the objective the objective nobody chose found inside every camera profile, with the same property: nobody decided the squared error was what mattered.

Tuned searches for the drives that bring the training population’s ninety-fifth percentile reader closest to the print — the reader a specification would name, and what a proofing workflow with a model of its readers could actually do.

Oracle does the same search on the two hundred readers it is scored on, which would need to know in advance who will look. It bounds what any choice of three drives can buy these readers.

The best tuning there could be buys under a seventh

The oracle is the number that settles the question, and it is small.

On the OLED panel the ninety-fifth percentile reader’s mismatch, median over thirty patches, moves from 5.12 for the exact proof to 4.86 for the oracle — 5 per cent. On the wide-gamut LCD it moves from 4.57 to 4.39, 4 per cent. On the laser projector, where the shared offsets were largest, it moves from 7.54 to 6.50, 14 per cent. On no display does the best three-drive proof the readers could have been given reach even a fifth better than the exact proof they were given. Single patches do better, and none does much better: the most improvable patch on each display gains between 18 and 21 per cent.

The tuning a workflow could actually perform does less. Tuned on a hundred other readers, the proof reaches 4.99 on the OLED panel, a gain of under 3 per cent, and 6.81 on the projector, 10 per cent. On the LCD it is worse than the exact proof, at 4.74 against 4.57: the drives that best served the hundred training readers’ tail were fitted to those hundred and did not transfer. It beats the exact proof at the ninety-fifth percentile on 13 of the 24 patches the LCD can reach, which is barely more than chance.

This matches what a mean is not a worst case found about adaptation, from the other side. There, a figure averaged over objects hid the object that fails. Here, a proof tuned to a population’s average cannot reach the readers in its tail, because what puts them in the tail is how they differ from everyone else, and the average is the one thing they do not differ from.

A different two hundred readers moves the number as much

Whether gains of this size mean anything has a free check. The exact proof was first scored on two hundred readers drawn with one seed, and is scored here on two hundred drawn with another. Its ninety-fifth percentile reader was 5.45 on the OLED panel and is 5.12; 4.80 on the LCD and is 4.57; 7.47 on the projector and is 7.54.

Two draws of the same readers differ by 0.33 on the OLED panel, and the best tuning there could be buys 0.26. The ninety-fifth percentile of two hundred is set by the ten readers beyond it, and which ten they are moves the number as much as the tuning does. On the projector the gain of a colour difference stands well clear of the 0.07 between samples, and there tuning is real. On the two broader displays a specification could not tell a tuned proof from an exact one by the readers it happened to have.

Least squares buys the median reader from the tail

The least-squares proof deserves its own look, because it does something none of the others do: it makes the ninety-fifth percentile worse on two of the three displays while making the median reader better on all three.

What least squares over a population does to the median reader and to the tail. Each dot is one printed patch on one display, proofed with drive levels that minimise the squared mismatch over a population of a hundred readers instead of matching the reference observer exactly. Across is the change in the median reader's difference and up the change in the 95th percentile reader's, both measured on two hundred other readers. On a wide-gamut LCD the median reader is better served on 17 of 24 patches and the 95th percentile reader worse served on 16. On an OLED panel the median reader is better served on 21 of 30 patches and the 95th percentile reader worse served on 21. On a laser projector the median reader is better served on 27 of 30 patches and the 95th percentile reader worse served on 16. Of the 84 dots, 41 sit in the shaded quarter, where the median reader gains and the 95th percentile reader loses.
Fig. 6 Each dot is one printed patch on one display, least squares over a population against the exact proof. Across is the change in the median reader’s mismatch, up the change in the ninety-fifth percentile reader’s. Nearly half the dots sit in the shaded quarter, where the median reader gains and the tail loses.

On every display least squares serves the median reader better on most patches and the ninety-fifth percentile reader worse on most: 21 and 21 of 30 on the OLED panel, 17 and 16 of 24 on the LCD, and 27 and 16 of 30 on the projector. Of the 84 patch-and-display pairs, 41 have both at once. At the median patch the median reader on the OLED panel improves from 2.01 to 1.90, and the ninety-fifth percentile reader worsens from 5.12 to 5.72.

The mechanism is the objective. Least squares removes the population’s mean offset, in the space it is solved in. When readers’ offsets are lopsided — a crowd a little off in one direction and a few readers far off in the other — the mean lies between them and nearer the crowd. Cancelling it moves the colour part of the way towards the crowd, which is where the median reader is, and the same distance further from the few, who are the tail. Least squares does not decide how much of the tail to give up for the middle; the square does, which is the same point the objective nobody chose made about a camera matrix, arriving at a reader.

Unprinted paper on the OLED panel shows the trade at its largest. Least squares moves its median reader from 5.81 to 4.35, a quarter better, and its ninety-fifth percentile reader from 11.04 to 11.43, slightly worse. A proofing workflow judged by how most people see the paper white would call that a clear improvement. One judged by how many readers reject the proof would call it a regression. Both are correct about the same drives.

An exact match for a forty-five-year-old makes the same trade

Since the shared part is the readers’ age, there is a proof that removes it with no search at all: match exactly, as before, but for an observer whose lens is forty-five years old, the readers’ mean. It is still three equations in three unknowns and still exact for a nameable observer, and it leaves the readers sharing 0.3 per cent of their squared mismatch on the OLED panel and the LCD and 0.4 per cent on the projector.

It does not rescue the tail. On the projector the ninety-fifth percentile reader improves from 7.54 to 6.90, most of what the tuned proof’s search found. On the OLED panel it worsens from 5.12 to 5.26, and on the LCD from 4.57 to 4.69. The arithmetic of ages says why. A proof matched for a thirty-two-year-old is twelve years wrong for the youngest reader and thirty-eight for the oldest; one matched for a forty-five-year-old is twenty-five years wrong for both. The extremes are evened out, and on the two panels, where most of the mismatch is the other variates’ scatter, evening them out buys the ninety-fifth percentile nothing and costs the young readers what it gives the old. The instrument, reading as the standard observer does, finds this proof wrong by a median of 1.01 colour differences on the OLED panel and 1.63 on the projector.

The approver pays for every reader served

Every tuned proof stops being exact for the standard observer, and that is not a technicality.

A soft proof is checked by measuring the screen and the print with an instrument, and the instrument reports through one fixed set of colour-matching functions: the instrument is one observer exactly. An exact proof measures as perfect. A tuned proof measures as wrong, by a median of 0.95 colour differences for the tuned proof on the OLED panel and 1.23 on the projector, and by up to 2.9 and 5.0 on their worst patches. Least squares departs further: a median of 1.35 on the OLED panel, 1.88 on the projector, and 7.0 on the projector’s worst patch.

So the person approving the proof, and the instrument verifying it, both see a proof that is off by about a colour difference, in exchange for a ninety-fifth percentile reader who is an eighth of a colour difference better off on the OLED panel, worse off on the LCD, and three quarters of one better off on the projector. That is a bad trade for whoever signs the proof, and invisible to the readers it was made for. A three-primary proof has one exact answer, and every other answer has to be defended against a measurement that says it is wrong.

The shared part does not follow where a proof fails

A tuning’s reach is not spread evenly across the patches, and it would help a proofing workflow if it reached furthest where proofs fail worst. It does not.

On the OLED panel the four corners of the table are all occupied. Solid yellow has the largest shared part, a third, and is among the most robust patches, with its ninety-fifth percentile reader at 1.69. Yellow with forty per cent magenta has almost none, two tenths of one per cent, and is just as robust, at 1.90. Paper has a quarter shared and is the most fragile patch on the panel, at 11.04; solid cyan with forty per cent yellow has two tenths of one per cent shared and is fragile too, at 7.50. Ranked by shared part and ranked by fragility, the patches are unrelated: the rank correlation is 0.07 on the OLED panel, −0.16 on the LCD and 0.04 on the projector.

How fragile a patch is depends on how large the readers’ offsets are; how tunable it is depends on whether they point the same way. Nothing connects the two, so a proof that could be tuned would be tuned on the patches that happen to be lopsided, not on the ones that need it.

Paper is where tuning pays most in absolute terms, because it is both the most fragile patch and a quarter shared. The oracle moves its ninety-fifth percentile reader from 11.04 to 10.35 on the OLED panel and from 16.00 to 13.86 on the projector. Both are the largest absolute gains on their displays, and both leave paper the worst patch, ahead of the stone grey at 9.45 and 11.86.

A fourth primary adds the freedom three do not have

The argument so far says what a three-primary proof cannot do. It also says what would change that.

With three primaries, matching the standard observer uses every degree of freedom, and tuning for readers can only be bought by giving up that match. Four primaries have a choice: matching three numbers with four drives leaves a one-parameter family of drive settings, every member exact for the standard observer, and the members differ in how different readers see them. A fourth primary lets a proof keep the approver’s measurement exact and still choose the member readers agree about, which is the trade this essay found impossible with three. A fourth primary is a design showed that choosing where to put that primary, against how far apart two hundred readers are about the display’s white, brings them two and a half times closer together.

That is not tuning in the sense used here. It changes the light the display can produce, so it changes which spectra readers are asked to compare with the print, and the scatter between readers is a property of how different those spectra are. The two other remedies a soft proof is exact for one reader named work the same way. Broader primaries make the display’s light more like the print’s, and a hard proof on the job’s own inks makes it identical. Every remedy that works on the scatter works on the spectrum; none of them works on the drives.

What was computed, and how

The displays, the thirty printed patches under D50 and the press model are those of the exact soft proof; the LCD cannot reach six patches with non-negative drives, and they are left out on that display. Readers are drawn from the same population model, a hundred for training on one seed and two hundred for scoring on another, and the held-variate populations reuse the scoring seed so that the held variate is the only change. Each reader’s readings are adapted from their own view of D50 by CAT16 and compared in ΔE₀₀.

Least squares solves the normal equations over the training readers’ white-normalised tristimulus values. The tuned and oracle proofs search from the exact drives by Nelder–Mead, since a percentile has no useful gradient, and refuse negative drives; a search of that kind can stop in a local minimum, which would loosen the oracle bound but not the decomposition’s cap. The shared offset and the scatter are the mean and root-mean-square deviation of each reader’s CIELAB offset.

What a panel of readers would settle

The population is a model of eyes. It varies what happens before the cones and in them, and nothing afterwards: readers also differ in how completely they adapt, how they weigh a hue difference against a lightness difference, and how large a difference they call a failure. Those add scatter, and more scatter strengthens the conclusion rather than threatening it, unless readers’ judgements happen to share an offset the eye model does not.

A panel of observers comparing an exact and a tuned soft proof against the same print, patch by patch on a broad-primary and a narrow-primary display, would say whether the tuned proof’s small gain on the projector is visible at all, and whether the readers who reject the exact proof are the same readers who reject the tuned one.

An allowance, not a correction

The tempting reading of the reader’s mismatch is that it is an error, and errors are for correcting. The decomposition says it is mostly not an error in that sense. An error has a direction, and most of this has none: it is the population’s spread, the same thing a tolerance is a probability found when two hundred people read one colour difference.

The move is to ask, before correcting anything aimed at a population, what share of the population’s mismatch points in one direction. That share is the most a single correction applied to everyone can remove, it is computable from the readers’ offsets without searching for a correction at all, and when it is found it usually has a cause that can be named — here, that the readers are older than the observer.

What cannot be corrected can be stated. A soft proof on three primaries is best left exact for the instrument that verifies it, with the population’s spread written beside it as an allowance, as a tolerance with an observer in it argued a delivery tolerance should carry — and the spread itself reduced, where it matters, by changing the light rather than the drives.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Colour managementIndividual variationLeast-squaresObserver metamerismOptimisationPaper whitePopulationPrimariesStandard observerTolerance