Matching and measuring

A narrow primary buys a disagreement

The observer audit decomposes what a display costs a population. Narrowing the primaries raises the pigment-peak departure monotonically, moving one raises or lowers the macular departure, and the two respond to different design variables — so a wide gamut and an observer-robust display are bought with the same money.

Assumes One wavelength is everyone's colour, A gamut has a population and A primary is chosen for four things.

This collection already knows that narrow primaries cost a population something. What the observer audit adds is which of six terms each display design variable moves, and the terms do not move together.

The same white, matched at six primary widths. At every width the three primaries are solved to match D65 exactly for the reference member; the bands are what the population sees. A broad primary integrates the observer differences over a band and averages them away; a narrow one samples them at a point and passes them straight through. From 40 nm to 2 the ninety-fifth percentile rises from 11.3 to 17.9 ΔE00, monotonically, and the technology has been moving from left to right for thirty years.
Fig. 1 What a population makes of a match as the primaries narrow. The spread rises monotonically, and this round decomposes it.

The claim

A display’s observer cost decomposes, and its two largest terms respond to two different design variables — bandwidth and position — which are not traded against each other in any datasheet.

  • Narrowing a primary raises the pigment-peak departure monotonically: 2.37 ΔE₀₀ under daylight, 2.84 under a three-emitter LED, 3.60 under a three-laser projector.
  • Positioning a primary decides the macular departure: a blue pump at 452 nanometres sits on the macular absorption band and gives 3.37; daylight gives 1.71.
  • The two are independent controls and neither appears in a display specification, which lists coverage, peak luminance, contrast and colour temperature.
  • And the trade against gamut is structural. A wide gamut needs primaries on the steep flanks of the cone sensitivities, and steepness is what makes a peak shift matter.

What this collection already established

A gamut has a population computes the whole thing directly: two hundred eyes drawn from a pigment template, a match made for the reference observer, and the fraction of the population who would reject it. The spread rises as the primaries narrow, and the figure at the head of this essay is that result.

That answers the question a manufacturer asks — how many of my customers will be unhappy — and it is the right question. What it cannot do is say which property of the eye is producing the unhappiness, because a population statistic aggregates over all of them at once.

The audit’s six departures are that decomposition. Each is a single physiological parameter moved by a stated amount, so a display’s cost can be attributed rather than merely measured, and attribution is what turns a measurement into a design input.

Which variable moves which term

Two display design variables and two observer terms, and the pairing is clean.

Bandwidth excites the pigment peaks. A peak shift is a derivative rather than a displacement, largest where the sensitivity curve is steep, and a narrow source samples a curve at a point while a broad one integrates it. So the peaks departure rises monotonically as the emitters narrow: 1.77 ΔE₀₀ under a broad phosphor, 2.84 under a three-emitter LED, 3.60 under three lasers.

Position excites the macular pigment. Macular absorption is a band centred near 460 nanometres and forty wide, so what matters is whether a primary lands in it. A blue pump at 452 is inside; one at 470 is not. Under a white LED the macular departure is 3.37 against 1.71 under daylight, and the difference is almost entirely the pump’s position.

The lens departure follows neither cleanly; it tracks the source’s overall blue content and is worst under warm light, which is a room property rather than a display one.

So a display has two knobs and a room has a third, and a specification that reported the first two would let a purchaser reason about the third.

The same white, matched at six primary widths. At every width the three primaries are solved to match D65 exactly for the reference member; the bands are what the population sees. A broad primary integrates the observer differences over a band and averages them away; a narrow one samples them at a point and passes them straight through. From 40 nm to 2 the ninety-fifth percentile rises from 11.3 to 17.9 ΔE00, monotonically, and the technology has been moving from left to right for thirty years.
Fig. 2 The same sweep at a different set of widths. The rise is monotone across every set tried, which is what a bandwidth effect looks like.

Most of the cost is already spent. The measurement has a shape a manufacturer would want to know and it is not the shape the marketing suggests.

Going from a broad phosphor backlight to a three-emitter one — twenty to thirty nanometre emitters — raises the peaks departure from 1.77 to 2.84 ΔE₀₀, a factor of 1.6. Going from there to three lasers at a fifth of a nanometre raises it to 3.60, a further factor of 1.27.

The large step is the first one and the industry took it years ago. Quantum-dot and three-emitter backlights are ordinary consumer hardware; the move to laser primaries adds a quarter to a term that has already increased by more than half.

That reframes the usual worry. Laser projection is treated as the observer-metamerism problem, and by this decomposition it is a modest increment on a problem that arrived with the previous generation of displays. A viewer who has trouble with a laser projector has been having a smaller version of the same trouble with their television.

The trade nobody can engineer away

There is a structural reason a manufacturer cannot buy their way out, and it is worth being precise about.

A wide gamut needs primaries far from the white point, which means far out along the spectral locus, which means at wavelengths where at least one cone’s sensitivity is small and falling steeply. Steep is exactly what makes a peak shift matter.

A primary placed at a cone’s own peak would be maximally observer-robust and would be a poor primary. The middle-wavelength cone peaks at about 541 nanometres, and a primary there is a yellowish green whose chromaticity is well inside the locus — a small gamut contribution.

So the primaries that make a wide gamut are the primaries on the steep parts of the curves, and the two goods are bought with the same money. A primary is chosen for four things — efficiency, gamut, the inverse’s conditioning and manufacturability — and observer agreement is a fifth that pulls against the second — which is one more objective in a set that already refuses to be optimised together.

The one escape is a fourth primary. Four primaries have a choice: a colour can be rendered by more than one drive combination, all colorimetrically identical for the reference observer, and the combination minimising the population spread can be selected. That is a free optimisation once the hardware exists, and it is one of the things a fourth primary is a design for.

A fourth primary, swept — every setting an exact match, none of them the same. Four primaries matching three numbers leave one degree of freedom. Along the horizontal axis it is the fourth primary's share of the white's luminance; at each value the other three powers are solved exactly, so every point on this plot is a floating-point-exact match for the reference member — worst residual 1.3e-15 — and no colorimeter can tell them apart. What the population sees runs from 14.2 ΔE00 at the ninety-fifth percentile to 16.9, a factor of 1.19. The best setting is the largest share the arithmetic admits, so what stops it is not colour but the requirement that four powers stay positive.
Fig. 3 A fourth primary swept through its free parameter. Every setting is an exact match for the reference observer and the population sees them differently, so one of them is best and nothing colorimetric distinguishes them.

The white is free and the colours are not

There is a consequence of the round’s identities that a display engineer can use and it is a curious one.

Every observer agrees about a display’s white, exactly. The white is what the eye adapts to, so it is a stimulus judged against itself, and a stimulus judged against itself is observer-invariant — not nearly, but to the floating-point floor, whatever the primaries are.

So a laser projector’s white looks right to everybody and its saturated colours do not, which is the reported field experience and is usually credited to careful white-point calibration. It is not the calibration; it is an identity, and no calibration was needed for it.

That says where calibration effort is worth spending. Adjusting a display’s white to please a population is adjusting the one thing every member already agrees about. The disagreement is at saturation, and it is at saturation that no consumer display offers an adjustment.

One tolerance decision, at four spectral distances. Every column is a pair of samples ΔE00 1.0 apart for the observer a colorimeter models — solved to that value by bisection, so the instrument would report the same number for all four. What differs is how far apart the two spectra are, which is achieved by adding a metameric black the reference observer cannot see. The bands are what two hundred people report: from 1.13 at the ninety-fifth percentile when the spectra are the same shape to 2.26 when they are not, and the worst case reaches 3.7. No specification records the quantity on the horizontal axis.
Fig. 4 What fraction of a population accepts a match at a stated tolerance. The aggregate a manufacturer needs, which the decomposition explains and does not replace.

That figure is the aggregate the decomposition sits under, and putting the two side by side is the honest way to present either. A manufacturer deciding whether to ship needs the acceptance fraction; an engineer deciding what to change needs to know that bandwidth moves one term and position moves another.

Neither report is sufficient alone. An acceptance fraction with no attribution leads to the wrong optimisation — narrowing everything and then compensating with calibration, which cannot work because the disagreement is at saturation and calibration acts on the white. An attribution with no aggregate leads to over-engineering a term that turns out to affect two per cent of viewers.

Where the disagreement comes from, one variate at a time. The same measurement with one source of variation live and the other three held at their medians. The largest is the lens — entered as an age, because that is what it is a function of — at 7.7 ΔE00 against lens density, entered as age 20–70. So the biggest single reason two people disagree about a colour is the one thing about an observer that is written on their passport, and no colour specification has a field for it. The shares do not sum to the whole and are not expected to: the variates enter a nonlinear function of the spectrum, so this is a ranking rather than a decomposition.
Fig. 5 Which of the four varying parameters contributes most to the population’s spread. The attribution the audit performs by construction, done here by suppressing one variate at a time.

The population machinery can attribute too, by holding one variate at its median and re-running, and it broadly agrees with the audit’s ordering. That agreement is worth having because the two methods share almost nothing: one moves a parameter by a stated amount between two constructed eyes, the other suppresses a variate across two hundred sampled ones.

Two methods with different failure modes agreeing about an ordering is a good deal stronger than either alone, and it is the reason the decomposition can be offered as a design input rather than as an illustration.

What a specification could carry

The decomposition suggests two numbers a display specification could report, and both are computable from data manufacturers already have.

The emitter bandwidths, which decide the peaks term. Full width at half maximum for each primary, which is on every emitter’s datasheet and is not on any display’s.

The emitter centres, which decide the macular term. Also on every datasheet, also absent from display specifications, which report a colour gamut coverage percentage instead — a number that is a function of the centres and throws away everything else about them.

From those two a purchaser could estimate the observer spread for their own population and application. A signage display seen by the public needs a different answer from a grading monitor used by three people whose eyes could in principle be characterised.

Coverage of a gamut is the wrong summary for this purpose, because two primary sets with identical coverage can have quite different bandwidths and positions. It is the right summary for the question it was invented for, which is how much of a standard’s colours can be shown.

Where the audit’s numbers and the population’s differ

The two approaches give numbers that are not comparable and it is worth saying why rather than reconciling them.

The audit measures two observers differing in one parameter at a stated strength, and reports a distance. The population work measures two hundred observers differing in four parameters at once, and reports a percentile of a distribution of match errors.

A distance between two specified observers and a percentile of a population are different objects. The audit’s 3.60 ΔE₀₀ for the peaks departure under a laser projector is not a prediction that any particular fraction of viewers will see 3.60; it is the distance between two constructed eyes six nanometres apart in long-wavelength peak.

What transfers between them is the direction and the attribution. Both say narrow primaries are worse; only the audit says which parameter is doing it; only the population says how many people are affected. They are complementary and a reader should not average them.

One exact match, handed to two hundred people. Three primaries 30 nm wide are solved so that their sum has exactly the tristimulus values of D65 for the reference member of the population — three equations, three unknowns, residual 4.3e-16 of the white's luminance. The histogram is what everybody else sees: a median of 7.0 ΔE00, a ninety-fifth percentile of 13.3, and a worst case of 15.0. Two members picked at random disagree with each other by 3.8 units at the median. Nothing about the two spectra changed between the reader of this caption and the next one.
Fig. 6 What a population makes of a match at a stated primary width. The aggregate a manufacturer needs, at one point of the bandwidth sweep.

The histogram at a fixed width is worth seeing beside the sweep, because a distribution says something a percentile cannot: the population’s disagreement is not symmetric, and its upper tail is what produces complaints.

The same white, matched at six primary widths. At every width the three primaries are solved to match D65 exactly for the reference member; the bands are what the population sees. A broad primary integrates the observer differences over a band and averages them away; a narrow one samples them at a point and passes them straight through. From 40 nm to 2 the ninety-fifth percentile rises from 11.3 to 17.9 ΔE00, monotonically, and the technology has been moving from left to right for thirty years.
Fig. 7 The same sweep over a different set of widths. The rise is monotone in every set tried, which is what makes bandwidth a design variable rather than an artefact of one choice of points.

Two sweeps over two sets of widths give the same monotone rise, which is the check that the trend is the bandwidth rather than the sampling of it.

What was computed, and how

The bandwidth sweep is this collection’s existing population machinery, drawing two hundred eyes from the pigment template with age, macular density, cone optical density and peak wavelengths varying together, and computing what fraction of them accept a match made for the reference member.

The per-departure numbers are the audit’s, computed on a quarter-nanometre grid for the laser projector because the site’s own five-nanometre grid turns a three-line spectrum into a one-line one and reports every observer departure as exactly zero.

The three lamps compared — a broad phosphor, a three-emitter LED, a three-laser projector — are constructions with stated emitter widths of 118, 18 and 0.2 nanometres respectively, so the bandwidth axis spans nearly three orders of magnitude.

A grading suite could do what a television cannot. There is one setting where the whole problem is soluble and it is worth naming because it is the only one.

A colour grading suite has a small, known set of viewers. Their lens densities follow from their ages, their macular densities are measurable in a few minutes by flicker photometry, and their cone peaks could in principle be estimated from a Rayleigh match. A personalised observer is available for three people and not for three million.

With a four-primary display and a characterised viewer, the free parameter a fourth primary supplies could be set per viewer rather than per population — every drive combination colorimetrically identical for the reference observer, one of them chosen to be correct for the person in the chair.

That is a real product and nobody makes it. What stands in the way is not the optics or the arithmetic; it is that the characterisation has no standard method and the benefit is invisible to anybody who has not been shown the alternative. It is also, by the numbers in this round, worth two to three colour differences to the one person it is set for — larger than any calibration adjustment a grading suite currently makes.

A last note on what a purchaser can actually do with any of this. Almost nothing, at present: no display specification carries emitter bandwidths or centres, and a purchaser cannot compute an observer term from the numbers on a datasheet.

What they can do is prefer a display whose primaries are broad where a choice exists — which for a monitor means preferring a white-LED backlight over a quantum-dot one at equal gamut coverage, and for a projector means preferring a lamp over a laser. Both preferences run against the marketing and both are defensible on the audit’s numbers.

The more useful lever is the one a facility rather than an individual has: characterising the small number of people whose judgements matter, which is available and is described above, and which no equipment purchase substitutes for.

Where the model stops

The lamps are constructions and the emitter shapes are Gaussians. Real quantum-dot emitters have asymmetric tails, real laser projectors are deliberately broadened for speckle reduction, and a real phosphor has a long red tail no Gaussian carries.

The departures are one parameter at a time, so nothing here says what a real pair of observers differing in all six would see on a narrowband display. The population work answers that and does not decompose it.

And the gamut trade is argued from where the cone sensitivities are steep rather than computed. A proper treatment would optimise primary positions jointly for gamut area and population spread, which is a small optimisation and is not in this round.

One more number belongs beside the design advice, because it bounds how much any of it is worth. The largest observer departure this round measures under a laser projector is 3.60 ΔE₀₀ for the pigment peaks, and the largest under a broad daylight source is 2.38 for the lens. So the whole of what a display’s primary design controls is the difference between those, which is about one and a half units.

That is worth having and it is not the dominant term. Most of a viewer’s observer departure travels with them — their lens, their macular pigment, their peaks — and the display decides how much of it is expressed rather than how much of it there is. A manufacturer can move a viewer between 2.4 and 3.6 units of disagreement with the standard, and cannot move them below 2.4.

The generalisation

The habit is about decomposing an aggregate before trying to improve it.

A population statistic — a rejection rate, a mean opinion score, a failure fraction — is the number a decision needs and it is the wrong number to design against, because it aggregates over every mechanism at once and therefore responds to every design change by an amount nobody can predict.

The move is to decompose it into terms that each respond to one design variable, even when the decomposition is cruder than the aggregate. A crude attribution that says this knob moves that term is more useful for design than a precise aggregate that says the total is 4.2.

The failure mode is to optimise the aggregate directly, which works and teaches nothing: the optimum is found by search, it does not transfer to the next product, and nobody learns which variable mattered. An aggregate is for deciding and a decomposition is for designing, and the two are different reports.

Who found it, and when

Observer metamerism on displays became a production concern as wide-gamut backlights reached consumer hardware, and the CIE published a technical report on assessing it in 2016. The industry’s response has largely been to characterise the spread rather than to design against it, because the design trade is the one described above and has no comfortable resolution.

The observation that a display’s white is observer-invariant by identity does not appear to be stated in that literature, though the field experience it explains is common.

Where the ladder goes next

The other place the observer term lands hardest is where a camera’s colours come from: the chart was measured by an observer too, and a profile fitted against measured tristimulus values inherits every departure in this round.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Display gamutIndividual variationMacular pigmentNarrow band displaysObserver metamerismPopulationPrimariesSpecificationTrade-offVisual pigment