What the brain does

The audit, read as appearances

The previous round priced six observer departures in a matching unit and left them there. Read in the appearance unit an appearance prediction would actually be judged by, they are not the same six numbers scaled by a constant — the model amplifies the smallest by 1.58 and the largest by 1.03, so the range between the largest and smallest term narrows from 3.34 to 2.17. The published ranking survives in the mean and changes on fourteen of the forty-two surfaces it was averaged over.

Assumes Where the model's curve does not matter, The appearance model takes XYZ and A model judged in another model's unit.

The previous round’s largest single result is a table of six numbers: what each of six physiological differences between two observers costs, priced in colour differences, over forty-two surfaces under three lights. The appearance model takes those numbers as input and the round said so, and then stopped — because carrying an average through a nonlinear model was an operation with no error bound on it, and its own shortfall list recorded the gap.

The audit's six departures, read twice. Each departure priced in the matching unit it was published in and in the appearance unit an appearance prediction would be judged by. The model does not scale them by one factor: it amplifies the smallest by 1.58 and the largest by 1.03, so the range between them narrows from 3.34 to 2.17. The mean ranking is unchanged and 14 of the 42 surfaces reorder.
Fig. 1 The six departures priced in the matching unit they were published in and in the appearance unit an appearance prediction would be judged by. The two sets of bars are not proportional.

The claim

The model does not scale the audit’s six departures by one factor, and the direction of the reweighting narrows the range between them.

  • The amplification runs from 1.03 to 1.58: the lens departure is barely changed and the cone-density departure is half again larger.
  • The range between the largest and smallest term falls from 3.34 to 2.17.
  • The published mean ranking survives — lens, macular pigment, pigment peaks, rods, field size, cone density — in both units.
  • And it changes on 14 of the 42 surfaces it was averaged over, which is a third of them.

Why this can now be done

The measurement rests on the previous rung and would have been indefensible without it.

Reading the audit in the model’s unit means taking a mean over surfaces of a quantity computed through the model, and the question of whether such a mean is legitimate is exactly what the previous rung answered: over the spread a population produces, the mean of the model’s answers is its answer for the mean to within 1.4 per cent, with a known sign.

So the operation has a bound on it, and the bound is two orders of magnitude below the effects measured here. That is the sequence a round should have — establish the licence, then use it — and it is the reverse of the order the work was actually done in, which is the ordinary situation.

Where averaging the model's answers is and is not averaging its argument. Each row takes a spread of situations, averages the model's predictions across them, and compares that against the model's prediction for the average situation. The bar is the gap as a share of the spread itself. Over the differences between observers it is 1.4 per cent — the model is very nearly linear there. Over the range of adapting luminance one room covers in a day it is 59 per cent, and from indoors to outdoors 74.
Fig. 2 The licence, from the previous rung. The first bar is the population case and it is what makes everything in this essay a measurement rather than an assumption.

What the two units are

The audit’s unit is ΔE₀₀ between two CIELAB readings: what a matching experiment would report about two stimuli seen under identical conditions by two different observers.

The appearance unit is a distance in CAM16-UCS, which is what a difference between two reports comes to — not two stimuli but two accounts of how a stimulus looks, made by two observers in the same room.

The distinction between the two is not bookkeeping: an appearance shift is a change in what an observer would report rather than a change in a stimulus, and measuring one with a matching difference means first asking what stimulus a settled observer would need to be shown to give the same report. This collection has already found that step costs a factor of 1.7 across the menu of difference formulae.

Here the step is the other way round. The departures are stimulus differences, priced in a matching unit, and the question is what they come to as differences between reports.

What the model does to them

The amplification is the ratio of the two prices, and it is not one number.

The lens departure — a twenty-year-old eye against a seventy-year-old — comes in at 2.40 colour differences and 2.46 CAM16-UCS units, a factor of 1.03. The cone optical density departure comes in at 0.72 and 1.13, a factor of 1.58. Between them: rods at 1.52, field size at 1.45, pigment peaks at 1.33, macular pigment at 1.19.

The pattern is clean and it is the opposite of what a reader would guess. The departures the model amplifies most are the smallest ones. Ordering the six by their matching cost gives almost exactly the reverse of ordering them by their amplification.

That is a compression, and it is the model’s compression. A larger stimulus difference sits further along a concave response, where the slope has fallen, so it buys proportionally less report; a smaller one sits where the slope is steeper. The model narrows the range between its inputs, and the narrowing is by a factor of 3.34 over 2.17, which is 1.54.

Which of the model's correlates the room actually moves. One stimulus, unchanged, read at adapting luminances from 8 to 20000 candelas a square metre, each quantity against its own largest value. Brightness rises by a factor of 5.1 and colourfulness by 2.0. Lightness moves by 2.5 per cent and chroma by 3.3, because both are ratios to the white and the white moved too.
Fig. 3 Which of the model’s correlates the room moves, from two rungs back. None of the reweighting here is a room effect: the viewing condition is held fixed and identical for both observers in every pair.

The six, one at a time

The pattern is clean enough that each row is worth a sentence, because the mechanism differs between them and the mechanism is what transfers.

The lens, at 1.03. The largest departure in the matching unit and the one the model barely touches. A yellowed lens is a filter in front of everything, so it moves the reading a long way and lands far along the model’s response, where the slope has fallen. Almost all of its matching size survives into the appearance unit and none of it is amplified.

The macular pigment, at 1.19. The same mechanism at a smaller size and a correspondingly larger amplification.

The pigment peaks, at 1.33. A peak shift is a derivative rather than a displacement, so it produces a smaller reading change of a different character — one that sits nearer the steep part of the response.

The field size, at 1.45, and the rods, at 1.52. Both middling departures, both amplified by about half.

The cone optical density, at 1.58. The smallest departure in the matching unit and the most amplified, which is the extreme case of the same rule.

So the ordering by amplification is very nearly the reverse of the ordering by size, and it is not exactly the reverse: the field size and the rods swap. That imperfection is the useful part, because it says the amplification is not purely a function of magnitude — the direction of each departure in the model’s own space matters too, which is the finding two ladders across arriving here.

What the narrowing means for a design decision

An audit’s ordering exists to tell somebody what to spend effort on, and narrowing the range changes the answer.

In the matching unit the largest term is three and a third times the smallest, which reads as a clear instruction: worry about the lens and ignore the cone density. In the appearance unit the ratio is two and a sixth, which reads as a spread rather than a hierarchy — six terms of comparable size, the largest about twice the smallest.

Those are different instructions. A designer given the first would characterise their viewers’ ages and stop; given the second they would want all six or none.

The honest position is that both are true and they answer different questions. A display manufacturer worrying about whether two viewers will accept the same match wants the matching unit. A lighting designer worrying about whether a room looks the same to two people wants the appearance unit. The audit was computed for the first and is quoted as though it settled both.

And the second reading is the one that generalises worse. In the appearance unit the six terms are close enough together that their ordering is not robust: it changes on a third of the surfaces, and the mean ordering survives by a margin that a different surface set could remove.

What survives, and what does not

Two of the audit’s conclusions are worth checking against the new unit and they come out differently.

The ranking survives in the mean. Lens, macular pigment, pigment peaks, rods, field size, cone density — the same order in both units. That is not guaranteed by anything: a reweighting by a factor of 1.5 could easily have swapped two adjacent terms, and the two closest pairs are field size against rods and rods against pigment peaks.

It does not survive surface by surface. On 14 of the 42 surfaces the two units rank the six differently, which is a third of them. The pairs that swap are the ones close together in both units, and the number of them says something about how much confidence a ranking of six means over one sample.

The winner survives the census's own construction; the middle of it does not. One row per perturbation of a constant the adaptation census is built from — the imaginary wall's centre wavelength, its width, its depth, its base, the macular filter's density and the two lens ages — each moved by an amount plausible for that quantity in its own units, up and down, and then all of them together. Each row shows where the five published transforms rank under it. Bradford holds the first column in all 14 rows. The second and third columns, which the table as built separates by six parts in a thousand, change places in 2 of them — so that ordering was never a fact about the transforms.
Fig. 4 The adaptation census, this collection’s largest computed result. It is the next object to carry across, and the previous rung’s licence covers it too.

And a conclusion about relative importance is weaker in the new unit. The audit’s headline is that the lens is three and a third times the cone density; in the appearance unit it is two and a sixth. Both statements are true about their own unit and a reader who takes away the lens dominates has taken away a claim that is a third weaker where it would actually be applied.

Why the mean ordering survives at all

The ranking being unchanged in the mean is worth an explanation, because a reweighting of half again could easily have broken it and did not.

The reason is that the audit’s six terms are not evenly spaced. In the matching unit they are 2.40, 1.70, 1.37, 1.05, 0.95 and 0.72, and the gaps between adjacent terms are 0.70, 0.33, 0.32, 0.10 and 0.23. The reweighting multiplies each by between 1.03 and 1.58 — a spread of 1.53 — and a swap needs the reweighting’s ratio between two adjacent terms to exceed the ratio of the terms themselves.

The dangerous pair is field size against rods, at 0.95 and 1.05: a ratio of 1.11, against an amplification ratio of 1.05 between them. It survives by a margin of six per cent.

So the ranking survives by luck rather than by structure, and it is the closest pair that decides. A surface set that moved either of those two by six per cent would break it, and fourteen of the forty-two individual surfaces do.

That is worth saying plainly because the opposite reading is available and is wrong. The ranking is robust to the change of unit would be a reasonable summary of the mean result and would be a claim about a margin of six per cent on one adjacent pair.

What this does not say

It does not say the audit’s numbers were wrong or should be restated. A colour difference between two stimuli is the right quantity for a matching question, and a great many of the audit’s uses are matching questions: whether two paints will be judged the same, whether a display’s white is acceptable, whether an instrument’s report transfers.

It says that the audit’s numbers are in a unit, that a second unit exists, and that the conversion is not a constant. A table of six numbers with a stated unit is a complete report; a table whose reader will apply it in a different unit is not, and the missing column is the six amplifications — the same shape as the second number a tolerance needs.

Those amplifications are cheap. They are one call to the model per row, they depend on the viewing condition, and they would let a reader in an appearance setting use the round’s results without redoing them.

What a stated lightness pins down, and where. A lightness quoted to 0.05 of a unit, inverted, and the luminance it fixes. Read as a fraction of the colour's own luminance the requirement is 0.90 per cent at J 10 and 0.097 at J 95, a factor of 9.3. Read in absolute luminance it is the other way round, by a factor of 6.5. Both readings are true and they answer different questions.
Fig. 5 What a stated lightness pins down. The same compression that reweights the departures makes a stated precision uneven, which is the last rung of this ladder.
The corner the hue scale has at each unique hue. How fast hue quadrature runs against hue angle, all the way round. It is piecewise, with a corner at each of the four unique hues, and its slope runs from 0.47 near 238 degrees to 1.92 near 90 — a factor of 4.1. The largest corner is at blue, where the slope changes by a factor of 0.41 across a single point.
Fig. 6 The model’s hue scale, from two rungs back. Every departure’s appearance distance passes through the opponent plane rather than through the quadrature, so this particular unevenness is not inside the numbers here — which is worth knowing, since it would have been the obvious suspect.

That the hue scale is not implicated is worth checking rather than assuming, and it follows from how CAM16-UCS is built: the uniform coordinates are formed from the opponent signals directly, not from the quadrature. So a distance in the appearance unit carries the model’s compression and its cone matrix and does not carry its hue table.

Three of the model’s four irregularities are in these numbers and one is not, which is the kind of statement that is only available once each has been measured separately.

The check that the reweighting is real

A reweighting by a factor of 1.5 between two units could be an artefact of how the two units are scaled against each other, and the check is to look at a quantity the scaling cannot touch.

The ratio of amplifications — 1.58 over 1.03, which is 1.54 — is dimensionless and survives any rescaling of either unit. Multiplying every CAM16-UCS distance by two would change all six amplifications and leave that ratio exactly where it is.

So the finding is the 1.54 rather than the 1.35. The absolute factor is a convention about how two scales were normalised; the spread between the six is a property of the model’s response and is what a reader should carry.

The same applies to the narrowing. The range falling from 3.34 to 2.17 is a comparison of two dimensionless ratios — largest over smallest, in each unit — so it too survives any rescaling. A model that scaled all six departures by any constant whatever would leave both ratios at 3.34.

What was computed, and how

Each departure is a pair of constructed observers differing in one stated argument, exactly as the previous round constructs them: age twenty against seventy, macular pigment at two standard deviations either side, cone optical density likewise, the long-wavelength polymorphism on the peaks, the CIE’s own second observer for the field size, and a tenth of the cone response for the rods — the seven ways to change an observer the round constructed.

Each observer forms their own tristimulus values for the same surface under the same light, normalises to their own white, and adapts to D65 through CAT16. The pair is then priced twice: once as a CIELAB colour difference and once as a CAM16-UCS distance at the reference viewing condition.

The audit's six departures, read twice. Each departure priced in the matching unit it was published in and in the appearance unit an appearance prediction would be judged by. The model does not scale them by one factor: it amplifies the smallest by 1.58 and the largest by 1.03, so the range between them narrows from 3.34 to 2.17. The mean ranking is unchanged and 14 of the 42 surfaces reorder.
Fig. 7 The same table again, with the reorder count on it. Fourteen of forty-two surfaces put the six departures in a different order in the two units.

The viewing condition is identical for the two observers in every pair, which is the assumption the whole measurement rests on: the two are looking at the same thing in the same room, and what differs is their eyes. An observer difference that also changed the room would be a different quantity.

The surfaces are the audit’s own forty-two, which are all pale, so the reweighting is measured over a set that does not span lightness. Over a set that does, the amplifications spread further, since the model’s compression is steepest where the audit’s surfaces are not — the same fault, in the same set, that the first rung of this round found.

Where the model stops

One viewing condition, one light and one route to the adapted values. The amplification is a property of where on the model’s response the departures land, so all three move it.

The two units are not commensurable and the ratio between them is not dimensionless in any deep sense: CAM16-UCS and ΔE₀₀ are both scaled so that one unit is roughly a just-noticeable difference, which makes the comparison meaningful in practice and not in principle — a unit rests on a space that was ranked, and this is two of them. A reader should treat the ratios between the six amplifications as the result and the absolute factor of about 1.3 as a scaling convention.

And nothing here says which unit an observer’s actual judgement follows. That is a psychophysical question, it would need an experiment, and this collection has none.

The generalisation

The habit is about a result that will be used in a different unit from the one it was measured in.

A measurement is reported in whatever unit the measuring apparatus produced, and it is used wherever the quantity matters. When the two are separated by a nonlinear map — a perceptual scale, a log, a model, a rating — the conversion is not a constant, and a table of results converts row by row rather than by a factor.

The move is to publish the conversion alongside the results. It is one extra column, its entries are cheap, and without it every downstream reader either redoes the measurement or applies a constant that does not exist.

The failure mode is that the constant gets applied anyway, silently, because a reader who needs the numbers in another unit will convert them somehow. A missing conversion column is not an absence; it is an invitation to assume proportionality, and proportionality is exactly what a nonlinear map does not have.

Who found it, and when

That appearance differences and matching differences are different quantities is the founding distinction of appearance modelling and is not in dispute. That converting between them is not a constant follows immediately and does not seem to be stated as a caution anywhere.

The reweighting’s direction — smaller departures amplified more — is a straightforward consequence of a concave response and would be predicted by anybody who thought about it. Its size, and the fact that it narrows the audit’s own range by half again, is what needed measuring.

Where the ladder goes next

Four rungs have taken the model’s coordinates, its scale, its curvature and its unit apart. What is left is the thing a specification actually contains, which is a number with a stated precision — and a lightness quoted to a tenth of a unit is two different requirements at the two ends of the scale, by a factor of nine.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Appearance modelAuditCIECAM16Colour appearanceColour differenceDeclared inputObserver variabilitySensitivitySpecificationUncertainty