What the brain does

A model judged in another model's unit

An appearance shift is a change in what an observer would report, not a change in a stimulus, so measuring one with a matching difference means first asking what stimulus a settled observer would need to be shown to give the same report. That step is not bookkeeping — it is the whole distinction the field rests on, and it costs a factor of 1.7 across the menu.

Assumes Matching is not appearance, The room settles after the eye does and A choice with no magnitude.

Five of the six quantities in this audit are printed by their own machinery in the unit the audit is about. The sixth is not, and it is the one that makes the audit worth running.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit.
Fig. 1 Six published quantities under six colour-difference formulae. The bottom row is an observer two seconds into a new room, and it is the only row whose own file publishes it in CAM16-UCS rather than in ΔE2000 — because the model it comes out of defines that unit.

The claim

Measuring an appearance difference with a matching formula requires running the appearance model backwards first, and the requirement is the distinction between the two kinds of model rather than a detail of implementation.

  • An appearance shift is a difference between two reports. Two seconds after a change of light, an observer’s report of a fixed stimulus is a lightness, a chroma and a hue that a settled observer would not give.
  • A matching formula takes two tristimulus values, and a report is not one. The conversion is cam16Inverse: what stimulus would a settled observer have to be shown to report this?
  • The step is not neutral. It carries the whole viewing condition — surround, adapting luminance, background — and a different room gives a different answer for the same report.
  • The six units then spread the answer by a factor of 1.705, from 1.88 to 3.21.
  • And the formulae that weight chroma read it high by fourteen to twenty-one per cent, which is the reverse of every other row in the inventory. The count of how many read it high at all turns out to depend on which of two published values for its own unit is used, and that is a section of its own below.

Two values for one reading, and the claim holds against one of them

The quantity is introduced as 2.65 CAM16-UCS units — the figure the clock’s own machinery publishes — and the table’s CAM16-UCS row reads 2.788. Those are two numbers for one reading in one unit, five per cent apart, and the essay’s headline claim depends on which is used.

Counting against the table’s own row at 2.788: Oklab reads low by 32 per cent and ΔE*uv by 2, and only ΔE*ab, ΔE*94 and ΔE2000 read high. Three of six, not five.

Counting against the 2.65 the quantity is introduced as: everything except Oklab reads high — ΔE*uv by 3 per cent, the appearance unit’s own recomputed value by 5, ΔE*ab by 6, ΔE*94 by 20 and ΔE2000 by 21. Five of six, exactly as claimed.

So the claim block and the table are computed against different reference values, and the percentage column beside the table is against the second. The sentence and the column it sits under do not agree, and the disagreement is invisible because the two reference values differ by only five per cent and the claim is about signs rather than sizes — a five per cent shift in the reference flips the sign for the two rows that sit closest to it.

The five per cent is itself informative rather than an error to be picked. The clock’s file computes the shift in its own reference room with its own stimulus; this essay recomputes it in the audit’s reference condition, and the audit’s viewing parameters are stated a few sections down as an adapting luminance of 100, a background of 20 and an average surround. Two computations of one quantity, in one unit, under two viewing conditions, five per cent apart — which is a measurement of exactly the sensitivity the essay warns about when it says the reading moves with each of those three parameters, arriving as a discrepancy rather than as a result.

That makes the repair worth more than tidiness. Quoting both values with their conditions turns a contradiction into a datum: the appearance shift at two seconds is 2.65 in the clock’s condition and 2.79 in the audit’s, so the viewing condition is worth five per cent on this quantity — which is a fifth of what the choice of unit is worth and belongs in the same table.

And the sign question has a clean resolution once the reference is fixed. Against either value, the two formulae that divide by chroma read the appearance shift high by fourteen to twenty-one per cent, and Oklab reads it low by thirty. Those two facts are what the section below is about, they do not depend on which reference is chosen, and they are the part of the claim worth carrying. What does not survive is the count — five of six is a statement about where the reference happens to sit among six closely spaced readings, and three of the six sit within six per cent of it.

What the quantity is

Two seconds into a new room, according to this collection’s own clock.

The setting: an observer looks at a fixed stimulus, the light changes from one white to another, and adaptation takes time — a fast receptor component with a time constant of about a second and a slow one of about a minute, splitting roughly two thirds to the fast. At any moment the observer is working from a white part-way between the two, and the appearance model can be asked what they would report.

At two seconds the report differs from the settled one by 2.65 CAM16-UCS units, which is the number the clock’s own machinery publishes. It is a real quantity — an observer walking into a shop makes judgements in that state, and the model has no argument for them.

The unit is the model’s own. CAM16-UCS is the uniform space built on CAM16’s outputs, and reporting an appearance shift in it is the only thing that is not a category error, because the two things being subtracted are two of that model’s reports.

Why a matching unit cannot take it directly

A colour-difference formula’s arguments are two colours. A colour, to such a formula, is three numbers describing light: a tristimulus value, or something computed from one.

CAM16’s output is not that. It is J, a lightness; C, a chroma; h, a hue angle — descriptions of what somebody would say, computed from a stimulus and a room. The same stimulus in a dim surround gives a different J. The same J in two rooms corresponds to two different stimuli.

So there is no tristimulus value to hand to ΔE2000. What there is instead is a well-defined question with an answer: what stimulus, shown to a settled observer in the reference room, would produce this report? The appearance model runs backwards and gives it. That is cam16Inverse, and it is a standard part of the model, used routinely for gamut mapping between viewing conditions.

Running it converts an appearance into a stimulus, and then a matching formula can be applied. But the conversion is the whole of the distinction the field rests on, and it is worth being explicit about what has been assumed by making it:

That the appearance has a stimulus. Not every report does. A J above what the reference room can produce inverts to a tristimulus outside the reachable range, and this collection marks rather than clips — the two-second report does invert cleanly, and a shift twice as large would not.

That the reference room is the right one. The inverse is taken in the settled condition, so the answer is the stimulus difference an observer in the new room’s settled state would need. Take the inverse in the old room instead and the number changes.

And that the two reports are comparable at all. A report from a partly adapted observer and one from a settled observer are two states of one person, not two colours, and the assumption that their difference is a colour difference is exactly the assumption the appearance field exists to question.

What the six units say

unit reading against its own unit
Oklab 1.884 −32%
ΔE*uv 2.735 −2%
CAM16-UCS 2.788
ΔE*ab 2.804 +1%
ΔE*94 3.172 +14%
ΔE2000 3.212 +15%

The spread is ×1.705, which is the smallest of the six rows in the audit’s inventory. That is worth noting before anything else: the quantity most exposed to a category error is the least exposed to a change of unit.

The reason is that the inverse model has already done most of the work. Two reports two seconds apart differ mainly in one direction — the adapting white is part-way moved, so the whole scene is shifted along the blue-yellow axis in a fairly clean way — and a shift concentrated in one direction is measured similarly by any formula that is not badly wrong about that direction. The units disagree most about differences with mixed components, which is why the observer-gap row spreads by 2.23 and this one by 1.7.

A single ellipse is where a unit’s anisotropy stops being an average, and the one at x 0.390 is nearly the roundest in MacAdam’s set.

One of MacAdam's ellipses as each unit sees it, at x 0.390, y 0.237. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 1.43; the least round is ΔE*94 at 2.68. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 2 One of MacAdam’s ellipses at x 0.390, y 0.237, drawn as its perimeter distance in each of the six units and scaled to its own mean radius so the six compare as shapes. The roundest reading is ΔE′ at an anisotropy of 1.43, and a unit in which one step meant one thing in every direction would have drawn a circle.

The row that reads it low

Oklab reads the shift at 1.884, a third below its own unit and 41 per cent below ΔE2000. Nothing else in the inventory has that shape — in five of the six rows Oklab is at or near the bottom by a modest margin, and here it is at the bottom by a large one.

The direction the adapting white moves is towards yellow, and Oklab was constructed with particular attention to the blue-yellow axis: the space’s own documentation cites the hue-uniformity problems of CIELAB in the blues as a motivating case, and its cube-root-of-cone-like coordinates spread that region differently. A shift concentrated along the direction a space was built to handle differently is a shift that space measures differently, which is the same mechanism as the lens row in the adaptation census.

That is a specific and slightly awkward conclusion. The quantity is an adaptation shift, the axis is blue-yellow, and the axis is exactly the one the six formulae disagree about most. Every adaptation number in this collection is exposed to that, and this row is where it is visible because it is the row with the cleanest single direction.

The weighted formulae read it higher, and that is the reverse

Every other row in the inventory has CAM16-UCS at the top and matching formulae below. This one has CAM16-UCS in the middle, with the two chroma-weighted matching formulae above it by fourteen and fifteen per cent.

The explanation is that this is the one row where the appearance unit is not being used out of its element. In the other five, CAM16-UCS is being asked about a stimulus difference — a change of light on a surface, a camera error, a metameric pair — and it answers by running a model with a room in it, which adds an incomplete adaptation the matching formulae do not have and reads high.

Here the appearance model is being asked about the thing it is for, and there is no extra adaptation to add: both reports come from the same model. So the appearance unit reads normally and the matching formulae, working on stimuli reconstructed by the inverse model, read high instead.

Which is a two-way result. It says the appearance unit’s tendency to read high in this audit is a consequence of being used on matching questions, not a property of the unit. And it says that measuring an appearance difference with a matching formula overstates it by about 15 per cent, in the published unit, for the same reason in reverse.

Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu.
Fig. 3 Where on the scale the units disagree. The two-second shift sits at about 2.8 units, in the band where the appearance unit has almost converged on the others — which is why this row’s spread is the smallest in the inventory.

The blue end of the set is where every one of the six units does worst, and it is the comparison that decides how much of the ranking is the unit and how much is the colour.

One of MacAdam's ellipses as each unit sees it, at x 0.187, y 0.118. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 2.15; the least round is ΔE*ab at 8.60. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 4 The ellipse at x 0.187, y 0.118, in the same six units. The roundest reading here is 2.15 against 1.43 at the other centre, so which unit is least anisotropic is a question with a different answer in different parts of the diagram.

The inverse model is not invertible everywhere

One property of the conversion deserves a section, because it is the reason the step cannot be made routine.

cam16Inverse takes a J, a C and an h and returns a tristimulus value, and for most inputs it does so cleanly. But the forward model is not a bijection onto anything convenient: it maps a cone of physically realisable stimuli onto a region of report space, and report space is larger than that region. A report can be perfectly well-formed and correspond to no light.

Two ways that happens here. A J above 100 asks for something lighter than the reference white in the reference room, which the inverse returns as a tristimulus with Y above the white’s — physical, but outside what a reflective surface can be. And a high C at a hue where the room’s gamut is narrow inverts to a tristimulus outside the spectrum locus, which is not a colour at all.

The two-second report used here inverts cleanly and well inside both bounds, which is checked rather than assumed. A shift twice as large would not, and that is a real limit on the method rather than a caveat: an appearance difference large enough to be interesting is an appearance difference a matching formula may not be able to be given. The collection’s rule is to mark what cannot be shown rather than clip it, and the same rule applies to a report that has no stimulus.

A second ellipse near the blue end is where the six units are furthest from each other, and it is worth drawing beside the first.

One of MacAdam's ellipses as each unit sees it, at x 0.253, y 0.125. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 1.46; the least round is ΔE*ab at 7.90. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 5 One measured ellipse drawn as its perimeter distance in each of the six units, with every outline scaled to its own mean radius so the six compare as shapes. A unit in which one step meant one thing in every direction would have drawn a circle here, and none does.

What this says about mixing the two

The practical question is whether an appearance quantity should ever be reported in a matching unit, and the answer has a shape.

For a comparison inside one room, never — and there is no reason to. Two appearances in one viewing condition can be differenced in the model’s own unit, and doing anything else adds an inverse model for no gain.

For a comparison against a matching quantity, sometimes, with the conversion stated. This audit needed exactly that: a table of six quantities cannot have one row in different money. The conversion is legitimate, the number is defensible, and what makes it defensible is that the inverse model is named — a reader who disagrees about the reference room can redo it.

And for a tolerance, no. A tolerance is a boundary, the six formulae draw six different boundaries, and an appearance-based tolerance converted into a matching unit inherits both the boundary disagreement and the inverse model’s dependence on the room. Two independent sources of disagreement in a quantity whose whole job is to decide a yes or a no.

The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2.
Fig. 6 The adaptation census in six units. Every row of it is a change of light acting on a surface, judged by a matching formula — the same operation this essay performs on an appearance shift, in the opposite direction and without an inverse model in the way.

The last figure and the inventory together make a point about the audit’s own shape. Five rows are matching questions asked in an appearance unit, and one is an appearance question asked in matching units, and both directions are category errors of the same kind committed deliberately in order to get one table. The five read high because the appearance model adds an incomplete adaptation; the one reads high because the inverse model has to invent a stimulus. In both cases the error has a sign and a size, and knowing both is worth more than refusing the comparison would have been.

A third, nearer the middle of the diagram, says the ordering of the units is not the same everywhere.

One of MacAdam's ellipses as each unit sees it, at x 0.258, y 0.450. A single discrimination ellipse from MacAdam's 1942 measurement, drawn as the distance from its centre to each point of its perimeter in each of the six units, with each outline scaled to its own mean radius so the six can be compared as shapes. A unit in which a step of one size meant the same thing in every direction would draw a circle here. None of them does. The roundest is ΔE′ at an anisotropy of 1.38; the least round is ΔEok at 2.45. What the outlines have in common is their orientation: every unit agrees about which direction this ellipse is long in and disagrees only about how long.
Fig. 7 Another of MacAdam’s ellipses in the same six units. Which of them is roundest at this centre need not be the one that wins on the average, which is what makes a ranking of units a statement about a set of colours rather than about the units.

Where the model stops

CAM16 has no clock. The two-second figure comes from a time course bolted onto it — two exponentials of stated constants, applied to the adapting white — and the model itself has no argument for a partially adapted state. The quantity is therefore a composition of a standard model and a non-standard extension, and the extension’s constants have their own ranges.

The inverse is taken in the settled condition at the reference viewing parameters: an adapting luminance of 100 cd/m², a background of 20, an average surround. All three are choices with continuous ranges standing on tabulated values, and the reading moves with each.

And the whole quantity is one stimulus. A different colour two seconds into the same room gives a different shift, and the distribution over stimuli is wider than the spread over units — which is the same warning every other row in this audit carries and is worth repeating here because a single stimulus makes it easy to forget.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one.
Fig. 8 The six units by how far each is from a rescaling of the published one. The appearance unit is second closest, which this audit found surprising and which this row explains: it is close on the questions it is suited to and reads high on the ones it is not.

That closeness is an average over pairs, and an average over pairs is exactly the statistic that cannot say whether one unit is a rescaling of another or a compromise between two behaviours. The pair-by-pair picture separates them.

ΔE′ against ΔE2000, on the pairs both were calibrated over. A scatter of 374 pairs of surfaces. The horizontal position is the pair's difference in ΔE2000 and the vertical is the same pair in ΔE′, multiplied by the single factor that best carries one onto the other. The diagonal is where a pure rescaling would put every point. 281 of the 374 pairs sit above it and the rest below, and the departure grows with the difference — the scatter is 24.4 per cent of the mean and the rank correlation is 0.972. Every point off the line is a pair the two units disagree about, and a pair of points on opposite sides of it is a comparison they would decide differently.
Fig. 9 The appearance unit against the published one, pair by pair, after each has been carried onto the other’s scale by its single best factor. A pure rescaling would put every point on the diagonal; the departures are the pairs on which a model of appearance and a formula for matching are answering different questions about the same two colours.

Who found it, and when

The forward and inverse forms of CIECAM are both in the standard, and the inverse is not an afterthought — gamut mapping between viewing conditions is one of the model’s stated purposes and cannot be done without it.

What is not standard is treating the inverse as a measurement decision. In gamut mapping it is machinery: a colour goes in, a colour comes out, and nobody asks what has been assumed. Used as a step in reporting a difference, it becomes an assumption with a name — that the appearance in question has a stimulus, in a stated room, that a settled observer would see the same way — and the assumption can then be argued with.

The distinction between matching and appearance is Hunt’s and the field’s, and this collection draws its foundation boundary at exactly that line. This essay is that boundary crossed deliberately, once, with the toll recorded.

Where the ladder goes next

The unit is one of three structural choices this collection named and could not reach. The second is that adaptation is a diagonal at all — a model with three numbers where the exact answer has nine — and the interesting question about it is not how far the diagonal is from exact, which is published on fourteen rows, but where those nine numbers would have to come from.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 10 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Appearance modelCalibrationChromatic adaptationCIECAM16Colour differenceColourfulnessInverse modelLightnessSurroundViewing condition