What the brain does

Where the model's curve does not matter

This round has been about what a nonlinearity does to an average, and CIECAM16 is the most nonlinear thing in the collection. Over the spread a population of observers produces, the average of its predictions is its prediction for the average to within 1.4 per cent — the nonlinearity is there and the excursion is too small to reach it. Over the range of adapting luminance one room covers in a day, the same gap is 59 per cent of the spread, and from indoors to outdoors 74.

Assumes The hue scale has four corners, The census under another observer and A viewing condition is a moment.

The round this essay belongs to has one recurring finding: an average taken on one side of a nonlinearity is not the average taken on the other. It has now been measured on a mixture of two lights, on the reconstruction of a mosaic and on the composition of two departures, and each time the gap was large enough to matter.

Where averaging the model's answers is and is not averaging its argument. Each row takes a spread of situations, averages the model's predictions across them, and compares that against the model's prediction for the average situation. The bar is the gap as a share of the spread itself. Over the differences between observers it is 1.4 per cent — the model is very nearly linear there. Over the range of adapting luminance one room covers in a day it is 59 per cent, and from indoors to outdoors 74.
Fig. 1 The gap between the mean of the model’s answers and its answer for the mean, over five different spreads. The first is negligible and the last is most of the spread itself.

The claim

The same gap is 1.4 per cent of the spread over the argument this collection has spent a round measuring, and 74 per cent over an argument nobody varies.

  • Over eighty observers drawn from a population, the gap is 0.020 CAM16-UCS units against a spread of 1.41 — 1.4 per cent, and negligible.
  • It has one sign on every surface tested: the mean of the predictions is below the prediction for the mean, on 42 of 42, which is Jensen’s inequality rather than noise.
  • Over the adapting luminance one room covers in a day the gap is 59 per cent of the spread, and from indoors to outdoors 74.
  • What decides the difference is how far along the curve the argument moves, not how curved it is.

Why the population case is the surprising one

A model with a 0.42 exponent in it, two compressions and a division by the white is about as nonlinear as anything in this collection, and the natural expectation is that averaging over it is nowhere safe.

The previous round measured the population’s own spread carefully. The census under another observer found that observer departures on ordinary surfaces are between one and two and a half units, which is the size of the effects the census reports; the observer audit put six named departures at between 0.7 and 2.4 colour differences. So the spread is real and it is not small against a delivery tolerance.

It is small against the model’s curvature, and that is the whole of the explanation. A nonlinearity’s effect on an average is set by how far the average’s arguments spread along it, not by how bent it is: a very curved function is very nearly straight over a short enough interval, and a population’s disagreement about one surface is a short interval.

Which of the model's correlates the room actually moves. One stimulus, unchanged, read at adapting luminances from 8 to 20000 candelas a square metre, each quantity against its own largest value. Brightness rises by a factor of 5.1 and colourfulness by 2.0. Lightness moves by 2.5 per cent and chroma by 3.3, because both are ratios to the white and the white moved too.
Fig. 2 Which correlates the room actually moves, over three and a half decades of adapting luminance. Brightness rises fivefold and lightness moves by two and a half per cent, because one is absolute and the other is a ratio to a white that moved too.

The measurement is direct. Eighty observers, drawn with their four measured variates, each producing their own adapted tristimulus reading of the same surface; the model run on each and the answers averaged in the model’s own uniform coordinates; the model run once on the average reading. The two land 0.020 apart against a spread of 1.41.

So the previous round’s numbers survive the model. An audit’s result stated as a mean over observers can be carried through CIECAM16 as a mean, and the error in doing so is under two per cent of the quantity’s own spread. That is a licence the round did not have and now has.

And why it has a sign

The gap is small and it is not random.

On all forty-two surfaces the mean of the individual lightness predictions is below the model’s prediction for the mean stimulus. That is Jensen’s inequality: the model’s lightness is a concave function of the signal over this range, and the average of a concave function is below the function of the average, always, with no exceptions and no error term.

An inequality is a stronger statement than a measurement, and it is worth having for the same reason the exact zeros elsewhere in this round are worth having. The direction is guaranteed; only the size had to be measured, and the size is what makes the direction not worth correcting for.

The winner survives the census's own construction; the middle of it does not. One row per perturbation of a constant the adaptation census is built from — the imaginary wall's centre wavelength, its width, its depth, its base, the macular filter's density and the two lens ages — each moved by an amount plausible for that quantity in its own units, up and down, and then all of them together. Each row shows where the five published transforms rank under it. Bradford holds the first column in all 14 rows. The second and third columns, which the table as built separates by six parts in a thousand, change places in 2 of them — so that ordering was never a fact about the transforms.
Fig. 3 The adaptation census, this collection’s largest computed result, from an earlier round. Every number in it is an average, and this essay’s finding is that carrying such an average through the appearance model is safe.

The number that decides it

The two cases differ by a factor of fifty and the quantity that separates them can be written down.

A twice-differentiable function averaged over a spread has a gap given, to leading order, by half its second derivative times the variance of the argument. So the gap as a fraction of the spread goes as the second derivative times the spread itself — linear in how far the argument moves, not in how curved the function is.

That is why the two cases land where they do. The model’s curvature is the same in both; what differs is the excursion. A population’s disagreement about one surface moves the model’s argument by a small fraction of its range. A room’s luminance over a day moves it by three decades.

The rule of thumb that follows is worth carrying: compare the spread against the scale over which the function’s derivative changes appreciably. For the appearance model’s lightness that scale is most of the luminance range, so a population’s spread is short against it. For the adapting luminance, entering through a fourth root, the derivative changes appreciably within one decade, and a room covers three.

Neither number is available from the phrase the model is nonlinear, and both are one loop away.

What the negative result buys

A finding that something does not matter is worth as much as one that something does, and it is worth being explicit about what this one licenses.

The previous round’s averages can be carried through the model. Every number in the observer audit is a mean over surfaces and observers; every number in the adaptation census is a mean over a hundred and twenty-five surfaces and fourteen changes of light. Carrying such a mean into an appearance prediction was, before this measurement, an operation with an unknown error on it — and the round’s own shortfall list recorded that the census had never been run through the model.

It is now an operation with a bound: under two per cent of the quantity’s own spread, in the population argument, with a known sign.

And a bound with a sign is better than a bound. The gap is always in the same direction, so a mean carried through the model is always a slight over-estimate of the mean of the individual predictions, never an under-estimate. A conclusion that survives the over-estimate survives.

That is the shape a licence should take, and it is the reason to measure a negligible effect rather than to assume it. An assumed negligible effect is a hole in an argument; a measured one is a step in it.

Where it is not small

The same question over the room’s own argument gives a different answer, and the numbers are large enough to change what a calculation may do.

An office’s adapting luminance over a working day runs from perhaps eighty to three hundred candelas a square metre. Over that range the gap is 21 per cent of the spread — noticeable and not yet serious. A grading suite’s own settings, eight to forty, give 23 per cent.

One room over a day — twenty in the early morning to two thousand with the sun on the wall — gives 59 per cent. Indoors to outdoors, thirty to twenty thousand, gives 74.

The mechanism is the same one and the excursion is different. The adapting luminance enters the model through a fourth root, so a range spanning three decades spans a large part of a strongly curved function, and averaging over it is averaging across the bend rather than along a nearly straight piece of it.

The corner the hue scale has at each unique hue. How fast hue quadrature runs against hue angle, all the way round. It is piecewise, with a corner at each of the four unique hues, and its slope runs from 0.47 near 238 degrees to 1.92 near 90 — a factor of 4.1. The largest corner is at blue, where the slope changes by a factor of 0.41 across a single point.
Fig. 4 The hue scale’s own unevenness, from the previous rung. It is a different kind of nonlinearity — a table rather than a curve — and it does not participate in any of this, because quadrature has no viewing condition in it.

The three rooms, and which one a specification names

The four ranges measured are not four hypotheses; they are four real situations, and which of them a specification is about is usually not stated.

A grading suite holds its own luminance deliberately, and the range measured — eight to forty candelas a square metre — is the span between a dim reference environment and a slightly brighter one. Its gap is 23 per cent, which is the smallest of the rooms and is not negligible: a facility averaging its own conditions is making an error a quarter of the size of the variation it is averaging over.

An office over a working day runs from perhaps eighty to three hundred, and gives 21 per cent. That is the case a colour management chain is implicitly about when it names an average surround, and the number says that treating the day as one condition costs a fifth of the day’s own effect.

One room from early morning to full sun gives 59 per cent. That is the case a shop, a showroom or a domestic living room is in, and it is where the approximation stops being an approximation: the average of the appearances across the day is more than half a spread away from the appearance at the average light level.

Indoors to outdoors gives 74 per cent and is the honest bound rather than a situation anybody specifies for.

The practical reading is that a specification naming a viewing condition is naming a point, and the point stands for a range. How well it stands for the range is the number above, and it is between a fifth and three quarters depending on which range is meant. No standard states a range; every standard states a point.

What moves and what does not

The luminance sweep says which of the model’s outputs carry the room, and the answer is a clean split.

Across a factor of 2,500 in adapting luminance — eight to twenty thousand — brightness rises by a factor of 5.1 and colourfulness by 2.0. Lightness rises by 2.5 per cent and chroma by 3.3.

That split is by construction rather than by accident. Lightness and chroma are ratios to the white, and the white is being lit by the same lamp as the sample, so most of the change divides out. Brightness and colourfulness are absolute, so none of it does.

So a calculation reported in lightness and chroma is nearly immune to the room and one reported in brightness and colourfulness is not, and the choice between the two pairs is usually made on grounds of familiarity. A colour management chain reports lightness; a lighting calculation reports brightness; and the second is where the averaging question bites — which is one more thing a lamp is measured by and not specified in.

Where a stated appearance has no stimulus under it. The chroma the inverse can still return a light for, at lightness 50, all the way round the hue circle. Below the curve a stated appearance corresponds to a stimulus; above it the formula still returns three numbers and one of them is negative. Over a lattice of 10488 appearances spanning the whole space a specification is written in, 10.8 per cent are of that kind, and only 44 per cent are colours a display could show.
Fig. 5 The model’s own domain, from two rungs back. Nothing about the averaging question depends on it, and a mean that landed outside it would be a mean of things that exist reported as a thing that does not.

One caution belongs here rather than in the limits, because it is about the operation rather than about its accuracy. The mean of several appearances is itself an appearance, and an appearance is not always a stimulus: a mean can land outside the model’s own domain even when every member of the set is inside it, since the domain is not convex in the model’s coordinates. Nothing measured here does so, and a set spread wider than a population’s might.

What a stated lightness pins down, and where. A lightness quoted to 0.05 of a unit, inverted, and the luminance it fixes. Read as a fraction of the colour's own luminance the requirement is 0.90 per cent at J 10 and 0.097 at J 95, a factor of 9.3. Read in absolute luminance it is the other way round, by a factor of 6.5. Both readings are true and they answer different questions.
Fig. 6 What a stated lightness pins down, at each end of the scale. The same compression that makes the averaging question interesting makes a stated precision uneven, which is the last rung of this ladder.

The compression under all of this is one function and it produces four separate findings in this ladder: the empty part of the coordinate box, the uneven pinning of a stated lightness, the amplification of the audit’s small departures, and the averaging gap measured here. Four measurements, one exponent, and the exponent is 0.42.

What was computed, and how

The population is eighty observers drawn with a stated seed from this collection’s own variate model — age, macular density, cone optical density and three peak wavelengths — which is the same population the previous rounds used.

Each observer forms their own tristimulus values from the same spectral power distribution, normalises to their own white, and adapts to D65 through CAT16, which is the audit’s own route. The model is then run at the reference viewing condition and the answers are averaged in CAM16-UCS, where an average is a well-defined operation; averaging the correlates directly would be averaging a hue angle, which is not.

The audit's six departures, read twice. Each departure priced in the matching unit it was published in and in the appearance unit an appearance prediction would be judged by. The model does not scale them by one factor: it amplifies the smallest by 1.58 and the largest by 1.03, so the range between them narrows from 3.34 to 2.17. The mean ranking is unchanged and 14 of the 42 surfaces reorder.
Fig. 7 The audit’s six departures read in the model’s own unit, which is the next rung’s subject. That measurement rests on this one: it is only legitimate because carrying an average through the model is safe.

The rooms are ranges of adapting luminance rather than measured light levels, chosen to span what a room plausibly does. Each is averaged uniformly over its stated values, which is a crude weighting — a real room spends more time at some levels than others — and the qualitative result is not sensitive to it.

The comparison is always the same: the mean of the model’s answers, against the model’s answer for the mean argument, both in CAM16-UCS, reported as a fraction of the spread so that the two cases are comparable.

Where averaging the model's answers is and is not averaging its argument. Each row takes a spread of situations, averages the model's predictions across them, and compares that against the model's prediction for the average situation. The bar is the gap as a share of the spread itself. Over the differences between observers it is 1.4 per cent — the model is very nearly linear there. Over the range of adapting luminance one room covers in a day it is 59 per cent, and from indoors to outdoors 74.
Fig. 8 The five spreads on one axis. The first bar is a population and the last is the world, and the difference between them is a factor of fifty in how far the model’s argument travels.

Putting the five on one axis is the honest presentation and it is the one that makes the finding usable. The population bar is short not because observers agree but because the model is nearly straight over the range in which they disagree, and the world bar is long not because the model is unusually curved but because the world is unusually wide. The bar’s length is a property of the pair — a function and a range — and neither of them alone.

What this says about the round’s own thesis

A round with a thesis should say where the thesis fails, and this rung is that place.

The thesis has been that everything downstream of the tristimulus values is nonlinear and that the previous round’s methods do not survive it. Five essays have found large effects and this one finds a small one, in the argument the previous round cared most about.

So the correct statement is narrower than the thesis. Nonlinearity does not by itself invalidate an average; it invalidates an average taken across a wide enough excursion, and wide enough is a comparison between the spread and the curvature’s own scale. Two of this round’s findings pass that test and one does not.

The two that pass are the mixture line, whose endpoints are as far apart as colours get, and the pricing of a departure, whose excursion is the whole object-colour solid. The one that does not is the population, whose excursion is a couple of colour differences.

That is a better result than a uniform verdict would have been. A round that found the same thing everywhere would have found a property of its own method; a round that finds a large effect in four places and a negligible one in a fifth has measured something about the subject.

Where the model stops

The population case holds one surface at a time and averages over observers. Averaging over surfaces as well — which is what a census does — is a different question and would need the surfaces’ own spread, which is much larger and would give a much larger gap.

The room case holds one stimulus. A real scene has a range of luminances in it and the model would be applied to each with the same adapting luminance, which is a different structure from the one measured here.

And the reference viewing condition’s own arguments are held fixed throughout: the surround is average, the background is twenty per cent, and the degree of adaptation is the model’s own 0.94 rather than one. Moving any of those moves every number.

The generalisation

The habit is about knowing which of a nonlinearity’s arguments is the one that matters.

A model with several arguments is nonlinear in all of them and is not equally nonlinear over the range each argument actually varies. The two facts get conflated: a model is described as nonlinear, and every operation on it is then treated as suspect, when in most of its arguments the range is short enough that a linear approximation is accurate to a per cent.

The move is to measure the gap in each argument, as a fraction of that argument’s own spread. The fraction is the transferable number: it says whether an average may be taken, and it is dimensionless.

The failure mode is not an error but a paralysis. A model treated as uniformly nonlinear is a model nobody may average, aggregate or summarise, which removes most of what a model is for — and the measurement that lifts the restriction in one argument is the same measurement that establishes it in another.

Who found it, and when

Jensen’s inequality is 1906 and its application to averaging model outputs is standard practice in every field that does it. The specific observation that an appearance model is nearly linear over a population’s spread does not appear in the sources consulted here, and its being nearly linear is the sort of negative result that is not usually published.

The strong dependence of the model’s absolute correlates on the adapting luminance is the Hunt effect and the Stevens effect, both measured long before any appearance model existed and both built into CIECAM16 deliberately. What is new here is reading them as a statement about when an average may be taken — which is a mean’s own set asked one level up.

Where the ladder goes next

If an average survives the model, then the previous round’s results can be carried through it as they stand — and the obvious thing to carry is the audit itself. The six observer departures read in the model’s own unit turn out not to be the same six departures scaled by a constant.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AdaptationCIECAM16Colour appearanceConvergenceDeclared inputLuminanceObserver variabilityPopulationSensitivityViewing condition