Where the model breaks

The audit that changed the object

Three rounds audited this collection's numbers, its sets and its conventions, and each one ended by naming the same thing it could not reach — that a surface is a reflectance. This round reached it, found that four departures share one algebraic form, had two of its predictions refused, and discovered that its own hemisphere quadrature had been passing a convergence check by luck.

Assumes The model has six arguments, Either factor being zero and A departure is not a unit.

Two rounds ago this collection audited every declared number it rests on. One round ago it audited three conventions. Both ended by writing down the same boundary, and the second one wrote it in a sentence: the decision that a surface is a reflectance rather than a bidirectional distribution.

The six arguments a surface's response has, and the one this model keeps. A surface's response to light is a function of six arguments: the wavelength, direction and place the light arrives with, and the wavelength, direction and place it leaves with. The model every colour here is computed from keeps one number per wavelength, which means it takes the diagonal of the first pair, integrates the second away, and assumes the third pair equal. Each departure drawn here restores one of them. The fourth departure is not on the diagram: the wavelength grid is the range of the index that was kept rather than an index that was dropped, which is why it is the cheapest of the four to fix and was still not fixed.
Fig. 1 The decision, opened. A surface’s response has six arguments and the model keeps one — a diagonal of the first pair, an average over the second, an assumption about the third.

The claim

The reflectance model is a projection of a six-argument function, and every departure from it is a pairing of two deviations.

  • Four departures were measured on one footing, in ΔE₀₀, on stated samples under stated lights: 6.98, 6.70, 1.96 and 1.00.
  • All four have one algebraic form. Each is an inner product of something the sample does with something the light does, and either factor being zero makes it exactly zero.
  • Two predictions were refused. The departure’s size is not the product of the two factors’ sizes, and two departures on one sample do not compound.
  • The machinery caught a defect its own convergence check could not, and the defect was in the check rather than in the model.
  • And the instrument does not reach far. Two of six published quantities admit a departure where all six admitted a unit.

The four, and what each cost

Three of the four restore an argument the model dropped and the fourth widens the range of the one it kept.

The wavelength index is a matrix whose diagonal everybody calls the reflectance, and on a coated printing paper measured with and without the ultraviolet of D50 it is worth 6.98 ΔE₀₀.

The range — the 380-nanometre lower bound of every integral on this site — is worth 6.70 on the same paper, and exactly nothing on anything unbrightened.

The place index is a kernel over the plane, and a pigmented plastic through a four-millimetre radius against an infinite one is worth 1.96. The same measurement on marble is worth 12.65.

The direction index is a function of two directions, and an eggshell paint at a window against the same paint under an overcast sky is worth 1.00. Under a lamp on a stand it is 3.01.

The ladder is a ranking of four examples, and every bar can be moved by a factor of ten by choosing differently. What survives that objection is that all four exceed the tolerance a specification is written in, on samples nobody would call unusual.

The form all four share

Each departure can be written as a pairing:

departure = ⟨ what the sample does that the model cannot hold , what the light does that the model assumed away ⟩

and it is an identity rather than a first-order approximation. Nothing is expanded, no small parameter appears, and the two sides agree to a part in a thousand billion where they share a quadrature.

Three consequences follow, and the second essay of the round works through all of them. Either factor being zero makes the departure exactly zero, which turns four departures into eight conditions. It is exactly bilinear, so doubling either factor doubles it. And the conditions are not all the same kind of statement: three of the eight are identities in floating point and five are limits, each of which is checked by pushing it and requiring the residual to fall.

That distinction was not expected and is the round’s quietest useful finding. A limit and an identity are indistinguishable when the numbers are small, and one is a statement about the world while the other is a statement about a construction.

Eight conditions under which the model equation is exact, and how exact each one is. Each of the four departures vanishes if either of its two factors is empty, which is eight conditions. The axis is logarithmic in the residual that is left when the condition is imposed. Three of the eight are identities: the fluorophore's loading is zero so the emitted term is an empty sum, and a Lambertian surface or a uniform field makes the pairing's second argument identically zero. The other five are limits — a Gaussian excitation band has no edge, an opaque sample still has a kernel a few microns wide, a four-metre aperture is still finite, and the observer is small rather than absent at 380 nanometres. Each limit is drawn with the sequence its residual falls along as the condition is pushed, because a small number is not evidence of a limit and a falling sequence is.
Fig. 2 The eight conditions, on a logarithmic axis. The three round marks are identities; the five sequences are limits, drawn as the path each residual falls along.

Two predictions the arithmetic refused

Each was written down before the number existed, and each is recorded where it was made rather than deleted afterwards.

That the departure’s size would be the product of its two factors’ sizes. It is not, and the failure is the one an inner product always produces when mistaken for a product: a pairing depends on the alignment of its arguments, and two deviations of a given size pair to anything between plus and minus their product. Across twenty-four cells the fraction of the bound actually used runs from −0.101 to 0.365, with six cells positive and eighteen negative. A viewing booth reads a gloss sample high and a window reads the same sample low, at non-uniformities a factor of three apart, because one lights from where the lobe points and the other does not.

That two departures on one sample would compound. They cancel. An aperture removes body light that travelled too far; an interface returns light that never entered the material; and on all six materials the two together cost less than the two apart — by between 118 and 286 per cent of the smaller of the two. Measuring either alone overstates what both do, which is a budget wrong in the least useful direction.

Two for two, in two different directions, on two unrelated questions. The previous round’s rate of correct prediction about its own subject was zero out of four; this one is zero out of two, and the pattern is now worth stating as a habit rather than as a coincidence: the round’s predictions are about the shape of an answer, and the shape is exactly what the arithmetic is for.

Which way the refusal runs

The twenty-four cells are six roughnesses in each of four non-uniform fields, and the sign is not scattered across them. Every one of the booth’s six cells is positive and every one of the other eighteen is negative — window, lamp and sun-and-sky, all six roughnesses each, without an exception. The sign of the error is therefore decided by the field alone, and decided with certainty: knowing which of the four lights a surface stands in predicts the direction of the reading on twenty-four cells out of twenty-four, while knowing the surface predicts nothing at all.

That is worth separating out, because six positive and eighteen negative reads like a scatter and is not one. What a pairing refuses is the claim that a departure’s size follows from its two magnitudes; it does not refuse structure. The structure here is that the alignment belongs to one factor and the magnitude to the other, which is the identity restated as an experiment rather than as algebra.

And the direction of it is the opposite of the intuitive one. The booth is the least non-uniform of the four fields — 0.212 against the window’s 0.665 and the lamp’s 0.931 — and it produces the smallest products of the two deviations, from 0.003 to 0.080 where sun-and-sky reaches 1.82. It is nevertheless the field that uses most of the bound: 0.365 at its peak, where the lamp reaches 0.066 and sun-and-sky 0.025. A booth is built to put light where a gloss lobe will send it into the aperture, so the alignment is not an accident of the arithmetic; it is the design, measured.

The fraction of the bound used also peaks in the middle of the gloss range in all four fields — at a roughness of 0.2 in three of them and 0.1 in the fourth — and falls away at both ends. A near-mirror at 0.02 and a near-diffuser at 0.6 both pair badly. The booth’s six cells run 0.056, 0.137, 0.252, 0.365, 0.277 and 0.015, a rise and fall of a factor of twenty-four across a range of roughness the eye would call all gloss. Neither limit is where the field’s shape reaches the reading, which means the surface most exposed to where the light happens to be standing is the semigloss in between — most of the paint in a house.

What the refusal left standing

A prediction can fail in more than one way, and which way it failed decides what is left. Ranking the twenty-four cells by the product of the two magnitudes, and separately by the size of the error each produced, Spearman’s coefficient between the two orderings is 0.845. On twenty-four cells that is t of 7.4 on twenty-two degrees of freedom, which no arrangement of chance supplies. The same coefficient between the error and the fraction of the bound used is −0.18, which is nothing.

So the product is a bad equality and a good ordering. It does not give the error, it does not give the sign, and it very nearly gives the rank. The division is clean: the two magnitudes carry the size, the alignment carries the sign, and mistaking a pairing for a product is a mistake about the second rather than the first.

The scale error is worth seeing beside the rank agreement, because the two coexist. Sun-and-sky’s largest product is 1.82 and buys an error of 0.029; the booth’s largest is 0.080 and buys 0.014. A predictor twenty-three times smaller produced an error half the size, and both cells sit where the ranking says they should. That is what a rank statistic is for, and it is also the whole of what the product may be trusted with.

This changes what the refusal is worth to whoever is budgeting. A product cannot be used to say how large a departure will be. It can be used to say which of two situations is worse, which is the question a specification usually asks. The claim that failed was the stronger one and the useful one survived it, which is not what a refusal ordinarily leaves behind.

The cancellation is stronger than “partly”

The second refusal has more in it than the sentence that recorded it. If two departures together cost less than the two apart by more than the smaller of them, then they cost less than the larger one alone — and the recorded cancellations run from 118 to 286 per cent of the smaller, so that holds on all six materials without exception. Marble’s aperture is worth 12.65 ΔE₀₀ and its interface 1.33, and the two together are 11.13. Candle wax: 17.55 and 1.12 apart, 15.46 together.

On two of the six it goes further still. Wall paint’s two departures are 1.155 and 1.118 and together they are 0.944; opal plastic’s are 1.961 and 1.866, together 1.538. On those two materials, adding a second departure to the first reduces the error, so the pair costs less than either member of it does by itself.

The practical form of that is a rule about budgets rather than a curiosity. The sum of two departures is not an upper bound on the pair, and neither is the larger of the two: the honest bound is the larger, less between 18 and 117 per cent of itself. A budget assembled by adding measured departures together overstates the total on every material here, and one assembled by taking the worst of them still overstates it.

What the machinery caught that the prose had wrong

The hemisphere is integrated by Gauss–Legendre in the cosine and a uniform grid in the azimuth, and the grid was set at forty by ninety-six on the strength of the standard test: refine it and see whether the answer moves. It did not move — forty by ninety-six and ninety-six by two hundred and fifty-six agreed to a part in ten thousand on the sharpest lobe in the file.

Both were wrong and they agreed because they were wrong identically. Each put a node at the azimuth the light arrived from, so each sampled the narrow lobe at its peak. Rotating the light to an azimuth of 137 degrees moved the answer by 2.5 per cent, on a quantity that cannot depend on azimuth at all, because every surface in the file is isotropic by construction.

The lesson is not about quadrature. A convergence check that refines one knob tests one knob, and what found this was a symmetry the model has and the grid does not. Every module in this round now has such a check: the kernel’s profile is integrated and compared against a total derived separately; the pedestal’s fast path is compared against the slow one; each pairing is computed twice.

What was built

Three modules, each a departure made computable, and each carrying its own assertions.

The first implements the dipole approximation to the diffusion equation, whose closed-form profile and closed-form total are two derivations of one number and are required to agree. A second implements a Lambertian body under a microfacet lobe, from which the two named instrument geometries fall out rather than being declared — and the pedestal between them turns out not to be a constant. The third puts the four on one footing, imposes each condition twice, and draws the boundary of what the instrument reaches.

Two findings came out of the building rather than out of the design. The interface’s map from internal to measured reflectance is a Möbius function rather than a pedestal, and mixing paint in the reported variable rather than the internal one costs between 3.35 and 7.86 ΔE₀₀ — on this collection’s own figures, from a convention nobody ever stated. And the same boundary constant appears in both new modules, which is asserted rather than arranged: what the dipole needs to place its virtual source is what Saunderson calls k₂.

What did not change

An audit that found everything exposed would be an audit of its own instrument, so the negative results matter.

The collection’s census barely moves. Replacing every one of its hundred and twenty-five test surfaces with what an instrument with an aperture reports moves the residuals by at most 5.0 per cent; with what a room reports, 9.3. A departure that does not depend on the light is largely absorbed by an observer’s own gain, because the gain is applied afterwards.

And the departure that does depend on the light cannot be handed to the census at all, because its content is that the sample has no reflectance. That refusal is asserted rather than described, so the boundary is one the build maintains.

The collection's adaptation census, with its surfaces departed. Each row is one of the fourteen changes of light in this site's adaptation census, and the bar is what a von Kries gain leaves behind. The open marks are the published numbers; the filled ones are the same computation with every one of the hundred and twenty-five test surfaces replaced by what an instrument with an aperture, or a room with a direction in it, actually reports. Nothing moves by more than 9 per cent. A departure that does not depend on the light is very largely absorbed by the observer's own gain, because it changes the reflectance and the gain is applied afterwards. The fourth departure is not on this chart and cannot be: a fluorescent sample has a different curve under every light, so there is no set of reflectances to hand the census at all.
Fig. 3 The census with its surfaces departed. Two of the four can be pushed through and neither moves anything by ten per cent; the fluorescent one has nothing to hand over.

What each essay of the round contributes

Twenty essays is a large slate for one argument and it is worth saying what the divisions are, because a reader arriving at any one of them should know where it sits.

Four set out the frame. The model has six arguments states the projection; either factor being zero gives the pairing form and the conditions; a departure is not a unit draws the boundary of the instrument; and this one closes.

Eight build the two new departures. The kernel, the aperture, the hue reversal, the two-aperture symmetry, the field, the boundary’s Möbius map, the mixture variable, and the pair no observer separates.

Four measure consequences elsewhere in the subject: what a camera does differently, what a chart’s substrate does to a profile, what the eye does with an object’s own blur, and what an appearance model cannot hold.

And four are about the audit rather than the physics: the spectral departures’ shared form, the grid’s separate status, the interaction, and the comparison against a tolerance.

What the round is worth as a method

Three shapes are transferable and none of them is about colour.

A model with a missing slot is not approximate. A dropped argument cannot be wrong by a small amount, because it cannot be wrong by any amount, so it does not show up as a residual anywhere. The way to find one is to write the fuller object down and integrate it deliberately, which turns an assumption into a weight that can be varied.

Rank the corrections by the cost of removing them, not by their size. The two largest departures here are removed by specifying a lamp and by widening a table; the two that need hardware are the two smallest. That reversal follows from the pairing: removing a departure means zeroing either factor, and some factors belong to the apparatus while others belong to the world.

And an audit’s cost is set by where the change enters. A metric applied to the answers is an afternoon; an object handed to the first stage is a rewrite of everything downstream, and the cost appears as stages that refuse the new object rather than as difficulty in the arithmetic.

Four departures from the model equation, each at an ordinary strength. What each of the four assumptions inside a colour integral costs, in ΔE₀₀, on a stated sample under a stated light. The wavelength index is a coated printing paper measured with and without the ultraviolet of D50; the range is the same paper integrated from 300 nanometres and from 380; the place index is a pigmented plastic through a four-millimetre radius; the direction index is an eggshell paint beside a window. The spread is a factor of 7.0. This is a ranking of four examples rather than of four departures — each of them can be made larger by choosing a more extreme sample, and the marble in the same collection of materials reaches 12.7 on the index that comes third here.
Fig. 4 The four at ordinary strengths, which is the round’s headline table and is a ranking of four examples rather than of four departures.
An aperture and a gloss lobe, apart and together. Six materials, each measured through a four-millimetre radius and each given a gloss lobe, alone and at the same time. The pale bar is what the two cost added together as if they were independent; the dark one is what they cost when both are present. Every material comes out below the sum, by between 0.8 and 3.6 ΔE₀₀. The two departures partly cancel: the aperture removes light that went into the material and came back out too far away, and the interface returns light that never went in at all. Measuring either one alone therefore overstates what both together do, which is the opposite of the way interacting errors are usually assumed to behave.
Fig. 5 The second: two departures on one sample, apart and together, with the cancellation larger than the smaller of the two on every material.

The first two of those are results about the object. The last two are results about the instrument, and an audit that cannot state its own boundary is one whose reach a reader has to guess.

Which of the collection's published quantities a departure can be pushed through. The six quantities the previous round recomputed under six different colour-difference units, and whether the same treatment works for a departure. Two do: the adaptation census and the metameric pair both take reflectances and a light, which is what a departure acts on. Four do not, and the reasons are different in each case rather than a single obstacle. A unit is a function applied to the answers, so it can be swapped at the end of any computation; a departure changes the object at the start, so it has to be accepted by every stage in between. That is the practical difference between auditing a convention and auditing a structure.
Fig. 6 The boundary of the instrument, which is narrower than the previous round’s and is narrower for a structural reason.
The pairing against the direct computation, for each departure that admits both. Each departure can be computed twice: directly, by taking the difference between the fuller model and the integral one, and as a pairing — an inner product of the sample's deviation with the light's. The bar is how far apart the two answers are, relative to the answer, on a logarithmic axis. The three directional rows agree to a part in a thousand billion, which is the arithmetic of one shared quadrature. The lateral row agrees to three parts in a hundred thousand, and the gap there is the radial quadrature rather than the identity: the two integrals are taken over different grids. The pairing is not an approximation to the departure. It is the departure, written so that its two factors are separate.
Fig. 7 The identity checked against the direct computation, which is the round’s foundation and is asserted at two different precisions for two different reasons.

What is left, measured

Four things, and the round’s standard is that a deferral rests on a measured number rather than an assumed one.

The grid is still 380 nanometres. It costs 6.70 ΔE₀₀ on a brightened stock, it is the cheapest of the four to fix — a table lookup and a longer array — and it is the one still unfixed, because widening it moves every spectral integral on the site and each of those movements has to be measured rather than assumed.

The radiosity solver still assumes Lambertian surfaces, which is not a parameter but a requirement of the method: a radiosity solution rests on radiance being independent of direction. Every scene result in this collection — the corner, the bounce, the green wall — is computed that way, and giving those walls a real lobe is a different algorithm.

Five of the six interaction pairs are uncomputed. One was measured and refused a prediction; the most likely of the remaining five to be interesting is fluorescence with an aperture, because emitted light has its own kernel and its emission band is blue, which scatters further than it arrived.

And nothing here is measured. Every material is a construction from stated coefficients, every field is a stated shape, and the fluorophore is a two-band caricature. The structural results — the pairing form, the conditions, the two refusals — do not depend on the numbers; the numbers do, and a real translucent sample measured at four apertures would settle in an afternoon what this round can only compute.

What it cost to build

The three audits before this one had a pattern in their costs and this one breaks it, which is worth recording for whoever plans the next.

The unit audit was mostly re-implementation: six published quantities had to be recomputed under six formulae, and each re-implementation had to reproduce its own source exactly. The set audit was mostly instrument design: a set has no magnitude, so four separate ways of interrogating it had to be invented. Both were expensive in the way an audit is normally expensive — the arithmetic was cheap and deciding what to vary was not.

This one was expensive in a third way. Deciding what to vary took an hour, because the six arguments of a response function are standard and the departures each have a literature. Building the machinery took most of the round: three modules, two of them implementing physics this collection had never had, each needing its own cross-checks before any of its numbers could be quoted.

The pattern is that auditing an object means acquiring the object, and the object here was three models from three different fields. That is not a cost an audit of a convention ever incurs, and it is the reason this round produced fewer numbers and more machinery than the two before it.

Where the ladder goes next

The list of six arguments has two more entries than this round audited. Polarisation is one: the interface’s reflection is polarised and the body’s return is not, which is why one measurement condition removes the interface optically rather than by an angle. Time is the other: a brightener is used up while it is being measured, so the response has a clock in it and is not linear in the light — which is the one departure on this site that the pairing form explicitly does not hold for.

Beyond those the next object is the observer. It is three curves, it enters at the beginning like a sample does, and this round’s own experience says that is the expensive kind of audit. What makes it the obvious candidate anyway is that its alternatives are already here: ten-degree functions, physiological fundamentals, and a population of two hundred. The substitution exists and only the plumbing is missing, which is exactly the situation the three reachable departures were in before this round started.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AuditBilinearityFalsificationMarginalisationModelling assumptionPredictionQuadratureReflectanceStructural choiceUncertainty