Theme

The thread: The instrument is the reader — page 6

This is the one subject where the page is displayed on the apparatus under discussion, and the reader's own eye is the measuring device. Several figures here are experiments rather than illustrations.
Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy. Where the model breaks

The instrument named the pair that moved

A standard error over a test set flagged four steps of the adaptation census as unresolved. Rebuilding the set three different ways reversed exactly one pair, and it was one of the four. Rebuilding it to a different rule reversed a pair the error separated by nearly nine standard errors — which is not a failure of the instrument but a statement of what it is about.

What a fourth reflectance dimension costs the theorem that a change of light is a matrix. Four rising curves on axes of the fourth dimension's amplitude, left to right, against what is left of daylight to tungsten after the exact 3×3 change-of-light matrix has been applied, in ΔE₀₀. All four begin at exactly zero: on the three-dimensional family the matrix is solved rather than fitted and there is no remainder at all, which is the theorem this collection's adaptation argument is built on. Adding a fourth reflectance dimension breaks it, and how badly depends far more on the fourth function's shape than on its size — at five per cent amplitude the four shapes cost 0.329, 0.572, 0.063, 0.124 ΔE₀₀ respectively, a factor of 9.1 between the dearest and the cheapest. For scale, the smallest von Kries residual anywhere in the census is 0.26 ΔE₀₀, so the cheapest of the four is a quarter of it and the dearest is twice it. What a scene does

A theorem about a family

A change of light acts on the test surfaces used here as an exact 3×3 matrix with no residual whatsoever, and the whole adaptation argument is built on that being exact. It is exact because the surfaces span exactly three dimensions, and they span exactly three dimensions because three basis functions were written down.

How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one. Difference and uniformity

Saturation is nearly everything

The set of test surfaces has three numbers describing it, and only one of them matters. How saturated the surfaces are carries an elasticity of about 0.7 on every result computed over them; how bright they are carries 0.10. A test chart's chroma range decides its answer and its lightness range does not.

How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one. Where the model breaks

The input nobody declared

An audit that swept every declared width in this collection found the largest elasticity anywhere to be about a half. The most elastic input turns out to be one that was never declared, never quoted with a range and never varied — and being undeclared is exactly why it escaped the audit that was looking for it.

A published residual is a mean, and the worst object in the room costs twice it. Three bars for each of the 14 changes of light in the adaptation census, ordered by how uneven the change is across surfaces. The first bar is the published mean residual. The second is the worst single surface in the audit's published test set. The third is the worst surface anywhere in the region that set is drawn from, found by search rather than by reading a maximum off a lattice. The mean-to-worst ratio runs from 1.90 to 4.02 and averages 2.43, so every published adaptation number has a worst case about twice it that no essay had ever quoted. The gap between the second and third bars is the other finding: a maximum over 125 sampled points understates the region's own maximum by up to 34 per cent. Where the model breaks

A mean is not a worst case

Every adaptation number this collection publishes is an average over objects, and the reader asking whether adaptation will fail them is asking about the object it fails on. That object costs between 1.9 and 4.0 times the published figure, and how uneven a change of light is across objects turns out to be a property of the change rather than a constant.

A published residual is a mean, and the worst object in the room costs twice it. Three bars for each of the 14 changes of light in the adaptation census, ordered by how uneven the change is across surfaces. The first bar is the published mean residual. The second is the worst single surface in the audit's published test set. The third is the worst surface anywhere in the region that set is drawn from, found by search rather than by reading a maximum off a lattice. The mean-to-worst ratio runs from 1.90 to 4.02 and averages 2.43, so every published adaptation number has a worst case about twice it that no essay had ever quoted. The gap between the second and third bars is the other finding: a maximum over 125 sampled points understates the region's own maximum by up to 34 per cent. Matching and measuring

An extremum is still not a sample

Two rounds ago three measurements turned up that took a maximum over a sample of a set and were short by up to a factor of two. The same error was live in a fourth place the whole time, on the set of surfaces every adaptation number is averaged over, and it is short by up to a third.

Every worst surface sits on a number somebody typed. The region the test surfaces are drawn from, in its own two modulation coordinates: a square of allowed depths with a diamond inscribed in it, the diamond being the requirement that the two depths sum to no more than 0.7. The 14 marked points are the worst surface for each change of light in the adaptation census, found by search over the whole region. Every one of them lies exactly on the diamond, and every one is also at the brightest level the region allows — both declared constraints active, on all 14 rows, with no interior maximum anywhere. That is the opposite of what bounding the wall gave: there the worst case turned over at a band width of six nanometres because a narrow band returns too little light, which is physics. Here the worst case is a reading of two numbers. The one constraint that is about the world — a paint's excitation purity may not exceed 0.6 — is slack everywhere: the most saturated surface the region admits reaches 0.459. What a scene does

Every worst surface sits on a declaration

Bounding the wall in a painted room produced a real worst case — the residual turns over at a band six nanometres wide because a narrower band returns too little light. Bounding the surfaces the residual is averaged over produces nothing of the kind, because all fourteen answers sit exactly on two numbers somebody typed and the one constraint that comes from the world never binds at all.

The same claim in nanometres of pigment, where no declared width can reach it. Five horizontal bars on a scale of nanometres, one per published chromatic-adaptation transform, each showing how far the medium-wave cone pigment's absorption peak would have to move for the receptors' own protan confusion point to land where that transform puts it. Zero is the measured peak. The bars run from -10.5 to 18.2 nanometres — in both directions, so two of the transforms want the pigment shorter and two want it longer. Drawn across them is the 25 nm separation between the L and M pigment peaks, which is the whole basis of red-green vision and is not a number this collection declared. The nearest transform asks for a displacement of 30 per cent of that separation, and the span across the table is 28.7 nanometres — larger than the separation itself. No population, cloud or standard deviation appears anywhere in the statement. What the eye does

The claim, in nanometres

For four rounds the claim here has been that every published adaptation transform puts the protanope's confusion point outside any real population of eyes, stated in standard deviations of a population whose widths were declared rather than measured. Restated as a pigment displacement it needs no population at all — and the nearest transform asks the medium-wave cone to move thirty per cent of the way to the long-wave one.

How far each census row moves when the test set's own description does. A grid of bars, one row per change of light in the census and one bar in each row per number that describes the region the test surfaces are drawn from: how saturated they are, how bright, and how far the two modulations may go together. A bar's length is the elasticity — the proportional change in the published residual for a proportional change in that number. Saturation runs from 0.49 to 0.91 and brightness averages 0.104, so a test set's chroma range is nearly everything and its lightness range is nearly nothing. For scale, the largest elasticity found anywhere among this collection's five declared population widths is about a half — and those at least have declared ranges, while these three numbers have never been quoted with one. What a camera does

A chart decides what a camera scores

A camera profile's reported error changes by a factor of five when the test chart's saturation changes, with the camera and its matrix untouched. The elasticity is 0.67 — the same figure, to two per cent, that an entirely unrelated measurement over an entirely unrelated set of surfaces gives.

A gradient of one hue, mapped into a press's gamut 3 ways. Chroma asked for along the bottom, chroma delivered up the side, for a ramp at lightness 55 and hue angle 25°. A colorimetric intent follows the diagonal until the press runs out and is flat afterwards — the flat part is a gradient arriving as a single colour. The perceptual intent is under the diagonal from the start, which is the price of never going flat. What it takes to deliver it

A budget drawn through one hue

The three-stage error budget this collection publishes for a colour-management chain is computed over twenty-four colours of a single hue at a single lightness. The quantity that actually varies with hue — how many distinguishable colours a rendering intent destroys — runs from nothing at all to more than a third, and the hue the budget uses is near the bottom of that range.

Every census row under five constructions of the same test set. A slope chart with 5 columns — lattice, coarse, fine, uniform, natural — and one line per change of light in the census, each line joining that row's mean residual under each construction. Four of the five columns describe the same region of surfaces walked at different densities or against different measures; the last is the clamped, realistic family, which is not linear in its parameters and is therefore answering a slightly different question. The levels move: between the coarse and fine lattices every row shifts by seven to nine per cent, in the same direction, which is a common-mode factor no published residual here has ever carried. The order almost survives. Inside the region exactly one pair crosses, and it is the pair the standard error had already flagged; under the clamped set two more cross, including one the error separates by nearly nine standard errors. The crossing lines are drawn heavy. Where the model breaks

What the audit still cannot reach

Two rounds have now swept every declared width in this collection and one of its structural choices. Three structural choices remain, none of them has a multiplier to sweep, and the reason each resists is different — which makes the list a description of where this kind of audit ends rather than a queue of work.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit. Where the model breaks

A choice with no magnitude

An audit can multiply a width by 1.25 and report an elasticity. It cannot multiply CIEDE2000 by anything. Auditing a structural choice needs a different instrument, and building one shows that six published numbers in this collection each carry a factor of about two of unit-choice — after the change of scale has been taken out.

The adaptation census in six units, calibrated onto one scale. Each line is one of the fourteen changes of light in the adaptation census, drawn across the six units the results could have been published in. Every unit is multiplied by the single factor that best carries it onto ΔE2000 over a reference sample of surface pairs, so the vertical axis means the same thing in every column and a sloping line is a disagreement rather than a change of scale. The levels move by up to a factor of two. More to the point, the lines cross: ΔEok puts 10 of the 91 pairs of rows in the other order, and CAM16-UCS, the only appearance unit here, puts the fewest — 2. What light is

The census in six units

Recomputing every change of light in the adaptation census under six colour-difference formulae, with the scale factor divided out, leaves a table whose levels move by up to a factor of three point seven. The rows that move most are the mild ones, which is the opposite of what a reader would guess and is a property of where each formula was fitted.

How far each unit is from being a rescaling of the one this collection publishes in. One row per unit on the menu. The bar is the root-mean-square scatter about that unit's own best rescaling of ΔE2000, over 374 pairs of surfaces differing by a fraction of a unit to about ten. A bar of zero would mean the unit is ΔE2000 in different money — every printed number would change and no conclusion would. ΔE2000's own row is zero by construction and is the check that the table is computed the right way round. The two units that divide a chroma difference by the chroma it was measured at, ΔE94 at 15 per cent and CAM16-UCS at 24, are closer to it than the three that do not, which run from 28 to 35. The split is by weighting and not by whether the unit is a matching difference or an appearance one. Difference and uniformity

The weighting is the disagreement

Five colour-difference formulae, three decades and two committees, and the single property that predicts which of them agree is whether a chroma difference gets divided by the chroma it was measured at. It sorts the menu exactly, it cuts across the distinction between a matching difference and an appearance one, and it halves the census's largest sensitivity.

The census along the line from ΔE76 to ΔE94, and past it. ΔE94 is ΔE76 with two weighting constants in it, and at zero those constants make every weight exactly one, so the two formulae are joined by a line rather than separated by a choice. The horizontal axis is how much of the published weighting is applied: 0 is exactly ΔE76, 1 is exactly ΔE94, and 3 is three times more weighting than anybody has proposed. The falling curve is the census's mean elasticity to how saturated its test set is, which drops from 1.11 to 0.74 — most of the fall happening before the published value is reached. The other curve is Kendall's τ against ΔE2000's ranking, and it peaks at w = 0.5, not at 1: the weighting that best reproduces the published ordering is about half the published weighting. There is no value of this dial that reaches ΔE2000, whose rotation term is not on this line at all. Difference and uniformity

A dial through a discrete menu

ΔE*94 is ΔE*ab with two weighting constants in it, and at zero those constants make every weight exactly one — so the two ends of the oldest disagreement in colour difference are joined by a line rather than separated by a choice. Walking it gives a derivative where a menu gives only a spread, and the derivative says the published weighting is on the far side of the interesting part.

Which steps of the census ranking a change of unit reverses. Every adjacent pair in the published census ranking that at least one unit puts the other way round. The bar counts how many of the five other units reverse it. The marker on the left says whether the test set had already declared the pair unresolved — a gap smaller than twice its own paired standard error, which is a statement about sampling over 125 surfaces and shares no arithmetic with a change of ruler. The two pairs every unit reverses are both flagged, which is the agreement. The pair at the bottom is the disagreement: the test set resolves it at 9.1 standard errors and four of the five units reverse it anyway, because a sampling error cannot see a change of ruler and a change of ruler cannot see a sampling error. Where the model breaks

Two instruments and one ranking

A sampling error over a hundred and twenty-five surfaces and a change of colour-difference formula share no arithmetic at all, and they were asked the same question of the same table. Every adjacency the whole menu reverses had already been flagged as unresolved. And one the test set settles at nine standard errors is reversed by four of the five formulae, which is what makes them two instruments rather than one.

Where on the scale the units disagree. The reference pairs split into bands by how far apart they are in ΔE2000, with each unit's root-mean-square relative departure from the published one plotted per band. Every unit is calibrated once, over the whole sample, so a band is not refitted and the shape is the effect rather than an artefact of fitting. Every one of the five falls: the disagreement is proportionally largest on the pairs that are closest together, which is the opposite of what being fitted to threshold data would suggest. The appearance unit is the extreme case, at 91 per cent on the narrowest band and 17 on the widest, because CAM16-UCS raises its distance to the power 0.63 and a power below one inflates small differences against large ones. In absolute terms every curve here runs the other way — the widest band disagrees by 1.16 to 2.37 ΔE₀₀-equivalent against 0.14 to 0.68 on the narrowest — so which reading is right depends on whether the published quantity is a level or a ratio. This is the mechanism behind the census's own behaviour, where the mildest rows spread furthest across the menu. Difference and uniformity

The disagreement is at the near end

Every colour-difference formula on the menu was fitted to threshold data, so the expectation is that they agree about pairs an observer can only just tell apart and diverge on large differences. They do the opposite. Proportionally the disagreement is largest at the near end, by a factor of six for the appearance unit, and the cause is an exponent of 0.63.

MacAdam's twenty-five ellipses, measured in each unit. The uniformity instrument used here, applied to units rather than to spaces. The upper bar is anisotropy — the mean over the twenty-five of the largest radius divided by the smallest, where 1 would be a circle. The lower is spread — the largest mean radius divided by the smallest across all twenty-five, which asks whether a step of the same size means the same thing in different parts of the diagram. Reading down the three CIELAB-based formulae in the order they were published, the anisotropy falls 3.42 → 2.89 → 2.74 and the spread rises 3.24 → 3.59 → 4.12: the weighting divides a difference by the chroma it was measured at, which equalises directions at a point and unequalises magnitudes between points. Neither number is scaled, so no calibration is applied here. CAM16-UCS is ahead on both. Matching and measuring

A unit rests on a space that was ranked

This collection ranks three colour spaces by how nearly they make MacAdam's ellipses circles, and CIELAB comes last. It then publishes every difference it computes in a formula built on CIELAB. Turning the same instrument on the formulae rather than the spaces shows the repair works — and that it buys roundness by paying in evenness.

Pairs built to sit exactly on a ΔE2000 tolerance, read in every other unit. Twenty-four pairs of surfaces, each constructed by walking one member along a fixed direction until the difference is exactly 1.0 ΔE2000 under D65. The bar spans what those same pairs read in each unit, after calibration, with the tick at the mean. ΔE2000's own row is a point at 1.0 by construction. Every other unit spreads them: ΔE*uv reads them from 0.69 to 1.75, so a contract written at "one unit" accepts and rejects a different set of deliveries depending on which unit it means. CAM16-UCS rejects all twenty-four: it reads the closest of them at 1.43. What it takes to deliver it

A tolerance is a boundary through pairs

Twenty-four pairs of surfaces built to sit exactly on a ΔE2000 tolerance of one read from 0.69 to 1.93 in the other five units. A contract quoting "one unit" without naming the formula does not become slightly wrong in another; it accepts and rejects a different set of deliveries, and in one case rejects every single one.

Two sensitivities from two libraries, under every unit. Two quantities that share no code, no test set and no physical question: how much the adaptation census's residual depends on how saturated its surfaces are, and how much a camera profile's reported error depends on how saturated its test chart is. The first is a mean over fourteen changes of light built from cosine combinations; the second is one number about one silicon sensor scored on Gaussian bumps. Under the published unit they sit at 0.687 and 0.656. Across the whole menu they move together, from about 0.5 under the appearance unit to about 1.15 under plain CIELAB, staying within 12 per cent of each other at the worst point. Two numbers agreeing once is a coincidence; two curves agreeing at six points across a factor of two and a half is a shared mechanism, and the mechanism is the compression the unit applies to a chroma difference. What a camera does

The coincidence was a mechanism

Two sensitivities from two libraries with no shared code came out two per cent apart, and the claim made about them was that they share a mechanism rather than a number. That claim has a colour-difference formula inside it, so it can be tested by changing the formula — and both curves move together across the whole menu, from 0.5 to 1.15.

A camera matrix refitted to minimise each unit, rather than solved in XYZ. Every camera profile here, and as far as can be told every camera profile anybody ships, is a linear least-squares solve in XYZ. That is an objective and it is on nobody's menu: it weights a difference by how large the tristimulus values are. Each row here refits the same 3×3 by direct search to minimise one of the six units instead. The upper bar is how much better the fit gets in that unit; the lower is how far the matrix itself moves, as a relative Frobenius norm. Both matter and they do not agree: CAM16-UCS moves the matrix least, at 0.34 per cent, for the largest improvement of the six, while ΔE*94 moves it 6.8 times as far for less. A score that changes is a report changing; a matrix that changes is the camera rendering different pixels. What a camera does

The objective nobody chose

Every camera profile here, and as far as can be told everywhere, is a linear least-squares solve in tristimulus space. That is an objective and it is on nobody's menu — it weights an error by how bright the patch is. Refitting the same matrix to minimise a real colour-difference formula improves the fit in all six, and moves the matrix, which means different pixels rather than a different report.

What six of this collection's published numbers do when the unit changes. Six quantities, from six calculations that share nothing: a change of light after an observer has adapted, a camera profile's error, the gap between the two standard observers, a metameric pair under the lamp that breaks it, the same image on two papers, and an observer two seconds into a new room. Each is recomputed under all six units and every unit is calibrated onto ΔE2000's scale first, so the bar is not a change of units in the ordinary sense. The bar is the ratio of the largest reading to the smallest, and it runs from 1.71 to 2.30. Five of the six are printed in ΔE2000 by the essays that report them; the sixth is printed in CAM16-UCS, because the model it comes out of defines that unit. What the brain does

A model judged in another model's unit

An appearance shift is a change in what an observer would report, not a change in a stimulus, so measuring one with a matching difference means first asking what stimulus a settled observer would need to be shown to give the same report. That step is not bookkeeping — it is the whole distinction the field rests on, and it costs a factor of 1.7 across the menu.

What an observer is left with, by how much it is allowed to know about the room. Six ways of discounting a change of light, averaged over the fourteen changes in the adaptation census and 125 test surfaces each. The bar is what each leaves behind, on a logarithmic axis because the models span two orders of magnitude. The second line under each name is the count that matters: how many numbers about this room the model has to be given. Doing nothing leaves 15.7 ΔE₀₀. A single gain read off the two whites' luminances leaves 15.3. A matrix fitted across half the census and then applied everywhere, knowing nothing about the room at all, leaves 12.5. The published von Kries gain, which is told the white and nothing else, leaves 1.312 — and bolting a fixed correction onto it, at no cost in scene information, leaves 1.368, which is very slightly worse. The exact matrix leaves nothing and is not on the chart: its nine numbers are the change of light, which is the quantity being discounted. What a scene does

Three numbers the scene supplies

An adaptation model's parameters are not all the same kind of thing. Some are numbers an observer must estimate from the room it is standing in; others could have been settled once by evolution. Counting them separately turns the diagonal gain from a crude approximation into the only model of the set that gets a large answer from information the observer can actually have.

How much of the residual a partial correction removes. Between the diagonal gain and the exact matrix there is a line: apply the correction that would make a row exact, but only a fraction of it. The horizontal axis is that fraction and the vertical is the share of the row's residual it removes, for all fourteen census rows. The straight diagonal is where a correction worth exactly its fraction would fall, and in the published unit every curve lies on it to within 2.2 percentage points. The lower band of curves is the same interpolation measured in CAM16-UCS, which departs by up to 17 points — because its distance is a power of the Euclidean one and a power is not homogeneous along a ray, where every ordinary norm is. The straight line is therefore a property of the ruler rather than of the correction, and the exception is what says so. What a scene does

A partial correction is worth its fraction

Between a diagonal gain and the exact matrix there is a line, and a bounded observer's natural hope is that the first part of it is worth a disproportionate share. It is not. On all fourteen changes of light, at every setting, the share of the residual removed matches the share of the correction applied to within 2.2 percentage points — which closes the last way the gap could have been cheap.

258 essays on this thread, page 6 of 11 · all threads · all essays