Where the model breaks

Every threshold was measured with a grating

An earlier essay here claimed that the model cannot explain why dither works, and named two missing pieces. One of them was real and worth thirteen times the guess; the other was not needed. The piece nobody named was the detector — and reading the same model two ways changes the answer by a factor of fifty.

Assumes Banding is not a bit depth and A pattern has a direction.

A phase ago this site asserted an absence. Its spatial model said that adding a noise mask before quantising a gradient makes the banding worse, on five ramps out of five, where every practitioner knows it makes it better — and rather than patching the model, the assertion was written to require the disagreement to persist, so that a later revision could not quietly start agreeing with a textbook it had not earned.

The assertion also named what it thought was missing: a second spatial dimension, and a nonlinearity that prefers incoherent noise to a coherent edge.

This phase built the second dimension. The diagnosis was half right, the half that was right is worth thirteen times what was guessed, and the piece that mattered most was not in the list.

What a dither mask is worth, read as components, in two dimensions. Five luminance ramps, each quantised to 8 bits with and without a high-passed mask of the same power. The bars are the most visible single sinusoidal component of the error, as a multiple of the contrast that component needs to be seen: above the line at one it is visible. The mask lowers it by 20–22×, on every ramp — which the one-dimensional model on this site says it does not, and that disagreement is the finding.
Fig. 1 The measurement that closes it. Five ramps, quantised to eight bits with and without a high-passed mask of the same power, read as frequency components in two dimensions. The undithered shadow ramp sits at 3.4 times threshold and the dithered one at 0.17 — a factor of nearly twenty, on every ramp.

The claim

The model was being asked the wrong question, and the wrong question was not about dimensions.

Every threshold this site quotes was measured with a grating. An observer was shown a sinusoid of one frequency at one contrast and asked whether it was there. A model built on those thresholds is a statement about components, and asking it what happens at a point in space asks it something it was never fitted to answer.

Read four ways, the same mask on the same ramps gives:

reading dimension what the mask is worth
largest departure at a point line 0.54× — worse
largest departure at a point plane 0.54× — worse
largest component line 1.47×
largest component plane 19.7×
  • The point reading says dither is harmful, in both dimensions. Adding a second dimension changes it not at all, which is the finding that corrects the previous phase’s diagnosis.
  • The component reading says it helps, in both dimensions — and by very different amounts.
  • The second dimension is worth 10 to 19 times, across the five ramps, and that is the part of the old diagnosis that was right.
  • And the nonlinearity is not needed. No masking term appears anywhere and the advantage comes out at full size without one.

Why a point is the wrong place to read

A quantised gradient’s error is a staircase: a sawtooth whose period is the code step, confined to one axis because the ramp varies along one axis. Almost all of its energy is in a few frequency components.

A dither mask’s error is noise: the same total power, spread over every component there is.

At any single point in space, the noise is often larger — that is what noise does. So a model that reports the largest departure at a point sees the dithered version as worse, correctly, and irrelevantly. Nobody detects a gradient by looking at one pixel.

Read as components, the same two signals are not close. The staircase’s power is concentrated: in the shadow ramp measured here, its largest component sits at 1.85 cycles per degree at 3.4 times the threshold contrast for that frequency. The mask’s power is diluted across the whole plane, and its largest component reaches 0.17 times threshold — twenty times below.

That is not a different model. It is the same filter, the same quantiser, the same ramps, and a different question.

What the second dimension is actually worth

The old diagnosis said a mask’s power spreads into an annulus in a plane while a band edge stays a line. That is exactly right and it is the reason for the factor of thirteen.

A signal of n samples has n components. A field of n by n samples has n². The mask fills both; the staircase fills a line in either case, because a ramp varying along one axis has no structure along the other. So the mask’s power per component is diluted by a further factor of roughly n in a plane, and the amplitude of each by roughly the square root of n.

ramp on a line in a plane the dimension’s share
0.0008 – 0.006 1.47× 19.7× 13.4×
0.005 – 0.02 0.98× 18.4× 18.8×
0.02 – 0.06 1.48× 19.5× 13.2×
0.05 – 0.12 1.46× 19.6× 13.4×
0.16 – 0.24 2.12× 22.1× 10.5×

Two things about that column are worth extracting, because the range quoted for it comes from an unexpected place.

The plane readings are nearly constant and the line readings are not. Across the five ramps the plane column spans 18.4 to 22.1 — a factor of 1.20 — while the line column spans 0.98 to 2.12, a factor of 2.16. The share is the quotient of the two, so the whole of its 10-to-19 range is the line reading’s instability, and none of it is anything the second dimension is doing differently from ramp to ramp.

That reverses which of the three columns is the measurement. Read as printed, the table says the dimension is worth somewhere between ten and nineteen times depending on the ramp, which invites a search for what distinguishes the ramps. Read by column, it says the plane reading is about twenty times threshold-improvement on every ramp to within a fifth, and the one-dimensional reading — which the whole essay argues is the wrong question — is the erratic one.

And the erratic one is erratic in the direction that would be expected. A one-dimensional mask has n components rather than n², so each carries far more power and the largest one is close to threshold — 0.98 to 2.12 times it. A quantity sitting near a threshold is a quantity whose ratio to that threshold is sensitive to everything, which is why the line column wobbles by a factor of two while the plane column, sitting twenty times clear, does not.

The second thing is that the dimension is worth less than the naive arithmetic predicts. Spreading the same power over n² components rather than n should reduce each amplitude by n\sqrt{n}, and the field is 512 samples square, so the prediction is 22.6. The measured geometric mean is 13.6 — sixty per cent of it.

The gap has a cause and it is the mask’s defining property. A blue mask does not fill the plane; it fills an annulus. Everything inside a stated radius has been removed, so the power is spread over the high-frequency part of the plane rather than all of it, and a smaller region means more power per component than the naive count allows. The forty per cent shortfall is the price of putting the power where the eye is worst — the mask buys its position on the sensitivity curve by giving back some of its dilution, and the two effects are the same design decision seen twice.

That is worth having explicitly because it says what a whiter mask would do. White noise fills the whole plane, so it should recover the full n\sqrt{n} dilution and lose the positional advantage; the essay’s own figures show both masks helping and attribute the difference to position alone. The dilution is part of it too, in the opposite direction, and a comparison of the two masks’ dimension shares would separate them — which is a measurement this machinery could make and has not.

One arithmetic note while the table is open. The claim at the top of this essay puts the two readings a factor of fifty apart; the pair the four-way table actually prints — 0.54 at a point and 19.7 as components — is a factor of 36.5. The larger figure belongs to a different ramp or a different rounding, and the pair the reader can check gives the smaller one.

On a line a mask is worth almost nothing — and on one ramp it is worth slightly less than nothing. In a plane it is worth twenty. assertTheDimensionIsWorthAboutTen holds the ratio between a floor and a ceiling rather than merely above zero, because a change that made it enormous would be as much of a defect as one that made it disappear.

What a dither mask is worth, read as components, in two dimensions. Five luminance ramps, each quantised to 6 bits with and without a high-passed mask of the same power. The bars are the most visible single sinusoidal component of the error, as a multiple of the contrast that component needs to be seen: above the line at one it is visible. The mask lowers it by 16–21×, on every ramp — which the one-dimensional model on this site says it does not, and that disagreement is the finding.
Fig. 2 The same measurement at six bits rather than eight. Everything is further over threshold and the mask is worth the same factor, because the ratio is about where power sits rather than about how much of it there is.

What this does to the rest of the site

The correction is not confined to dither, and this is the part worth carrying away.

Every claim on this site that reads a filtered signal at a point inherits the same question. A difference has no size measured colour differences between filtered rows, position by position — a point reading. Banding is not a bit depth takes the largest response of a filtered error signal — a point reading. Both are answering how large is the departure here, and the thresholds underneath both were measured on gratings.

Neither result is wrong, and the reason is worth stating: both are about structured signals, where the point reading and the component reading agree. A halftone screen, a quantisation staircase and a subsampled edge all have concentrated spectra, so the largest point departure and the largest component are the same object seen twice. The two readings only diverge when one of the signals being compared is incoherent — which is exactly the case the dither comparison is.

So the standing rule this phase adds is narrow and checkable: when a comparison has noise on one side of it, read components.

What the old assertion still says

assertDitherIsNotExplainedHere has not been changed, and it still passes. It measures the one-dimensional model, reading peaks, and that model still says dither is worse — which is true of what it measures.

What has changed is its docstring, which now records what the plane found: that the dimension is real and large, that the nonlinearity is not needed, and that the missing half nobody named was the detector. The correction belongs where the measurement is, and the assertion belongs where it was.

That is a deliberate arrangement rather than an oversight. An assertion whose subject is this file, read this way should not be weakened when a different file reads a different way; it should be joined by a second assertion in the new file, which is what happened. There are now two, they disagree, and the disagreement is the finding.

What was computed, and how

The construction is bandingVisibility’s, step for step, in two dimensions: a luminance ramp, encoded, quantised, decoded, the error taken as a Weber contrast against the local luminance, filtered through the detection function, and compared with the threshold. The only change is that the ramp is a field — varying along x and constant along y — so its quantisation error is confined to a line in the frequency plane.

The sampling rate matters and got this wrong first. The first version computed on a 128-square field over twelve degrees, which is eleven samples per degree — so a blue mask’s power sat between 1.9 and 5.3 cycles per degree, right at the peak of the contrast sensitivity function. At a real display’s 45 samples per degree the same mask sits between 7.9 and 22.5, where sensitivity is a fraction of its peak. The measurement is now made on a 512-square field at 45 samples per degree.

The component reading is spectralVisibility, which takes the largest filtered Fourier amplitude of the error, doubled for the positive and negative frequency pair, as an equivalent contrast — and divides by the threshold at the frequency it sits at. Its one-dimensional counterpart exists so that the two can be compared like for like, and the ratio between them is the dimension’s worth.

The masks carry identical power. White noise, and the same field with everything inside a stated radius removed and the remainder rescaled to the same root mean square. That is what a blue-noise mask is as a spectrum, and it is the only property of one this argument uses — it is not a void-and-cluster mask and does not claim to be.

And the error still rises. The mask raises the root-mean-square error by 1.42× on every ramp, exactly as the previous phase measured. Dither has never been about reducing error; it is about moving it.

Where the model stops

The masks are spectra rather than constructions. A real dither mask is built by a spatial algorithm — void-and-cluster, error diffusion, an ordered matrix — and its defining property is the spectrum this file synthesises directly. What is not captured is that a real mask is a fixed pattern with a tileable structure, and a fixed pattern has correlations a random one does not.

There is no masking, still. A linear model has no term for a pattern being harder to see beside another pattern, and the practical advantage of dither on a picture rather than on a flat ramp is partly that.

The component reading has its own limit. It reports the largest single sinusoid, which is a detection model with one channel per frequency and no summation across them. A signal with many components each just below threshold is predicted invisible and is not.

And the ramps are one-dimensional. A gradient in a picture is also a gradient in colour rather than in luminance alone, and nothing here has a chromatic channel in it. A real gradient also runs at an angle and often in two directions at once, which by the orientation argument changes both the staircase’s visibility and the mask’s.

What a dither mask is worth, read as components, in two dimensions. Five luminance ramps, each quantised to 10 bits with and without a high-passed mask of the same power. The bars are the most visible single sinusoidal component of the error, as a multiple of the contrast that component needs to be seen: above the line at one it is visible. The mask lowers it by 15–20×, on every ramp — which the one-dimensional model on this site says it does not, and that disagreement is the finding.
Fig. 3 Ten bits, where the undithered banding is under threshold on every ramp and the mask has nothing left to do. The factor stays; the question stops mattering, which is one answer to how many bits are enough.
What a dither mask is worth, read as components, in two dimensions. Five luminance ramps, each quantised to 12 bits with and without a high-passed mask of the same power. The bars are the most visible single sinusoidal component of the error, as a multiple of the contrast that component needs to be seen: above the line at one it is visible. The mask lowers it by 10–16×, on every ramp — which the one-dimensional model on this site says it does not, and that disagreement is the finding.
Fig. 4 And twelve bits, where nothing is above threshold on any ramp with or without a dither. Somewhere between six and twelve the question stops being about the eye and becomes one about the encoding, and the crossing is not where a threshold measured on a grating puts it.

The generalisation

The sentence worth carrying: a model inherits the question its measurements were made with.

Every landmark on this site is a psychophysical threshold, and a threshold is the answer to a specific question asked with a specific stimulus. The contrast sensitivity functions were measured with gratings; the colour-matching functions were measured with a bipartite field of a stated size; MacAdam’s ellipses were measured with a specific arrangement and a specific instruction. Using any of them means asking the same question, and the failures on this site have almost all been cases of asking a different one.

The surprising connection is with the arrangement argument. That essay found a colour difference formula giving fifteen times the right answer because it was applied to a pattern when it was fitted to two large patches. This one finds a contrast sensitivity function giving a fiftieth of the right answer because it was read at a point when it was fitted to components. Same failure, two directions: a measurement used outside the arrangement it was made in, with no warning attached to it anywhere, and the size of the error in both cases larger than any of the effects being studied.

That is now the site’s most repeated finding, and it has arrived from four separate directions in three phases.

A grating at 0°, 8 c/°, and where its energy sits. Left, the pattern. Right, its power in the frequency plane with the zero frequency at the centre and the edges at the sampling limit of 23 cycles per degree, on a logarithmic scale over five decades. The closed curves are the visual system's own sensitivity at 5, 25, 60 per cent of its peak; they are not circles, because sensitivity is lower on the diagonals than on the cardinal axes by a factor of 2.0 at high frequency. Energy inside a curve is seen; energy outside it is not, whatever its size.
Fig. 5 What a threshold was measured with: one sinusoid, one frequency, two points in the plane. Every landmark on this site was fitted to an observer looking at something like this, and every claim made from one is a claim about how much of a signal is at some frequency.

Who found it, and when

Dither arrived in the 1950s in the analogue world — literally a vibration added to mechanical computers so their gears did not stick — and reached signal processing in the 1960s, where the theory of subtractive and non-subtractive dither is well developed and entirely one-dimensional.

Blue-noise masks for images are Ulichney’s, from 1987, with the void-and-cluster construction following in 1993. The argument for them was always spectral: put the mask’s power where the eye is worst, and the eye’s own filter removes it.

And the detection model underlying the component reading is older than any of it: the notion that the visual system analyses an image into spatial-frequency channels, and that detection is decided by the most responsive channel, is the framework the contrast sensitivity functions were measured within in the 1960s.

So all three pieces were in place. What this essay adds is a measurement of what each is worth in a case where they disagree, and a correction to a diagnosis this site published one phase ago.

What the pictures cannot show

They cannot show the ramps at the geometry the numbers assume. Every measurement here is computed at 45 samples per degree, which is a reader sitting 60 cm from a display of about a hundred pixels to the inch. The figures on this page are drawn as scalable vectors and will be rendered at whatever size a browser chooses.

And a figure of the dithered ramp would be a lie either way. Rendered small the mask is invisible; rendered large it is grain; and either way the reader’s own display is quantising and dithering the figure a second time. What the figures on this page draw is where energy sits, which is the quantity the argument is about and the only one a scalable drawing carries faithfully.

Where the ladder goes next

The nearest unfinished piece is summation. The component reading takes the largest single component, and a real detector pools across nearby channels — which would make a spread-out mask slightly more visible than this model says and would narrow the factor of twenty. The pooling rule is a standard one and adding it is a small change with a measurable consequence.

The second is a real mask. Everything here uses a synthesised spectrum, and a void-and-cluster mask has a tileable structure with correlations at the tile scale. Whether those correlations cost anything is a question this machinery could answer and has not been asked.

And the third is the audit this essay implies. Every claim on this site that reads a filtered signal at a point should be re-read as components, and the two answers compared. Most will agree, for the reason above; the ones that do not will be the ones with noise on one side, and nobody has made the list.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 11 that link here.

The objects this essay names

Each one links to every other essay that touches it.

BandingContrast sensitivityDisplay gamutDitherJust-noticeable differenceMoireQuantisationSamplingScreen angleSpatial frequencySpecificationThresholdTone curveTransfer function