How many colours are there
Assumes How far apart are two colours and A threshold is not a unit.
Two numbers get quoted in answer to this question and neither of them answers it.
16.7 million is . It counts the distinct values a 24-bit integer can hold, and it is a fact about a file format. A display with eight bits per channel can be commanded in 16.7 million ways; how many of those commands produce visibly different results is a completely separate question, and the answer is far fewer.
Ten million is the number usually offered as the perceptual one. It is a volume divided by the volume of a just-noticeable difference, and it therefore inherits whatever the difference formula says a just-noticeable difference is. The formulae disagree.
The count is a volume divided by a cell size, so both of its arguments can be moved. Moving the gamut first:
Pushing the gamut further makes the same point more sharply, since the two formulae were fitted on data that never went out that far.
The other argument is the lattice the count is taken over, and a number that moved with it would be a number about the lattice rather than about the eye.
Two more coarse counts say that the lattice is not what is producing the disagreement between the two formulae.
The widest of the three, counted the same coarse way, is the last of the six and the one where the two formulae are furthest apart.
The computation, and its one honest limit
The recipe is simple and every step of it is a choice.
Take a gamut — here sRGB, a solid in CIELAB whose volume this site computes two independent ways and requires to agree. Walk a lattice through it. At each point that is inside the gamut, ask how large a region around that point is indistinguishable from it, and add one over that volume, times the cell volume.
That integral is the count. Under a formula whose just-noticeable difference is the same size everywhere, it reduces to volume divided by a constant. Under a formula whose JND changes size and shape from place to place, it does not, and the difference between those two cases is the entire result.
The honest limit is stated before the numbers rather than after: this assumes perfect packing. No arrangement of cells fills space with non-overlapping ellipsoids at unit density, so both counts are upper bounds. They are upper bounds on a quantity nobody can measure directly, computed under two models that disagree, and the useful content is the ratio rather than either magnitude.
Why the two formulae disagree
ΔE*ab, the 1976 formula, is Euclidean distance in CIELAB. Its just-noticeable difference is a sphere of radius 1, everywhere, in every direction. That is a strong claim — it says CIELAB is perceptually uniform — and it is the claim CIELAB was built to make.
ΔE2000 is CIELAB distance with three weighting functions applied, plus a rotation term. It says that a unit of lightness difference at is not the same as at ; that a unit of chroma difference means less far from the neutral axis than near it; that hue and chroma differences interact; and that in the blues the whole thing is rotated.
Every one of those is a correction away from uniformity, and every one of them says CIELAB overstates some differences. So ΔE2000’s JND ellipsoids are, on average across the gamut, substantially larger than ΔE76’s unit spheres — and larger cells means fewer of them.
The factor of 4.68 is what that averages out to. It is not a subtlety and it is not a correction; it is most of the answer.
Which formula is right
Neither, and the question is the wrong one.
ΔE2000 is a better predictor of perceived difference than ΔE76 on the datasets it was fitted to, and on most datasets since. It is the formula an industrial tolerance should be written in, and this site’s own tolerance essay uses it. In that sense it is right and the 1976 formula is superseded.
But a count of distinguishable colours is not a prediction of perceived difference. It is an integral over a whole solid, and it is dominated by the regions where the two formulae disagree most — which are the regions where the fitting data is sparsest, because experimental effort went where the interesting cases were rather than where the volume is. So the count inherits the formula’s behaviour in exactly the places the formula is least well constrained.
The right conclusion is that the count is not a well-posed quantity, not that one formula computes it better. A number that changes by a factor of five depending on a modelling choice is a number about the modelling choice, and reporting it without the choice is what makes “ten million colours” a piece of folklore rather than a measurement.
It is worth noticing how the folklore survives. The figure is repeated because it is the only figure available, it is never recomputed because recomputing it requires choosing a formula and a gamut and stating both, and nobody who repeats it is claiming precision — it functions as “a great many”, and as that it is harmless. It becomes harmful the moment it is used comparatively, as in “this display shows a billion colours”, which multiplies a code count by a marketing factor and compares the result against a perceptual estimate.
The measurement behind the disagreement
Neither formula is a theory. Both are fits to psychophysical data — experiments in which people were shown pairs of samples and asked whether they could tell them apart, or asked to judge which of two pairs differed more.
MacAdam’s 1942 experiment is the ancestor: one observer, twenty-five reference chromaticities, and a great many matches at each. The ellipses vary in size by a factor of about twenty across the diagram and their orientations rotate systematically. Every uniform colour space and every difference formula since is an attempt to find coordinates in which those ellipses become circles of equal size, and none has succeeded.
That is worth stating plainly, because the count above is a consequence of it. If a perceptually uniform space existed, the number of distinguishable colours would be well defined — it would be the volume in that space, in units of the JND, and the formulae would agree because there would be nothing left to disagree about. The formulae disagree precisely to the extent that no such space has been found.
What the number is good for, and what it is not
The count is not useful as a count. It is useful as a comparison, in three specific ways.
Comparing gamuts. The same computation on P3 and Rec. 2020 gives larger numbers, and the ratios between them are more meaningful than the absolutes — because the packing assumption cancels. The metric choice does not, and a later section measures how badly.
Comparing bit depths. Eight bits per channel gives 16.7 million code values against a distinguishable count in the tens or hundreds of thousands, which sounds like ample headroom and is not, because the code values are not distributed the way the distinguishable colours are. The transfer function decides that distribution, and the reason eight-bit gradients band in the shadows is that the codes run out there while the eye’s discrimination does not.
Comparing metrics. Which is what this essay is actually about, and the only one where the number is the point rather than the vehicle.
The comparison that was supposed to cancel
The packing assumption cancels in a ratio between gamuts, being one constant multiplying both counts. The metric choice does not, because it is not a constant.
| gamut | ΔE*ab | ΔE00 | metric factor |
|---|---|---|---|
| sRGB | 195,720 | 41,819 | 4.68 |
| Display P3 | 293,680 | 51,561 | 5.70 |
| Rec. 2020 | 444,911 | 62,620 | 7.11 |
The metric’s factor is 4.68 on sRGB and 7.11 on Rec. 2020: it grows with the gamut. So the ratio between two gamuts — the quantity the comparison was supposed to make safe — depends on the formula as well. Rec. 2020 holds 2.27 times as many distinguishable colours as sRGB under the 1976 formula and 1.50 times as many under ΔE00. That is not a rounding disagreement about a headline; it is a factor of one and a half on the answer to how much is a wide gamut worth.
Where the two formulae part company
The reason is visible in the local cell volume, and it is almost entirely about chroma.
Comparing the two formulae’s ΔE = 1 ellipsoids point by point: at the neutral axis the ΔE00 cell is smaller than the unit sphere — 0.72 of it at a lightness of 55, 0.88 at 30, 0.98 at 80 — so near grey the newer formula counts slightly more colours rather than fewer. The ratio then climbs steeply with chroma: about 2.7 at a chroma of 20, 5.3 at 40, 8.6 at 60, 17 at 100 and 26 at 130.
So the 4.68 is an average over a solid in which the local disagreement runs from under one to over twenty-five, and where that average lands is decided by how much of the solid is saturated. sRGB is mostly not saturated. Rec. 2020’s extra volume is almost all at high chroma, which is exactly where the ΔE00 cell is tens of times the larger. A wider gamut adds its volume precisely where the two formulae disagree most, so the disagreement grows with the gamut by construction rather than by accident.
The hue dependence is smaller and points at the term it should. At a lightness of 55 and a chroma of 60 the ratio runs from 6.10 at yellow round to 12.50 at blue — a factor of two around the circle, with its maximum where CIEDE2000’s rotation term does its work.
What survives
Two of the three uses come through intact and one has to be qualified.
Comparing bit depths is unaffected. It sets a count against a count of codes under one metric, and the metric is the same on both sides of every such comparison.
Comparing metrics was the point rather than the vehicle, and it is strengthened rather than weakened: the factor is not one number but a field running from 0.7 to 26 across the solid, and 4.68 is what that field averages to under one particular gamut’s weighting. Quoting it as a property of the two formulae is the same slip as quoting ten million as a property of the eye.
Comparing gamuts needs its metric named, exactly as the count itself does. Saying that Rec. 2020 holds more distinguishable colours than sRGB is safe. Saying how many times more is one more number about a modelling choice, wearing the clothes of a measurement.
What was computed, and how
The local JND volume comes from the metric tensor, estimated numerically from this site’s own ΔE2000 implementation rather than by reading the formula’s weighting terms off by hand.
The reason is the rotation term. ΔE2000 includes a term that couples chroma and hue differences in the blues, which tilts the JND ellipsoid off the CIELAB axes. An estimate that took the three semi-axes along , and and multiplied them would be computing the volume of an axis-aligned ellipsoid, ignoring the tilt, and would overstate the volume and therefore understate the count — in exactly the region the formula’s authors added a term to fix.
So the tensor with is recovered from six probes: three along the axes for the diagonal entries and three along the bisectors for the off-diagonal ones. The ellipsoid volume is then , which is exact to second order and handles the tilt automatically.
Two assertions guard it. The tensor must be positive-definite everywhere the count is taken — a formula that said two distinct colours differ by zero could not be used to count colours, and the code refuses rather than returning an infinity. And the two metrics must give counts differing by at least a factor of 1.2, which is the structural claim: if they agreed, this essay would have no subject.
The lattice skips . The black point is a single point in the solid, every chromaticity converges on it, and the metric degenerates there. Including it contributes a vanishing volume divided by a vanishing cell size, which is a numerically unstable way of adding approximately nothing.
That redistribution is the reason the count is metric-dependent rather than merely uncertain. If the non-uniformity were noise, better data would shrink the disagreement. It is not noise: it is a systematic mismatch between the shape of human discrimination and the shape of any three-dimensional Euclidean space, and better data makes the mismatch more precisely known rather than smaller.
There is a result lurking behind that which is worth stating even though this site cannot prove it. The set of discrimination thresholds defines a Riemannian metric on colour space, and such a metric can be flattened into a Euclidean one only if its curvature vanishes. The measured colour metric has non-zero curvature. So a perceptually uniform colour space does not exist — not “has not been found”, does not exist — and every uniform space is an approximation whose error is bounded below by that curvature. The counting disagreement is one visible consequence.
The count is a lattice integral, so the check that matters is whether refining the lattice moves the ratio between the two formulae.
Where the model stops
Four places, and they compound.
Perfect packing. Real discrimination cells overlap and leave gaps; the densest sphere packing fills about 74% of space, and JND ellipsoids are not spheres and are not arranged optimally. A realistic count would be smaller by some factor nobody has established.
Pairwise discrimination is not identification. The count asks how many colours are pairwise distinguishable when compared side by side. The number a person can distinguish from memory, or name, or reliably identify in isolation, is smaller by orders of magnitude — a few dozen for naming, a few hundred for reliable identification.
One observer, one condition. The formulae are fitted to averaged data from small panels under controlled viewing. Individual thresholds vary; thresholds change with adaptation, surround and luminance; and none of that is in the arithmetic.
Threshold is not suprathreshold. A threshold is not a unit — the amount of difference at which two samples become distinguishable is not one-tenth of the amount at which they look ten times as different, and difference formulae are fitted to acceptability data rather than to threshold data. Dividing a volume by a “just-noticeable difference” when the formula was fitted to something else is a category slip that every published count makes.
A wider gamut is where the two formulae disagree by more, and the same lattice refinement says by how much.
What the picture cannot show
The hero figure is two bars. It could not have been anything else, and that is a limitation worth naming.
The natural figure for this essay would show the cells — the gamut solid packed with its discrimination ellipsoids, so a reader could see the varying size and count them by eye. That figure cannot be drawn. At ΔE = 1 there are tens of thousands of cells in a solid that has to fit in a few hundred pixels, so any honest rendering is a grey block. Drawing a hundred cells instead and labelling it “not to scale” would be drawing a different claim.
There is a second thing no figure here can show, and it is the subject. Every colour being counted is a colour, and the ones near the boundary of the solid are the ones a reader’s display is least likely to reach, so a figure showing the distinguishable colours would be showing a reader mostly colours their screen is approximating. The display is an unknown, and a count of what can be distinguished is a count made on an apparatus whose own reach is unmeasured.
So the argument is carried by two bars and a ratio, which is less satisfying and is what the evidence supports.
Who found it, and when
MacAdam’s ellipses are 1942. The first attempts to count distinguishable colours from them followed almost immediately — the figure of about ten million circulates from the 1940s onward and is usually attributed to Judd or to Nickerson, with a lineage that is hard to trace because everybody cites everybody.
Judd and Wyszecki’s estimate of around ten million surface colours, and Pointer’s later work on the gamut of real surface colours, are the more careful versions. Pointer’s gamut is much smaller than any display gamut, which is a separate and useful result: most of what a wide-gamut display can show does not occur on any surface, so the extra volume buys reach into a region occupied mainly by light sources and fluorescent materials.
The ΔE2000 formula is Luo, Cui and Rigg, 2001, fitted to a combination of several datasets. Its known defects — including discontinuities in the hue rotation term — were catalogued shortly after and are documented in the CIE’s own technical reports, which is unusually candid for a standard.
Rec. 2020 is the extreme case, and it is the one a bit-depth argument is usually made about.
The bit-depth question, answered properly
The comparison worth making from all this is not between two counts of colours but between a count and a count of codes, because that comparison decides how many bits an encoding needs.
Eight bits per channel gives 16.7 million codes against a distinguishable count of a few hundred thousand under either metric. That is a ratio of fifty to one, and it looks like enormous headroom.
It is not, because the codes are not distributed where the distinguishable colours are. The sRGB transfer function allocates codes roughly evenly in perceptual lightness, which is the right idea, but “roughly” is doing a lot of work: in the deepest shadows the steps between adjacent codes exceed one just-noticeable difference, which is why eight-bit gradients band there and only there. In the highlights the codes are wasted several to a JND.
Ten bits fixes the shadows for standard dynamic range. High dynamic range does not fix so easily, because an absolute luminance encoding has to cover several more decades and the eye’s discrimination does not scale with luminance in a way that any single power law tracks. That is the whole reason PQ exists and is shaped as it is: it is a transfer function derived from a threshold model rather than from a display’s response, which is a different kind of object from the sRGB curve.
Where this goes next
The metric’s shape is the thread. A formula whose JND ellipsoids vary in size by more than an order of magnitude across a solid, and rotate in one region, is a formula worth examining where it is least well behaved — and it turns out not to be smooth everywhere.
Downward, how far apart two colours are is where the formulae themselves are set out, and MacAdam’s measurement is the data all of them are fitted to.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A difference has no place ciede2000 · cielab · δe · just-noticeable difference · threshold · tolerance
- A name is not a threshold ciede2000 · δe · just-noticeable difference · perceptual uniformity · threshold · tolerance
- A colour has a name ciede2000 · cielab · δe · gamut · perceptual uniformity
- A difference is not a distance ciede2000 · cielab · δe · perceptual uniformity · tolerance
- No diagram makes them circles cielab · δe · just-noticeable difference · macadam's ellipses · perceptual uniformity
- Two units with a light level disagree about lightness ciede2000 · δe · just-noticeable difference · perceptual uniformity · tolerance
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
CIEDE2000CIELABΔEGamutJust-noticeable differenceMacAdam's ellipsesPerceptual uniformityQuantisationThresholdTolerance