There is no word for that colour
Assumes A colour has a name and A name is not a threshold.
This phase gave the site a model of colour naming: eleven quoted centroids, a distance, and the colour whose centroid is nearest. It produced several results worth having, and it fails in three specific places that are worth more than the results.
Each failure is a colour the model gives a name no speaker would give it. Together they say something about what a vocabulary is, and it is not eleven points and a metric.
The three failures
A saturated cyan is called green. At L* 60, a* −35, b* −12 — a clear blue-green that most English speakers would call turquoise, teal or “a greenish blue” — the nearest of the eleven basic centroids is green, and it is not close: green is at 12.9 units and the next candidate is grey at 24.3. The model is not confused; it is confidently wrong, because English has no basic term in that region and the model has no way to represent that fact.
Part of a chromatic ring is called grey. At L* 60 and C* 40, with all eleven terms competing, an achromatic term is nearest for four and a half per cent of the hue circle. That is a consequence of the metric rather than of the centroids: CIEDE2000 compresses chroma differences at moderate chroma, so a colour at C* 40 in the cyan region is closer to a grey centroid at C* 0 than to any chromatic centroid twenty units of hue away. Both the previous essays therefore restrict that ring to the eight chromatic terms, which improves the picture and is a patch.
And one focal colour cannot be shown. Blue’s quoted centroid is outside the sRGB gamut. It misses by a small margin and it still misses when the centroid is displaced by the five units it is quoted to, in either direction — so the failure is stable and is not an artefact of the rounding. The site’s rule about unreachable colours applies: the patch is hatched rather than clipped.
What each failure says
They are not three instances of one problem. They fail for three different reasons and each names a different thing a vocabulary has.
The cyan says a vocabulary has holes. A nearest-centroid partition is total by construction: every point in the space gets a name, because some centroid is always nearest. Real vocabularies are not total. There are regions where speakers hesitate, disagree, reach for a modifier or borrow a term from another domain — turquoise, teal, aqua — and the existence of a hesitation region is a property of a language that a partition cannot express.
The grey says a metric decides more than it should. Whether a colour at C* 40 is nearer to grey or to green ought to be a question about vocabulary, and here it is decided by CIEDE2000’s chroma weighting — a function fitted to matching judgements by people who were not asked to name anything. Ask the same question in plain Euclidean distance and 18.8 per cent of the displayable gamut changes name, most often over whether something is pink or purple.
The same number is published twice on this site, at two values
The metric-disagreement figure is the most consequential thing in this essay, and it appears elsewhere on the site at a different value. It is worth setting the two side by side rather than leaving a reader to find them.
Here it is 18.8 per cent, over a 6,574-point lattice. In the essay that puts the metric change against a change of room it is 20.5 per cent, over a lattice built at a six-unit step from about thirty thousand candidate points before the gamut test — a substantially finer sampling of the same gamut. That essay quotes both figures, one as its own measurement and one as a citation, without remarking that they are the same measurement.
So the quantity moves by 1.7 percentage points — nine per cent of itself — between two lattice resolutions used in one collection. Neither is wrong; they are two quadrature rules over one region, which is precisely the situation the census audit spent an essay on, and the lesson there was that refining a rule under a constraint moves the answer without necessarily improving it.
The reason the disagreement is not merely tidy-up is where the number is used. The room essay’s headline comparison is 20.5 against 26.8 — the metric against the surround — and the margin between them is 6.3 percentage points. The lattice disagreement is 27 per cent of that margin. A conclusion resting on a six-point gap, computed from a quantity that moves by 1.7 points depending on how finely the gamut is sampled, is a conclusion with a quarter of its evidence coming from a choice nobody stated.
It survives, on the ordering: 18.8 and 20.5 are both well below 26.8, so the room beats the metric whichever lattice is used. What does not survive is the precision. A reader carrying away the room renames 6 points more of the gamut than the metric does is carrying a difference known to about 1.7 points, and the honest form is the room renames noticeably more, by something between four and eight points depending on how the lattice is drawn.
And the repair is the collection’s own rule, applied to itself. Every residual in the adaptation family now names the set it was computed over, because a mean is a statement about a set. A renaming percentage is a mean over a lattice in exactly the same sense — the fraction of a sampled region satisfying a predicate — and it carries the same obligation. Two essays quoting one quantity at two values, neither naming its lattice in the sentence where the number appears, is the failure that rule was written to prevent, arriving in a family the rule had not been applied to.
The general form is worth keeping because it is not about naming. A percentage of a region is a quadrature result wearing a proportion’s clothing, and a proportion looks unitless and self-explanatory in a way a residual does not — which is exactly why it escapes the discipline a residual now gets. Any figure of the form x per cent of the gamut on this site inherits a lattice, and none of them says so.
That number is the sharpest statement of the problem in this essay. A model in which nearly a fifth of the answers depend on which distance function is used is a model in which the distance function is doing a fifth of the work, and no theory of colour naming assigns it any.
And the undisplayable blue says the focal colours are further out than the technology. That is a fact about sRGB rather than about naming — and it is the one failure that would be repaired by better hardware rather than by a better model, which is worth separating out.
Why the failures were worth building the model to find
There is a reasonable objection to this whole exercise: if the model is known in advance not to be a theory of naming, what was gained by building it and then cataloguing its errors?
Three things, and the third is the one that justifies the phase.
The results that do not depend on the failures. The territories are unequal by a factor of 4.6, the arcs by 4.5 in the metric’s own units, the radii by 2.8, and the naming resolution is about ten thresholds. None of those numbers needs the model to be a theory; they need it only to be an interpolation between quoted focal colours, which is what it honestly is. A result that follows from the quoted data plus arithmetic survives the model being wrong about the parts the quoted data does not cover.
A measurement of how much the metric is deciding. Nobody would have guessed 18.8 per cent. The number is only available because two partitions were computed and compared, and it is the strongest available argument that a colour difference formula is doing work in this area that nobody has authorised it to do.
And a set of failures precise enough to be useful. “A nearest-centroid model is not a theory of naming” is a sentence anybody could have written before the model existed. “It calls a saturated cyan green at 12.9 units against grey at 24.3, gives an achromatic name to four and a half per cent of a chromatic ring, and changes a fifth of its answers under a different distance” is a list of specific things a better model would have to do differently — and the second of the three is a defect nobody would have predicted, since it comes from the metric rather than from the vocabulary.
The fourth absence
This site asserts absences rather than leaving them implicit, and there are now four: binocular mixing, why a stabilised image fades, why dither works, and this one.
Naming is not explained here. A vocabulary is a partition that people agree on: learned, shared across languages that had no contact, and stable against the enormous variation in the eyes doing the learning. None of those properties is in this model. What is here is an interpolation between eleven quoted numbers, which reproduces the arithmetic of a partition and predicts nothing at all about why that partition rather than another.
assertNamingIsNotExplainedHere is written to fail if the model ever stops failing — if a future revision names every colour on the ring the way a speaker would, the site stops building and somebody has to decide whether the model has become a theory or has merely been fitted to the cases it was being embarrassed by. That distinction is the whole reason for asserting an absence rather than writing it in a comment.
What was computed, and how
The cyan is one call: a colour is chosen in the region English has no basic term for, and the model is asked what it is. The answer and the two runners-up are reported, which is the informative form — a model that was merely uncertain would show a near-tie, and this one does not.
The ring share is a scan. Every two degrees of hue at L* 60 and C* 40, inside the gamut, is named with all eleven terms competing, and the fraction receiving an achromatic name is counted. Four and a half per cent.
The undisplayable term is a gamut test on each centroid and on the same centroids displaced by the quoted rounding in both directions. Blue fails all three; two other terms fail one displacement each, which is why the assertion asks for the stable set rather than the set at the quoted values.
And the metric disagreement is the partition computed twice over the same 6,574-point lattice, once under each distance, with the differing points counted and the most common disagreeing pair reported.
None of it is difficult. The reason it is here rather than in a comment is that a list of a model’s failures, computed, is a different object from a list of a model’s failures, remembered: this one is re-measured whenever the site is rebuilt and changes when the model does.
Where a real model would differ
Worth sketching, because “this is not a theory” is more useful with an indication of what one would contain.
It would have a category that is not a point. The strongest empirical result in colour naming is that speakers agree far more about the best example of a term than about its boundaries. A model built from centroids reproduces that trivially and for the wrong reason — its boundaries are wherever two quoted points happen to be equidistant, which is a fact about the quotation.
It would have hesitation. A region where no term applies well, where speakers use modifiers, and where the naming probability is low for everything. The softmax here can produce a flat distribution but cannot produce a low one, because it is normalised: every colour gets a full unit of probability distributed among eleven terms, however badly they all fit.
It would have more than eleven terms available, with the basic ones distinguished by frequency and consensus rather than by being the only options.
And it would be fitted to naming data. All of the above is available in the World Color Survey and in the naming studies that followed it, and none of it is available to this site, which has eleven centroids quoted from a summary.
The generalisation
The sentence worth carrying is: a total function is the wrong shape for a partial phenomenon.
Nearest-centroid naming is total: every colour gets a name. Language is partial: some colours get a name easily, some get one with a modifier, and some get an argument. Building the first as a model of the second guarantees the failures in this essay, and no amount of moving the centroids removes them.
This site has met the same shape before. A colour vision deficiency simulation is a total function producing an image for every input, standing in for something — what another person sees — that is not an image at all; the repair there was to narrow the claim to which discriminations survive, which is total and is checkable. A dominant wavelength is a total construction over a region where a third of the answers do not exist.
The pattern in each case: a total function is easy to compute, a partial phenomenon is what is being modelled, and the difference shows up as confident wrong answers in the region where the phenomenon has nothing to say.
Who found it, and when
Berlin and Kay’s 1969 monograph is where the eleven come from, and the hesitation regions were visible in their data from the start: the boundaries between terms were the part that varied between speakers and between languages, and it was the focal colours that agreed.
The critiques through the 1970s and 1980s are largely about exactly this — that the strong reading of the result treats a set of centroids as though it were a partition, and that the partition is the part the data does not support.
Kay and Regier’s later statistical work on the World Color Survey put numbers on both halves: substantial cross-language agreement about the best examples, much weaker agreement about where the regions end. Any model that reproduces the first and fabricates the second will look right on the published summaries and wrong on the raw data.
And turquoise is the standing counterexample. Whether the blue–green boundary region is one category, two, or a gap between two is language-dependent, has been the subject of some of the field’s most-cited experiments, and is precisely the region where this model says “green” with no hesitation at all.
What the pictures cannot show
They cannot show hesitation. Every figure in this family assigns a name to every colour, because the model does. A picture of the model’s actual uncertainty would need an output the model does not have.
And they cannot ask the reader. The one experiment that would settle any of this is to show a reader the cyan patch and ask what it is, and a figure cannot do that — nor could it record the answer, nor compare it with the reader’s neighbour’s, which is where the whole phenomenon lives.
Where the ladder goes next
The nearest unfinished piece is a hesitation output. A model that reported how well its best term fits, rather than only which term is best, would mark the cyan region without needing more data — the distance to the nearest centroid is already computed and is simply thrown away.
The second is the metric question, which is answerable and uncomfortable. Nearly a fifth of the partition depends on which distance is used, so somebody could ask which distance best reproduces published naming boundaries — and the answer would be a statement about a colour difference formula derived from linguistic data, which is not how any of them were built.
And the third is the one this essay keeps circling: the raw naming data exists, has been public for decades, and would replace every quoted number in this phase’s naming work with a measurement. What this site has instead is eleven points, and three failures that are the shape of the gap between them.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- A catalogue is not a vocabulary basic colour terms · chroma · cielab · gamut · naming · perceptual uniformity · specification · standard observer
- A departure is straight in the excitations assertion · chroma · ciede2000 · individual variation · perceptual uniformity · standard observer
- A name in the model's own words assertion · basic colour terms · categorical perception · cielab · naming · perceptual uniformity
- A difference is not a distance assertion · ciede2000 · cielab · perceptual uniformity · specification
- A gradient is a path chroma · cielab · gamut · perceptual uniformity · specification
- Three constants nobody quotes chroma · ciede2000 · perceptual uniformity · specification · standard observer
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AssertionBasic colour termsCategorical perceptionChromaCIEDE2000CIELABGamutIndividual variationNamingPerceptual uniformitySpecificationStandard observer