What the brain does

Brightness is not luminance

Photometry is additive because the CIE defined it that way. Brightness is not, and a saturated colour looks as bright as a neutral of one and a half times its luminance — an effect this site's appearance model moves by two per cent, in a direction that depends on the hue.

Assumes Brighter looks more colourful and A viewing condition is an argument.

Luminance is additive. Two lights mixed have the sum of their luminances, exactly, for every pair of lights — which is what makes colour add up — and that is not a measurement. It is a definition: luminance is a spectrum integrated against one agreed curve, and an integral of a sum is the sum of the integrals.

Brightness, which is how bright something looks, is not additive. The two words are used interchangeably in almost every technical document that is not about colour appearance, and this essay is about the size of the gap.

Brightness against chroma, at exactly constant luminance. Seven stimuli of identical luminance and rising chroma at hue 25. The model's brightness moves by 2.2 per cent across the whole sweep, and at some hues it moves the other way. The Helmholtz–Kohlrausch effect — measured repeatedly, by several methods — is that a saturated colour looks as bright as a neutral of 1.3 to 2 times its luminance, shown as the band. The gap is the model's, and nothing here closes it: the term that would is not in CIECAM16 and is not invented for the occasion.
Fig. 1 Seven stimuli of identical luminance, rising in chroma at one hue, with the model’s brightness plotted against the measured Helmholtz–Kohlrausch range. The model moves by two per cent across the whole sweep. Observers report that a saturated colour looks as bright as a neutral of one and a half to twice its luminance, which is the band. The gap belongs to the model, and this site does not close it.

The claim

A saturated colour looks brighter than a neutral of the same luminance, and this site’s appearance model does not predict it.

Two separate statements, and the second is the one worth recording carefully:

  • The Helmholtz–Kohlrausch effect is measured, repeatedly, by several methods, at a size usually quoted as an equivalent luminance ratio between 1.3 and 2.
  • CIECAM16’s brightness correlate QQ, computed here at rigorously constant luminance across a chroma sweep, moves by at most 2.2 per cent — and the direction depends on the hue. At hue 25 it rises by 2.2 per cent; at hues 120 and 240 it falls by 1.1.

A model that moved the right way and too little would be a model with a weak version of the effect. A model that moves by two per cent in whichever direction the hue happens to send it has no version of the effect at all, and the two per cent is the arithmetic of its nonlinearity rather than anything about appearance.

Why the model has none

CIECAM16’s brightness is built from the achromatic response:

A=(2Ra+Ga+Ba/200.305)Nbb,J=100(A/Aw)cz,QJA = (2R_a + G_a + B_a/20 - 0.305)\,N_{bb}, \qquad J = 100\,(A/A_w)^{cz}, \qquad Q \propto \sqrt{J}

Chroma enters that chain only through the compressive nonlinearity applied to each adapted cone signal. A saturated stimulus has a different distribution of cone responses from a neutral of the same luminance, so the compressed sum comes out slightly different — and that residue is what the two per cent is.

Nothing in the model adds colourfulness into brightness, and the effect being described is exactly that: the chromatic channels contributing to apparent brightness — a second thing an appearance model has no slot for. The published extensions that do model it — a Helmholtz–Kohlrausch term added to CIECAM-class models — are additions rather than consequences, and no such term is in this site’s implementation.

So the absence is asserted rather than described. assertTheModelUnderpredictsHelmholtzKohlrausch sweeps four hues, requires the luminance to be held to a millionth, and requires the brightness movement to stay under five per cent. If a future revision of the appearance model starts predicting the full effect, the build stops and somebody has to say where the term came from. That is the same shape as the Stevens effect’s absence, recorded during the expansion phase for the same reason: an unexplained improvement is not good news.

Brightness against chroma, at exactly constant luminance. Seven stimuli of identical luminance and rising chroma at hue 240. The model's brightness moves by -1.1 per cent across the whole sweep, and at some hues it moves the other way. The Helmholtz–Kohlrausch effect — measured repeatedly, by several methods — is that a saturated colour looks as bright as a neutral of 1.3 to 2 times its luminance, shown as the band. The gap is the model's, and nothing here closes it: the term that would is not in CIECAM16 and is not invented for the occasion.
Fig. 2 The same sweep at a blue hue, where the model’s brightness falls by about one per cent as chroma rises. Saturated blues are the classic demonstration of the Helmholtz–Kohlrausch effect in the other direction — they look far brighter than their luminance — so this is the model at its furthest from the measurement.
Brightness against chroma, at exactly constant luminance. Seven stimuli of identical luminance and rising chroma at hue 120. The model's brightness moves by -1.1 per cent across the whole sweep, and at some hues it moves the other way. The Helmholtz–Kohlrausch effect — measured repeatedly, by several methods — is that a saturated colour looks as bright as a neutral of 1.3 to 2 times its luminance, shown as the band. The gap is the model's, and nothing here closes it: the term that would is not in CIECAM16 and is not invented for the occasion.
Fig. 3 And the same sweep at a green hue. The three hues together are the finding: at fixed luminance the model’s brightness moves with chroma, and which way it moves is a fact about the hue rather than about the light.

What the effect actually is

The clean statement is about equivalent luminance. Set a saturated patch beside a neutral one and ask an observer to adjust the neutral until the two look equally bright. The neutral ends up at a substantially higher luminance than the coloured patch — up to about twice, depending on hue, purity and the method used.

Three things follow, and each contradicts a piece of everyday practice.

A photometer does not measure brightness. It measures luminance, which is a weighted integral with V(λ) as the weight, and V(λ) was itself established by flicker photometry — a method chosen precisely because it makes the answer additive. The 683 lm/W ceiling and every efficacy figure on this site are luminance quantities, and a lamp that scores well on them can look duller than one that does not.

Equal luminance is not equal appearance, which makes “isoluminant” a colorimetric term rather than a perceptual one. An isoluminant grating is one the luminance channel cannot see; it is not one that looks flat.

And a contrast ratio is a luminance ratio. That last one has consequences well outside this subject.

What a contrast ratio cannot see

The accessibility rule that governs text on screens is

L1+0.05L2+0.054.5\frac{L_1 + 0.05}{L_2 + 0.05} \ge 4.5

with LL the relative luminance. It contains no observer state, no adapting level, no surround and no chroma, so every pair of colours with the same ratio is the same pair as far as it is concerned.

Sampling pairs at exactly that ratio and asking the appearance model what their lightness difference is:

ratio pairs found smallest ΔJ largest ΔJ spread
3.0 274 22.2 47.2 2.13×
4.5 115 34.0 56.5 1.66×
7.0 44 47.4 67.3 1.42×

At the pass mark, pairs the rule cannot distinguish differ by a factor of 1.66 in what the model says about them.

And the ordering inverts. Among pairs near the mark, the least-different passing pair is 34.2 lightness units apart at a ratio of 4.54; the most-different failing pair is 58.3 units apart at 4.41. The rule accepts the first and rejects the second, and the model says the second is nearly twice as distinguishable.

Four pairs at contrast ratio 4.5, and what the model says about them. Every pair here has the same luminance contrast ratio to within a twentieth, so the accessibility rule cannot tell them apart. The appearance model puts their lightness differences between 34 and 56 units — a spread of 1.66×. Worse, the ordering can invert: the passing pair below is 34 lightness units apart at a ratio of 4.54, and the failing pair is 58 apart at 4.41.
Fig. 4 Four pairs, all at the same luminance contrast ratio to within a twentieth. The top row is the widest and narrowest lightness differences the model finds at that ratio; the bottom row is the inversion — a passing pair the model puts closer together than a failing one.

The same model says a second thing about brightness that luminance has no way to say, and it says it as a colour rather than as a curve.

One light, seven rooms. The same stimulus — fixed in XYZ, unchanged throughout — shown against whites from 12 to 800 candelas per square metre. Its lightness falls from 152 to 16 and its brightness rises, because one of those is a ratio to the white and the other is not. Brown is the low-lightness end: a colour that exists only when something brighter is present, which is why no lamp is brown and no star is.
Fig. 5 One stimulus, fixed in every colorimetric coordinate, shown against surrounds from twelve rooms. Nothing about the light reaching the eye has changed between the patches, and their brightness has.
Lightness and brightness of one light, against the white it is judged with. The stimulus never changes. As the white in the field rises from 12 to 800 cd/m², the model's lightness falls from 152 to 16 — a factor of 9.4 — while its brightness rises. Lightness is a ratio to the white and brightness is not, and everything that can only be dark lives in the first family.
Fig. 6 And the same sweep as two curves. Lightness falls and brightness rises across it, which is a pair of statements no single luminance can carry.

The rule’s resolving power is a factor of 2.3

The three rows of that table are three points on two straight lines, and drawing them turns “a factor of about 1.7 at the mark” into something a designer could act on.

Both bounds are linear in the logarithm of the ratio, to within a third of a lightness unit at every point:

ΔJmin=10.6+68.5log10r,ΔJmax=21.0+54.6log10r\Delta J_{\min} = -10.6 + 68.5\log_{10} r, \qquad \Delta J_{\max} = 21.0 + 54.6\log_{10} r

The lower bound climbs faster than the upper one, so the spread narrows as the ratio rises — 2.13 at a ratio of 3, 1.66 at the pass mark, 1.42 at 7 — and the two lines meet at a contrast ratio of 191. Since sRGB’s own maximum is 21:1, black on white, the rule never becomes exact on any display: even at that extreme it still carries a spread of 1.17.

And the rule is worst where it is used most. The spread grows as the ratio falls, so the proxy is at its least reliable in the range between 3 and 4.5 — which is exactly the range a designer argues about, because it is the range where a choice passes or fails.

The sharper form is the resolving power. Two pairs are certainly ordered by the rule only when the weaker pair’s largest lightness difference is below the stronger pair’s smallest — and setting ΔJmax(r1)=ΔJmin(r2)\Delta J_{\max}(r_1) = \Delta J_{\min}(r_2) gives

r2r1=1031.6/68.52.3\frac{r_2}{r_1} = 10^{31.6/68.5} \approx 2.3

A luminance contrast ratio cannot order two colour pairs whose ratios differ by less than a factor of 2.3. A pair at 3:1 may be more distinguishable than a pair at 6:1; only past about 7:1 is it certainly beaten.

That is a stronger statement than the inversion example the essay already gives, and it is the same fact generalised. The inversion — a passing pair at 4.54 sitting closer together than a failing one at 4.41 — looks like a boundary curiosity, something that happens only at the threshold where two verdicts meet. It is not. The bands overlap over a range of 2.3 in ratio at every point on the scale, so the ordering fails everywhere and the threshold is merely where anybody notices.

Which is a precise account of what the rule is for. It is not a ranking and it should not be read as one: two designs at 4.6 and 5.5 have not been ordered by it, whatever their numbers say. What it is is a floor, and a floor is a much weaker instrument than a scale — it says a pair below the mark is probably bad without saying that a pair above it is better than another pair above it. Every use of the rule that compares two passing designs is using it a factor of 2.3 past its resolution, and the factor is now measurable rather than a matter of opinion.

None of that makes the rule useless. It is cheap, it is computable from a hex code without any viewing information, and it correlates with legibility well enough to have improved a great deal of typography. It is a proxy, its residual error is a factor of about 1.7 at the mark and 2.3 as a resolving power, and knowing the size of the residual is what separates using a proxy from believing it.

Where the gap shows up

Signage and safety colours. A saturated red sign and a grey one of equal luminance are not equally conspicuous, and the standards that govern warning colours specify chromaticity as well as luminance for exactly that reason. The equivalent-luminance correction is one of the few places the effect is written into practice.

Displays sold on peak brightness. A panel’s headline figure is a luminance, measured on white. What a viewer notices about a highly saturated highlight is not a luminance, and a display with wider primaries can look brighter at the same measured peak — the gamut and the brightness are not independent as far as appearance is concerned, although they are entirely independent as far as photometry is concerned.

Lighting design. A lamp’s efficacy is a luminance per watt, so a source optimised for it is optimised for the additive quantity. There is a known and awkward consequence: light sources with more chromatic content can be judged brighter at equal measured output, and a design rule written in lumens cannot express that.

And the blue channel of everything. Saturated blues have very low luminance — the blue primary of a display carries about seven per cent of white’s — while looking far from dark. That mismatch is where the effect is largest, and it is why blue text on a dark ground routinely passes a luminance-based contrast check while being difficult to read, and why a blue that looks bright measures as though it were nearly black.

What the pictures cannot show

The effect itself. A figure demonstrating the Helmholtz–Kohlrausch effect would need two patches that a reader judges equally bright, and this page cannot control the reader’s adaptation, surround or display. What is drawn instead is the model’s prediction and the measured band around it, which is a picture of a gap rather than of an appearance.

And the model’s two per cent is not visible either. A change of two per cent in brightness is below any threshold a reader could apply to two patches on a page; the number is meaningful only against the thirty to hundred per cent that was measured, and the whole content of the comparison is the ratio between them.

What was computed, and how

The sweep. Seven stimuli at one hue, with chroma from 0 to 60, each solved so that its luminance is exactly 30 — by iterating CIECAM16’s inverse on lightness until YY lands, which is checked to a millionth on every row before anything is reported. Holding YY rather than JJ is the whole experiment: a sweep at constant lightness would be answering a different question.

The brightness is CIECAM16’s QQ under an average surround at 100 cd/m², reported as a ratio to the neutral’s QQ at the same luminance.

The contrast pairs are random sRGB triples, filtered to those whose luminance contrast ratio lands within 0.05 of the target, with the lightness difference computed as J1J2|J_1 - J_2| under the same viewing conditions. Forty thousand candidates give a few hundred at any one ratio, which is why the tables report how many were found.

And the measured range is quoted, not computed. 1.3 to 2.0 is a band drawn around what the literature reports for the equivalent-luminance ratio of a saturated colour; no experiment is done here and none is claimed.

A magenta is the hue where the model’s brightness moves least, and it is the strongest form of the complaint.

Brightness against chroma, at exactly constant luminance. Seven stimuli of identical luminance and rising chroma at hue 300. The model's brightness moves by 1.4 per cent across the whole sweep, and at some hues it moves the other way. The Helmholtz–Kohlrausch effect — measured repeatedly, by several methods — is that a saturated colour looks as bright as a neutral of 1.3 to 2 times its luminance, shown as the band. The gap is the model's, and nothing here closes it: the term that would is not in CIECAM16 and is not invented for the occasion.
Fig. 7 Seven stimuli of identical luminance and rising chroma at hue 300. The model’s brightness moves by 1.4 per cent across the whole sweep and at some hues it moves the other way, while the effect being described is repeatedly measured.

Where the model stops

The model has no Helmholtz–Kohlrausch term and this essay does not add one. Fitting one would mean inventing a constant, and the site’s whole position is that a computed number should be downstream of something stated. What is offered instead is the measured absence and a pointer to the published extensions.

The two bounds are three points each. The straight lines fitted through them reproduce all six values to within a third of a lightness unit, which is a good fit and is a fit to three points — so the meeting-point at a ratio of 191 is an extrapolation nine times past the last measurement and should be read as far outside anything a display reaches rather than as a number. The resolving power of 2.3 is an interpolation between the measured rows and is on firmer ground.

The contrast-ratio comparison uses lightness, not brightness. JJ is the relative attribute and is the right one for text on a page, since a page has a white. Using QQ would change the numbers and not the conclusion.

One viewing condition throughout. 100 cd/m², average surround, D65 white, twenty per cent background. Text on a phone at night is none of those, and every number would move.

And no legibility is measured. Lightness difference is not reading speed; the claim here is that the rule and the model disagree about ordering, not that the model predicts who can read what.

A seven-to-one contrast ratio is what an accessibility standard asks for at its stricter level, and the pairs that satisfy it are not equally legible.

Four pairs at contrast ratio 7, and what the model says about them. Every pair here has the same luminance contrast ratio to within a twentieth, so the accessibility rule cannot tell them apart. The appearance model puts their lightness differences between 47 and 67 units — a spread of 1.42×. Worse, the ordering can invert: the passing pair below is 48 lightness units apart at a ratio of 7.01, and the failing pair is 67 apart at 6.50.
Fig. 8 Every pair here has the same luminance contrast ratio to within a twentieth, so the rule cannot tell them apart. The appearance model puts their lightness differences between 47 and 67 units, a spread of 1.42 times.

The generalisation

There is a family of quantities in this subject that are defined to be linear, and a matching family that are measured and are not, and almost every confusion in applied colour is a member of the first standing in for a member of the second.

defined, linear measured, not linear
luminance brightness
tristimulus values appearance
dominant wavelength hue
purity saturation
a spectral integral a judgement

The left column exists because somebody needed arithmetic that closes: photometry has to add, or a lighting calculation cannot be done at all. The right column is what people mean when they use the words. The two agree well enough for a great deal of engineering, and the places they part company are the places this site keeps finding its essays.

The practical rule is a question rather than a formula: is this quantity additive because it was measured to be, or because it was defined to be? For luminance the answer is the second, V(λ) was established by a method chosen to make it so, and every downstream use inherits the choice.

Who found it, and when

Helmholtz noticed it and Kohlrausch measured it, in the 1920s: a coloured light matched in luminance to a white one does not look equally bright, and the discrepancy grows with purity.

The reason it survives as a named effect rather than being absorbed into photometry is the reason above. Additivity was chosen. Flicker photometry, and heterochromatic brightness matching by other methods, give different luminous efficiency functions — and the CIE standardised the additive one, knowing it was not the brightest-looking one, because a photometry that does not add is not usable for calculation.

The modern form of the argument is in display engineering, where “equivalent luminance” corrections are applied to saturated colours in some signage and cockpit standards and to nothing else, and in accessibility, where the luminance-only contrast rule is periodically re-litigated on exactly the grounds measured above.

Colourfulness rises with light level; apparent contrast does not. Predicted colourfulness M for one stimulus across four decades of adapting luminance, and the exponent of the lightness curve over the same range. M rises by a factor of 2.24 — the Hunt effect, which the model does predict. The lightness exponent changes by -2.2%, and in the wrong direction — the Stevens effect, which it does not.
Fig. 9 And the effects the model does predict, for contrast: colourfulness rises with light level and the model says so. The Helmholtz–Kohlrausch effect is not in that list, which is why an essay is needed to say so rather than a figure.

What the pictures cannot show

The effect, again. Every figure here draws the model or the pairs; the phenomenon itself needs a matching experiment the page cannot run. What a reader can do is the informal version: find a saturated blue and a grey that a photometer calls equal, and notice that they are not.

And the contrast pairs are drawn as adjacent patches, which is the arrangement the accessibility rule is about — text on a ground — but at a size and a spacing the rule does not specify either. A difference has no size applies to this rule as much as to any tolerance: the same pair at nine points and at ninety is two different judgements, and the ratio is the same number for both.

Where the ladder goes next

The obvious next rung is the one the model’s structure points at: the difference between the relative attributes and the absolute ones. Lightness and chroma are ratios to a white; brightness and colourfulness are not, and a stimulus with no white beside it has only the second kind. That turns out to decide which colours can exist at all — there is no brown light — and it is the same distinction as this essay’s, applied to naming rather than to matching.

The other direction is the honest repair. Adding a published Helmholtz–Kohlrausch term to the appearance model would be a real improvement and a real risk: it would have to be checked against the same published test vector the rest of the model is checked against, and against data this site does not hold. Until then the absence is asserted, which is worth more than a term nobody can test.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 10 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AccessibilityBrightnessCIECAM16Colour appearanceColourfulnessContrast ratioLightnessLuminanceLuminous efficiencyPhotometry