A lamp is not a blackbody
Assumes The illuminant is half the answer and Blackbody and the colour of temperature.
Every idealised light source on this site is smooth. Planck’s law gives a continuous curve, the CIE’s daylight reconstruction gives another, and both vary gently with wavelength because the physics producing them does.
Almost nothing anybody actually reads under is like that.
Three phosphors is one construction among several, and the others fail the same assumption in different shapes.
A white light-emitting diode fails it in a third way, with no thermal radiator anywhere in the device at all.
And the furthest thing from a blackbody that still produces a white anybody would accept is three narrow emitters with nothing between them.
Where the structure comes from
A fluorescent tube is a mercury discharge inside a phosphor-coated envelope. The discharge emits mostly in the ultraviolet, at 254 nm, and the phosphor converts that into visible light. What the phosphor does not convert is the mercury’s visible emission, which passes straight through unchanged.
That is why the four lines are universal. They are not a design choice and not a defect; they are the lamp’s mechanism showing through. A warm-white tube and a daylight tube differ in their phosphor and are identical in their lines, which is why fluorescent spectra are instantly recognisable and why any two of them have more in common with each other than either has with daylight.
A white LED is a different mechanism with the same consequence. A blue indium-gallium-nitride die emits a narrow band near 450 nm; a cerium-doped garnet phosphor in front of it converts part of that blue into a broad yellow. Blue plus yellow reads as white. What it is not is broad: there is a pronounced trough around 480 nm between the spike and the hump, and a steep fall past 650 nm.
Constructed, not tabulated
The spectra here are built from a stated emission model rather than read from a table, and the choice costs something worth being explicit about.
These are not the CIE’s F-series illuminants and must not be called them. What they are is a phosphor continuum of stated band positions and widths, plus mercury lines at the wavelengths mercury emits at, plus a stated line strength. Nothing was fitted to a measured lamp.
What that buys is that the structure is visible as structure. A mercury line is in the spectrum because a line was added at 546 nm with a width, and the width can be changed to see what happens. A tabulated lamp is eighty-one numbers and explains nothing about why it looks the way it does.
The instrumental width is worth a note. Real mercury lines are of the order of a picometre wide. The spectra here are sampled every 5 nm, so a line narrower than the sampling interval is not representable at all — and any published fluorescent spectrum is already the lamp convolved with a spectroradiometer’s bandpass. The width used here is that instrumental width, which is the honest one: it is what a measurement of the lamp actually reports, and what an instrument reports is a subject in itself.
What a colour temperature leaves out
Every one of these lamps is sold with a number on the box, and the number is a correlated colour temperature: the temperature of the Planckian radiator the source most resembles.
“Most resembles” is doing a great deal of work. It means nearest point on the Planckian locus, nearest means perpendicular, and perpendicular is only well defined on one particular diagram — the 1960 UCS, superseded in 1976 and kept alive for this single purpose.
The length of that perpendicular is Duv, and it is the coordinate the number on the box throws away. Two lamps at the same colour temperature, one above the locus and one below, are visibly green and visibly pink respectively. Both are correctly labelled.
Duv is quoted in professional lighting specifications and essentially never on consumer packaging. A green cast in a room is one of the most commonly noticed and least commonly diagnosed lighting complaints, and the number that would have predicted it was measured and discarded.
Chromaticity does not determine rendering
Here is the claim that matters most, and it can be stated as a computation.
Take a halophosphate fluorescent tube. Take a white LED, and tune its phosphor conversion until its white lands on the tube’s — the two agree to within ΔE′ 0.31, comfortably inside a just-noticeable difference, so a person looking at the two lamps could not tell them apart.
Now put a surface under each. The worst disagreement across a set of saturated test reflectances is ΔE′ 5.6, which is not subtle — it is a different colour.
The mechanism is the collapse this site is built around, applied twice. The lamp’s spectrum is projected onto three numbers, which is where the holes stop being visible. Then the reflectance multiplies the spectrum before the projection happens again — and a surface reflecting where the lamp has no output has nothing to reflect. The first projection discarded exactly the information the second one needed.
This is metamerism with the arrow reversed. The usual case is two stimuli that match for one observer and not another. Here two illuminants match for every observer, and everything they fall on comes apart.
What a rendering index measures, and what decides it
Every colour rendering index works the same way: render a set of test reflectances under the source and under a reference of the same colour temperature, adapt both to a common white, and measure how far apart they land.
The index computed here is that method in CAM16-UCS, over reflectances constructed for the purpose. It is stated not to be CIE Ra: the CIE’s index uses eight tabulated samples in a 1964 space with a fixed scale factor, and this site does not have those samples. What it shares is the method.
Running it produces one result the method makes unavoidable, and it is not about lamps.
How saturated the test samples are decides the verdict at least as much as the lamp does. A source with holes in its spectrum is punished heavily by narrow reflectances and lightly by broad ones; a smooth source is barely affected either way. So an index computed on desaturated samples grades every lamp on the case it cannot fail.
This is not a novel observation — it is exactly why the CIE’s own index has a saturated red sample, R9, that is excluded from the headline average, and why every modern replacement reports a distribution rather than a scalar. A tube can score 85 on the eight samples that count and fail badly on the one that does not, and the number on the box is the 85.
Two lamps in the same room
The most common practical failure has nothing to do with indices, and it is the one everybody has seen.
Put two lamps of different spectra in one room — a fluorescent tube in the ceiling, an LED lamp beside the desk, both nominally the same colour temperature — and surfaces change colour as they move between them. Not because either lamp is wrong, but because a surface’s colour is a property of the surface and the light, and there are now two lights.
The visual system makes this worse by handling it well. Constancy discounts the illuminant and does so by estimating a single adapting white for the scene, so in a room with two different lights it produces one estimate that suits neither. The estimate is a compromise, and everything is slightly wrong everywhere rather than clearly wrong in one place.
The lighting design answer is to specify not just colour temperature but Duv and a rendering measure, and to use one lamp type throughout a space. The reason that advice is so often ignored is that the number on the box is a colour temperature, two lamps with the same number look identical when compared against each other directly, and the failure only appears once there is something in the room to look at.
What was not established
A phase that constructs lamp spectra and computes their rendering is under obvious pressure to reproduce the received ranking: triphosphor good, halophosphate bad. That claim was written into the gate first and it failed, repeatedly, under every retuning of the two spectra that kept them faithful to their known structure.
It was deleted rather than tuned until it passed. These are constructed spectra, not measured lamps, and a ranking between two constructions is a fact about the constructions. Asserting it would have been the kind of unearned claim this site exists to avoid.
What survives is what the machinery genuinely establishes: a blackbody scores 100 against its own reference, exactly; every source with spectral structure scores materially below it; two sources of identical chromaticity render differently by a wide margin; and the sample set decides the verdict. Those are asserted, and each is asserted in the direction that can fail.
The gate also checks that the index is capable of reporting a bad lamp as bad, by handing it a comb spectrum — a source designed to be terrible — and requiring it to score twenty points below a real LED. An index that gave everything a high number would be the commonest failure of this kind of measure and the hardest to notice, because a high number is what everybody wants.
What was computed here
The four mercury lines must survive the sampling grid as local maxima, or a figure claiming to show a discharge lamp would be showing a continuum. That is checked, at all four wavelengths.
Correlated colour temperature must recover a Planckian radiator’s own temperature exactly, since the definition is circular in the useful way: the CCT of a blackbody at T is T. Measured across the range the scale is used over, the worst relative error is 2 × 10⁻¹⁵ and the Duv of every blackbody is under 10⁻⁴ — which it must be, since a blackbody is on its own locus.
The rendering comparison uses the model from the appearance essay, and it has to: comparing two lamps requires adapting each to its own white first, which is an appearance calculation whether or not anybody calls it one. Doing it in CIELAB would confound the rendering difference with the white-point difference.
What the number on the box could say instead
Given all of this, a fair question is what a lamp’s packaging would have to carry to be useful.
Correlated colour temperature and Duv, together. The first alone is half a chromaticity and the missing half is the one that reads as a green or pink cast across a whole room.
A fidelity measure and a gamut measure, separately. A source can render every sample faithfully and be unpleasant, and a source that slightly increases chroma is generally preferred to one that is exactly faithful. Reporting one number conflates two properties that people want in opposite directions.
The worst hue rather than only the average. A lamp with one catastrophic region and a lamp with mild error everywhere score the same on an average and are not the same lamp.
Modern lighting specifications carry all of this, and consumer packaging carries a colour temperature and sometimes a colour rendering index whose sample set is not stated. The gap is not a failure of colour science, which settled these questions decades ago; it is that the box has room for one number and the number chosen was the one that existed first.
Eighteen to one
The essay’s central computation has two numbers in it and their ratio is the claim.
The tuned LED and the halophosphate tube agree on white to ΔE′ 0.31 and disagree about a surface by ΔE′ 5.6. That is a factor of 18: matching two lamps’ whites to a third of a just-noticeable difference leaves the surfaces under them eighteen times further apart than the whites are.
Which is the number to carry rather than either half of it. A reader told that two lamps at the same chromaticity render differently might picture a small residual; the residual is not small relative to the agreement, it is an order of magnitude larger. And the direction is fixed: making the whites agree better does nothing at all to the surfaces, because the two quantities are computed from different parts of the same spectrum.
It also says what a white-point match is worth as a specification. Specifying a lamp to a tenth of a ΔE on its white constrains its rendering by nothing, since the constraint acts on three numbers and the rendering depends on the eighty-one the three were made from.
A four-phosphor tube sold on its colour rendering is the case where the spectrum is deliberately made less spiky, and its white is worth putting beside the others.
The sample set nearly doubles the gap between lamps
How saturated the test samples are decides the verdict at least as much as the lamp does is stated as a comparison and the two sensitivities give it a number.
A smooth source loses 1.22 times as much on the saturated set as on the moderate one; a narrowband source loses 2.36 times as much. The ratio between the two sensitivities is 1.93.
So changing the sample set nearly doubles the gap between a good lamp and a bad one. A pair losing 5 and 20 points on moderate samples — a gap of four times — loses 6.1 and 47.2 on saturated ones, a gap of 7.7. A pair at three and fifteen goes from five times to nearly ten.
That is a stronger statement than the sample set mattering as much as the lamp. It means the sample set changes what the comparison between two lamps says, not merely what each of them scores, and it does so multiplicatively. Two indices computed on two sample sets do not differ by an offset that could be calibrated away; they differ by a factor that depends on which lamps are being compared.
Which is exactly why a headline average with an excluded saturated sample is the worst arrangement available. The excluded sample is the one carrying most of the discrimination between lamps, and removing it does not merely lower the resolution — it removes the part where the lamps differ most.
The old halophosphate tube is the other end of the fluorescent family, and its white is further off the locus than any of them.
The Stokes loss the mechanism implies
The two mechanisms are set out and the efficiency consequence that follows from them is not, though it is one subtraction.
A phosphor absorbs a photon and emits a longer-wavelength one, and the energy difference goes to heat. For a fluorescent tube the absorbed photon is mercury’s 254-nanometre resonance line and the emitted one is somewhere near 550, so 54 per cent of the energy is lost before the light leaves the phosphor. For a white LED the absorbed photon is the die’s 450 and the emitted one is near 560, and the loss is 20 per cent.
That is a factor of two and a half in the unavoidable conversion loss, and it follows entirely from where each mechanism’s pump sits. The fluorescent tube pumps from the ultraviolet because a mercury discharge is what it is; the LED pumps from the blue because a nitride die is what it is, and blue is inside the band the phosphor is emitting into.
It also explains a structural feature of both spectra that the essay describes and does not connect. The LED’s blue spike is unconverted pump light doing useful work — it is part of the white — while the fluorescent tube’s mercury lines are unconverted pump light that happens to be visible, arriving in four narrow places nobody chose. One lamp’s leakage is its blue primary and the other’s is four spikes, and the difference is whether the pump was inside the visible band to begin with.
And it prices the trough. The LED’s hole at 480 nanometres exists because the die emits at 450 and the phosphor starts near 500, so the gap between them is the price of pumping from inside the band. A fluorescent tube has no such trough because its pump is outside the band entirely — which is the same fact that costs it 54 per cent instead of 20.
The narrowband source under the ten-degree observer is the extreme case of a white with no shape behind it.
Where the model stops
The spectra are constructions and the essay’s conclusions are stated to hold for constructions. Anything about a specific commercial lamp requires measuring that lamp.
The fidelity index is fidelity only. A source can render every sample faithfully and be unpleasant to sit under, and a source that slightly increases chroma is generally preferred to one that is exactly faithful — which is why modern systems report a gamut measure alongside a fidelity measure. Nothing here measures preference, and fidelity is not a synonym for quality.
And there is no geometry. Real illumination arrives from a distributed source with a spatial structure, reflects from a surface with a bidirectional reflectance distribution, and none of that is here. Every calculation on this page assumes a reflectance and an illuminant multiply, which is the model a fluorescent sample breaks and which spatial structure complicates further.
Whether the correlated temperature and the distance from the locus depend on the observer is a fair question, and the LED is the lamp where it matters most.
What the pictures cannot show
The rendering figures pair a swatch under the source with a swatch under the reference, and any seam is a rendering failure. Whether the reader sees a seam depends on their display, their room, and their adaptation state — so the numbers are printed too.
There is a sharper problem. Both swatches in every pair are being shown on the same display, under the same lamp, in the same room. What is being drawn is what the calculation says each would look like, side by side, to an observer adapted to each in turn. Nobody has ever seen those two swatches next to each other, because doing so would require being adapted to two lamps at once.
That is not a defect in the drawing. It is the definition of a corresponding-colour comparison, and it is the same difficulty that makes chromatic adaptation transforms so hard to measure.
Who found it, and when
The mercury discharge lamp dates to the 1900s and the phosphor-coated fluorescent tube to the 1930s. Halophosphate coatings dominated until the 1970s, when rare-earth triphosphors arrived — driven by the oil crisis and the efficiency it made urgent, rather than by any complaint about colour.
The CIE colour rendering index was standardised in 1965 and revised in 1974, and its limitations were documented almost immediately by the people who built it. It survived because it was the only number available and because it was written into procurement specifications, which are considerably harder to revise than standards.
White LEDs arrived commercially in the mid-1990s following Nakamura’s blue InGaN die, and their rendering was poor enough at first to make the index’s inadequacies a practical problem rather than an academic one: lamps were being sold with respectable CRI values and were visibly bad in exactly the region the eight samples did not probe. TM-30, published in 2015, reports fidelity and gamut separately and gives a per-hue breakdown, which is the same conclusion this essay reaches by measurement.
Where this goes next
The model of an object colour that these lamps still obey, and the case that breaks it, is some paper is brighter than white. What happens when a spiky spectrum meets an instrument of finite bandpass is what the instrument reports. And the narrow primaries that make rendering worst are also what makes the standard observer least reliable.
What this makes readable
Essays that name this one as a prerequisite.
- A lamp has a direction
- A lamp has a waveform
- A lamp switched on is not the lamp measured
- A screen is a poor lamp
- Some paper is brighter than white
- The filter that makes colour possible
- The index is one observer's opinion
- The sun is not one illuminant
- Three numbers cannot see a line
- Two lamps do not average
- What a lamp cannot give back
- What the instrument reports
- Where the grid starts
- Two matrices do not reach a white LED
- Three lines spare a slow pigment
- A fourth emitter spends the gap it fills
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.
- Only one dimmer is invisible correlated colour temperature · duv · led emission · planckian locus
- A surface that is not a multiplication fluorescence · metamerism · reflectance
- Two sheets that match until the window fluorescence · metamerism · reflectance
- A fourth dimension has a shape metamerism · reflectance
- A narrow channel has to be read on its own fluorescent · led emission
- A reflectance is a diagonal fluorescence · reflectance
What links here
The 8 essays that link to this one and share the most of its objects, of 35 that link here.
The objects this essay names
Each one links to every other essay that touches it.
Colour renderingCorrelated colour temperatureDuvFluorescenceFluorescentLED emissionMetamerismPlanckian locusReflectance