What light is

The index is one observer's opinion

A colour rendering index is printed to a tenth of a point and computed for a single set of colour-matching functions that no standard names as a choice. Scored by two hundred eyes instead, one lamp's spread is wider than the whole range the standard observer puts five lamps in — and two lamps the standard separates by four tenths of a point come out the other way round for every member of the population.

Assumes A lamp is not a blackbody and Whose eyes.

Every lamp sold carries a number for how well it renders colour. The number is computed by rendering a set of test surfaces under the lamp and under a reference of the same colour temperature, and measuring how far apart they land. Everything in that recipe is stated in the standard except one thing.

The rendering is done through a set of colour-matching functions, and the standard does not present that as a choice. CIE Ra uses the 1931 observer; CIE 224’s Rf uses the 1931 observer; the index computed on this site uses the 1931 observer. None of the three documents contains a sentence beginning for an observer whose.

A colour rendering index, computed for everybody instead of for one observer. Each bar spans what 120 eyes make of one lamp: the same reflectances, the same reference illuminant, the same arithmetic, different colour-matching functions. The mark is the standard observer's own answer, which is the number printed on the box. One lamp's spread is 3.8 points wide while the whole range the standard observer puts these lamps in is 1.8 — so a ranking quoted to a tenth of a point is a statement about one set of tables. Two of the marks fall outside the population's range entirely.
Fig. 1 Four lamps and a reference, each scored by a hundred and twenty eyes. The bar spans what the population makes of one lamp; the mark is the standard observer’s answer, which is the number on the box. Two of the marks fall outside the population’s range entirely.
Two lamps, ranked by everybody. The difference in fidelity index between a phosphor-converted white LED and a triphosphor tube, one observer at a time. The standard observer puts the LED 0.40 points ahead — the marked line — and 100 per cent of the population puts it behind, by as much as 6.2 points. The spread inside the histogram is individual variation; the distance between the histogram and the mark is something else, and it is larger: it is the difference between two models of an observer, and neither lamp has anything to do with it.
Fig. 2 The consequence, on the pair the standard’s ordering is most often quoted for. The difference in index between a white LED and a triphosphor tube changes sign across the population, so which of the two renders better is a question about the reader.

A hundred and twenty observers is a sample, and the two measurements are worth repeating on a larger one before anything is concluded from either.

A colour rendering index, computed for everybody instead of for one observer. Each bar spans what 200 eyes make of one lamp: the same reflectances, the same reference illuminant, the same arithmetic, different colour-matching functions. The mark is the standard observer's own answer, which is the number printed on the box. One lamp's spread is 3.8 points wide while the whole range the standard observer puts these lamps in is 1.8 — so a ranking quoted to a tenth of a point is a statement about one set of tables. Two of the marks fall outside the population's range entirely.
Fig. 3 The same four lamps scored by two hundred eyes rather than a hundred and twenty. The spread does not narrow, which is what says it is a property of the population rather than of the sample size.
Two lamps, ranked by everybody. The difference in fidelity index between a phosphor-converted white LED and a triphosphor tube, one observer at a time. The standard observer puts the LED 0.40 points ahead — the marked line — and 100 per cent of the population puts it behind, by as much as 6.2 points. The spread inside the histogram is individual variation; the distance between the histogram and the mark is something else, and it is larger: it is the difference between two models of an observer, and neither lamp has anything to do with it.
Fig. 4 And the reversal on the larger sample. The fraction of observers who put the two lamps the other way round is stable, so the standard’s single number is naming a majority rather than a fact.

The claim

A colour rendering index quoted to a tenth of a point is a statement about one set of tables, and the tables are worth several points.

  • One lamp’s spread across a population is 3.8 points wide, while the whole range the standard observer puts four real lamps in is 1.7. The spread inside a single number is larger than the differences that number is used to make.
  • The reference scores a hundred for everybody, exactly, which is not a result — it is the definition, and it is the one row of the table that carries no information.
  • Two lamps separated by four tenths of a point come out the other way round for every member of the population, by as much as six points.
  • And that reversal is not individual variation. The spread inside the histogram is people differing; the distance between the histogram and the standard observer’s mark is two models of an observer differing, and it is larger.

What the index is

The method is common to every fidelity index and has three parts. Take a set of test reflectances. Render each under the lamp and under a reference illuminant of the same correlated colour temperature — a Planckian radiator below 5000 K, a daylight phase above it. Adapt both to a common white and measure how far the pairs land apart.

What differs between indices is the samples, the space and the scale factor. CIE Ra uses eight tabulated samples in a 1964 uniform space with a factor of 4.6; the index here uses constructed samples, CAM16-UCS, and a factor chosen so the numbers land on a familiar range. The scale factor is presentational and is stated rather than fitted; the quantity that means something is the mean shift in appearance units underneath it.

None of that is what this essay is about. All of it is held fixed. The only thing that changes from one score to the next is which eye the integrals are taken against.

Two hundred eyes, one lamp

The population is the site’s own: members differing in five measured variates — lens age, macular density, cone optical density and the three pigment peak wavelengths — each of which was already an argument to the retinal model that nobody had varied.

The reference illuminant is chosen once, from the standard observer’s correlated colour temperature, and held fixed across the population. That is the conservative construction and it is also what a laboratory does: the reference is picked from the measurement, and everybody then looks at the same pair of lights.

The result is a distribution per lamp. The widest belongs to the white LED at 3.8 points from the lowest score to the ninety-fifth percentile; the narrowest of the real lamps is the triphosphor tube at 2.7. The reference radiator scores 100.000000 for every member, exactly, because it is being compared with itself.

That last row is worth a moment. A perfect score is not evidence about a lamp; it is a tautology in the definition, and every fidelity index has it. The interesting question — how well does a lamp render colour — is answered relative to a source chosen to be its own answer.

The spread is wider than the ranking

Four real lamps, scored by the standard observer, span 1.7 points: 81.2 for the triphosphor tube up to 82.9 for three narrow emitters. Those 1.7 points are what a purchasing decision or a specification is made of.

One lamp’s spread across the population is 3.8.

A ranking is a statement about a difference, and a difference smaller than the spread of opinion behind it is not a ranking — the same objection a mean makes to a comparison from a different direction. There is nothing subtle about this arithmetic and it is not visible anywhere in the practice, because the practice has one observer in it and a single observer produces a single number with no error bar attached.

The reversal, and what causes it

Two lamps, ranked by everybody. The difference in fidelity index between a phosphor-converted white LED and a triphosphor tube, one observer at a time. The standard observer puts the LED 0.40 points ahead — the marked line — and 100 per cent of the population puts it behind, by as much as 6.2 points. The spread inside the histogram is individual variation; the distance between the histogram and the mark is something else, and it is larger: it is the difference between two models of an observer, and neither lamp has anything to do with it.
Fig. 5 The difference in index between a white LED and a triphosphor tube, one observer at a time. The standard observer puts the LED four tenths of a point ahead; the population puts it behind, unanimously, by up to six.

The histogram sits entirely on the other side of zero from the mark. That is a stronger result than a spread and it needs a more careful explanation than people differ, because a spread that straddles the mark would look like this only by accident.

It is not individual variation. It is that the median member of this population is not the 1931 observer, and cannot be: the cone fundamentals here are built from a pigment nomogram seen through modelled ocular media, and the best linear map from them to the standard functions leaves a residual — under a unit on natural reflectances, which the population machinery asserts and bounds rather than removing.

Under one lamp that residual costs 0.9 points. Under another it costs 4.6. The swing across the four lamps is 5.15 points, and the gap being ranked is 0.4.

So the honest statement is not the population disagrees with the standard observer about which lamp is better. It is:

Which model of an observer is used moves a fidelity index by several points, it moves it by a different amount for every lamp, and both of those are larger than the differences the index is used to decide.

That is a claim about the standard’s silence rather than about anybody’s eyes. Two laboratories using two defensible observer models would rank these lamps differently and neither would have departed from any published method, because no published method says which observer to use.

Narrow bands are scored less reliably

The spread across the population is not the same for every lamp, and the direction it varies in is the direction the industry has been moving for thirty years.

A lamp made of narrow emitters is scored less consistently than a broader one, because a narrow band samples the differences between observers instead of averaging over them. Two people whose long-wavelength pigment peaks differ by three nanometres see almost the same thing under a smooth radiator, where the difference is integrated away across the whole spectrum; under a source that is mostly a spike at 615 nanometres they do not, because the spike lands on a different part of each of their sensitivity curves.

This site has the same result from the display side, where narrowing the primaries from forty nanometres to two raises the population’s ninety-fifth percentile from 11.3 units to 17.9. The lighting side is the same mechanism with a different apparatus — though the four lamps here do not demonstrate it, for a reason the section below has to work out — and it points the same way: every generation of lamp and panel has narrower emission than the one before, and every generation is therefore a little more dependent on whose eye is looking at it.

The narrowness claim is imported and these four lamps do not test it

The section above argues that a narrow-banded lamp is scored less consistently than a broad one, and the mechanism it gives is sound. The evidence it offers is from the display side, where narrowing primaries from forty nanometres to two raises the population’s ninety-fifth percentile by more than half. The four lamps in this essay do not show it, and two of them show the opposite.

Reading the spreads off the table by which lamp is narrower: three narrow emitters span 3.3 points and a triphosphor tube spans 3.0 — the two most narrow-banded sources here, and the two tightest distributions. A halophosphate tube, which is a broad phosphor, spans 3.6. A phosphor-converted LED, which is a narrow pump line beside a broad converter, spans 4.7 and is the widest.

The essay’s own two quoted figures say the same thing more starkly, since they are the two extremes: the widest belongs to the white LED and the narrowest to the triphosphor tube, and of those two the LED is the one with most of its power in a broad phosphor.

That is not a refutation of the mechanism. Narrowing a display’s primaries really does widen the observer spread, this site measured it, and the reason is exactly the one given — a narrow feature samples the differences between people’s sensitivity curves instead of integrating over them. What the four lamps say is that narrowness is not the only property that does it, and on this small set something else dominates.

The candidate the spectra suggest is a gap. A phosphor-converted LED is not simply narrow or broad; it is a narrow blue line, a hole, and a broad yellow hump — and a hole in the spectrum is as good a sampler of observer differences as a spike, because what matters is that the light has structure at a scale comparable with the differences between two people’s curves. A triphosphor tube’s three bands are narrow but they are spread across the band and there is emission between them; a halophosphate’s phosphor is broad and smooth. On that reading the ordering comes out right, and it is a reading rather than a measurement.

And the spread is not the interesting column anyway. Setting it beside the bias — the gap between the standard observer’s answer and the population’s median — the two move together, at a rank correlation of +0.6 over the four: the LED and the halophosphate are the two with the widest spreads and the two with the largest biases, at 4.1 and 4.5 points, and the two narrow-banded lamps have the two tightest spreads and biases of 0.9 and −0.6.

That is worth more than either column alone, because it says the two silences in the method are not independent. A lamp whose score depends most on which observer model is used is also the lamp whose score varies most among observers of one model. Both are consequences of the same property — how much of the answer is decided by a few narrow regions of the spectrum rather than by the whole of it — and a lamp that is safe on one count is safe on the other.

Which gives the essay’s conclusion a cheaper form than it currently has. The full argument requires a population, a second observer model and two hundred renderings. The shortcut is to look at the spectrum: a source with structure at a scale of tens of nanometres carries a score that is soft in both senses, and a source without it does not. That is not a substitute for the measurement and it is available before anybody has one.

What a number with no observer in it looks like

It is worth being concrete about how the omission is written down, because it is not a slip and nobody hid it.

The recipe for CIE Ra is a page long and every quantity in it is specified: the eight test colours by their spectral reflectances, the reference by its correlated colour temperature and a stated switch from Planckian to daylight at 5000 K, the adaptation by a von Kries transform in a named cone space, the summary by a mean and a scale factor. The colour-matching functions appear as the colour-matching functions — the definite article doing the work, because in 1965 there was one set that anybody used and the alternative had not been published.

That is a reasonable way to write a standard and it has an unreasonable consequence sixty years later. There are now at least three defensible answers to which observer: the 1931 functions, the 1964 ten-degree functions, and the physiologically-derived cone fundamentals the CIE has published since. A laboratory that used the second would not be departing from the method, because the method does not mention the first.

The reference, which is the other silence

A fidelity index measures a distance from a reference, and the reference is chosen by the lamp’s own correlated colour temperature. That has a consequence the number’s users mostly know and its form conceals: two lamps with different colour temperatures are not being compared with each other. Each is compared with its own reference, so a score of 82 for a 2700 K source and 82 for a 5000 K one are two different measurements that happen to produce the same digit.

Colour temperature is itself a projection — the nearest point on a locus, with the distance off it thrown away — so the reference is selected by a number that has already discarded information about the lamp. And below 5000 K the reference is a Planckian radiator while above it is a daylight phase, which is a discontinuity in the definition rather than in anything physical.

None of that is this essay’s argument, and all of it compounds with it. The observer is one silence in the method; the reference selection is another; and both act before the lamp is compared with anything.

What was computed, and how

Each member’s tristimulus values come from their own cone absorptances through a fitted cone-to-XYZ matrix, and the matrix is the same for every member — so a difference between two people is a difference in what their cones caught and never a difference in bookkeeping.

The appearance calculation is CIECAM16 with each observer’s own view of both whites, so a member who sees the lamp as slightly warmer adapts to their own slightly warmer white rather than to the standard observer’s. Both the test and the reference rendering are done that way, which is what keeps the comparison inside one head.

The scale factor and the sample set are held at the values the site’s own index uses, so nothing here is a comparison between indices. assertARenderingIndexIsOneObserversOpinion requires the population spread to be points wide and requires the reference to score exactly a hundred for every member, because a reference that drifted would mean the machinery was wrong rather than the number being soft.

The five lamps, in one table

lamp on the box the population’s median its range the box minus the median
three narrow emitters 82.9 83.5 81.3–84.6 −0.6
a phosphor-converted LED 81.6 77.5 74.9–79.6 +4.0
a triphosphor tube 81.2 80.3 78.9–81.9 +0.9
a halophosphate tube 82.4 77.9 75.7–79.3 +4.6
a filament, at 2856 K 100.0 100.0 100.0–100.0 0.0

The last column is the finding compressed into five numbers. If the gap between the two observer models were a constant it would be a calibration and nobody would need to care which model was used; the standard observer would simply read a little high or a little low and every comparison would survive. It is not a constant. It is −0.6 for one lamp and +4.6 for another, which is exactly the situation in which a ranking cannot be trusted to survive a change of tables.

Where it stops

The population is a model. Its five variates are quoted with the ranges the literature reports and they are sampled independently, which is not right — macular density and lens density are correlated in real people, and both are correlated with age, which is one of the five. Treating them as independent widens some distributions and narrows none, so the spreads here are an upper estimate.

The reference illuminant is held fixed across observers. A more thorough treatment would let each member compute their own correlated colour temperature and pick their own reference, which would move the answer in an interesting direction and would also stop the members being comparable.

And nothing here says which observer is right. The claim is that the answer depends on that choice by more than the differences it is used to settle, which is true whichever way the choice is made, and would remain true if a better set of colour-matching functions replaced both candidates tomorrow.

Where the ladder goes next

Two directions, and they part at the question of what a rendering index is for.

If it is a specification — a number a manufacturer must meet — then the useful next quantity is a tolerance interval rather than a point: the score such that ninety-five per cent of observers report at least it. That is computable from what is already here and would change which lamps pass.

If it is a comparison — a way to choose between two products — then the useful next quantity is the probability that the ranking survives, which is what the histogram above is a picture of. Neither is a new measurement. Both are the same integrals taken against more than one eye, which is a decision somebody made once, in 1931, for a different purpose.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the index of named objects makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 10 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AcceptabilityCIECAM16Colour renderingCorrelated colour temperatureFluorescentIndividual variationObserver metamerismQuality controlSpecificationSpectral power distributionStandard observerWhite LED