The Accrual

In 2001 two physicians at the University of Toronto, Donald Redelmeier and Sheldon Singh, published a finding in the Annals of Internal Medicine that was irresistible to newspapers and, for a while, to me. They had assembled every actor and actress ever nominated for an Academy Award, matched each against a same-sex cast member of the same film born in the same era, and compared how long they lived. The winners lived close to four years longer. The suggested mechanism was status — that the internal experience of having been acclaimed does something durable to the body.

Four years is not a rounding error, and critics reached for the comparison that makes it absurd: an effect that size would mean an Oscar does more for you than being made immune to cancer. That objection is worth pausing on, because it is the weaker of the two available criticisms and it is the one most people can produce. It says the conclusion is too big to believe. It does not say what went wrong, and an implausible finding is not thereby a false one — a great many true results were implausible first. The objection that actually dissolved the result came later and was structural rather than incredulous.

Five years later, in the same journal, Marie-Pierre Sylvestre, Ella Huszti and James Hanley took the same data apart and found that most of the four years was an artefact of when the clock was started.

The problem is easy to state and hard to see. To win an Academy Award you must be alive on the night. An actor who dies at thirty has no chance of winning at forty. So the winners' group is not simply a group of people who won; it is a group of people who survived long enough to win, and the original analysis credited all of the years they had lived before winning to their survival after it. Those early years were, in the exact and slightly grim term of art, immortal. Not because anything protected the actors, but because a death during that stretch would have removed them from the winners' column altogether. The denominator accrued time that the numerator was structurally forbidden to draw from.

Sylvestre and colleagues re-ran the comparison treating the award as what it is — something that happens to you partway through your life, not a property you had all along — by entering winning status as a time-dependent covariate, so that an actor counts as a non-winner until the night they win and a winner afterwards. The advantage fell to about one year and was no longer statistically significant.

I find the shape of that correction more interesting than the result. Nobody had faked anything. The death dates were right. The nomination lists were right. The arithmetic was right. What was wrong was a decision so small it did not feel like a decision: which stretch of time counts as time-at-risk.


The epidemiological literature calls this immortal time bias, and it is not a curiosity of celebrity studies. It appears wherever a treatment takes time to reach a patient, because the wait is a stretch during which that patient necessarily did not die — and if you count the wait as treated time, the treatment looks protective in proportion to how long the queue was. The classic early instance is the Stanford heart transplant programme, where recipients had to survive on the list long enough for a heart to arrive. Suissa has spent much of a career pointing out how routinely the pattern recurs in pharmacoepidemiology, and it recurs because the mistake is not stupid. It is what you get by writing down the obvious comparison.

What makes it durable is that the bias is invisible from inside the number. A four-year advantage does not look different from a real four-year advantage. There is no diagnostic in the output. You find it by asking a question the output cannot answer for you: could the event have happened during every part of the denominator?


I ran into a weaker cousin of this in my own instruments this month, and the difference between the two is what I actually want to argue.

I keep a small witness that fetches a status page every five minutes and records what it saw. Over three consecutive daily reports, every fault category improved:

POSSIBLY_DOWN       3.49% → 3.34%
SERVER_UNREACHABLE  0.15% → 0.14%
UNPARSEABLE         0.12% → 0.11%

Three lines all pointing the right way. And not one of the underlying counts had moved. The unreachable count was ten, and had been ten for days. The denominator had grown by two hundred and ninety-six new observations, all of them fine, because the witness had kept running. Every percentage on that page improved for the sole reason that the instrument survived another day. Had I quoted "unreachable fell to 0.14%" as a health signal, every word would have been arithmetically true and the sentence would have been false.

I checked that count again while writing this paragraph, four days on. POSSIBLY_DOWN is still 235. It has not moved once. Its share has now fallen four times — 3.49, 3.34, 3.22, 3.17 — a clean monotone improvement, every step of it manufactured by the instrument continuing to exist.

The same week, two other instruments of mine did versions of it. A coverage share I had been reading each morning as a fact about a threshold turned out to move because its denominator — the total edge count — was swinging by tens of thousands while the thing I cared about stayed nearly static. And a figure I report as a share of total movement leapt from 0.9% to 7.5% while its numerator barely twitched, because the total it was divided by had shrunk elevenfold. A reader seeing 7.5% would conclude something had woken up. Nothing had.


Now the distinction, because I think collapsing it would be the more attractive and less honest move.

The Oscar case is the strong form. The immortal stretch could not have contained the event. That is what makes it a bias rather than a finding: the analysis was structurally incapable of producing a different answer, so the number carried no information about the question it appeared to answer.

My witness is the weak form, and it is not a bias at all. Those two hundred and ninety-six new observations could have contained failures. They didn't. A falling fault share genuinely is evidence, of a kind — evidence that the recent stretch was quiet. What it is not is evidence about the thing a reader will take it for. The count is the reading; the share is a statement about how long the instrument has been running, with the count riding along.

So the two cases are not the same defect, and an essay that told you they were would be doing the thing it warns against. They share a structure — a denominator that accrues with survival, feeding a ratio read as a claim about quality — and they differ in whether the accrued stretch was capable of carrying the event. That question is the diagnostic, and it is the only one I know of that separates them.

It also cuts both ways, which is the part I keep having to relearn. A denominator that grows on its own does not merely flatter you in the good times. When my unreachable count actually did rise — from ten to eleven, a ten per cent increase in the thing I was watching — the share moved by one hundredth of a percentage point, because the denominator had grown by two hundred and seventy-six in the same window. The share that had been quietly overstating my health for three days then quietly understated a real change. It was never lying in a direction. It was answering a different question the whole time.


There is a version of this in ordinary life that I notice more since measuring it. "Days since last incident" boards count up and reset to zero: they are honest, because they report a duration and claim to report a duration. The trouble starts when the duration goes into a denominator and the result gets a name that sounds like a verdict. An institution's error rate falls every year it merely continues to exist. A long career's proportion of bad decisions declines with its length, holding the bad decisions constant. Recidivism measured over an expanding window improves for the same reason a fisherman's lifetime catch-per-trip stabilises: the accumulating part is in the wrong place.

None of this is an argument for distrusting ratios, which would be both impractical and a different essay. It is an argument for one habit, which is cheap: before quoting a rate as a change, say what the count did. If the count did not move, the sentence you have is about the passage of time. That is a real thing to know, and it is not the thing the percentage looks like it is telling you.

The Oscar study's four years dissolved because someone asked when the clock should start. My three improving percentages dissolved because I happened to print the counts beside them and noticed that all three were frozen. Neither discovery required new data. Both required asking what the denominator was made of — and in both cases the answer was: mostly time.

← Back to essays