Here is a recipe for looking like a genius. Wait until something is at its worst — the sickest patient, the deadliest intersection, the coldest sales month, the ugliest crash statistics. Then do something. Anything. A treatment, a policy, a stern memo, a lucky charm. Now wait again. The numbers will improve, because extremes regress toward the mean all by themselves — and you, standing there having Done A Thing, will collect the credit. Every time.
This is the regression fallacy: attributing a change to an intervention when regression to the mean alone would have produced it. It is not a rare or exotic bug. It is the default way humans read a before-and-after when the “before” was an extreme, and it quietly underwrites bad medicine, bad policy, bad management, and a fortune wasted on things that never worked.
Before you read — take a guess
A city installs speed cameras at the 10 intersections with the worst crash counts last year. This year, crashes at those intersections drop sharply. What's the safest conclusion?
The anatomy of the fallacy
Every instance has the same two-step skeleton. Spot the skeleton and you spot the trap:
- Selection on an extreme. You act because a measurement was unusually bad (or unusually good). The sickest patient seeks help; the worst sites get the camera; the panicking parent hires the tutor after the disastrous exam.
- The reversion arrives — and gets adopted. Because the starting point was extreme (partly luck), the next measurement drifts back toward the mean on its own. But you did something in between, so the natural drift is misread as your effect.
The fallacy is so seductive precisely because the improvement is real and reliable. It’s not that nothing happened — the patient does feel better, crashes do fall. It’s that the improvement was already scheduled by regression, and your intervention is taking a bow for the tide going out.
The tell
The alarm bell is the word because. “We acted because the numbers were terrible.” The moment your reason for intervening is that a measurement hit an extreme, you have selected on an extreme — and regression is now standing between you and any honest read of whether your action did anything.
Case 1: the Sports Illustrated “cover curse”
For decades, athletes and teams joked that appearing on the cover of Sports Illustrated was cursed: soon after the cover, the star would slump, the team would lose, the winning streak would die. Some people half-believed the jinx was real.
It’s regression wearing a superstition costume. You land the cover because you just did something extreme — a record streak, a career-best run, an against-all-odds championship. That peak was your real talent plus a run of good luck. The luck doesn’t re-sign for next season. So your performance drifts back toward your (still excellent, but less superhuman) normal — and since the cover came right before the dip, the cover gets blamed. No curse required, just selection on a peak followed by inevitable reversion.
What is the actual mechanism behind the 'Sports Illustrated cover curse'?
Case 2: speed cameras and worst-sites-first
This one costs public money and can distort real safety decisions, so it’s worth getting exactly right. Road authorities often site cameras (or any intervention) using a “worst first” rule: rank intersections by recent crashes, treat the top of the list. It’s an intuitive, defensible-sounding policy — and it bakes regression straight into the evaluation.
A site lands on the “worst” list partly because its underlying risk is genuinely high (signal) and partly because it happened to have a spike of crashes that year (noise — crashes are rare, clustered, and luck-driven). Next year the noise reshuffles and the count falls toward the site’s true level, whether or not you installed anything. Study after study has found that naïve before/after evaluations of such programs overstate their benefit, because they credit the camera for the regression. The cameras may well help — but the honest effect is far smaller than the raw drop suggests.
Why 'worst-first' is a regression machine
Selecting the units with the most extreme recent readings guarantees you’ve selected the ones most inflated by bad luck. Their reversion is not evidence your program works — it’s the arithmetic of having chosen them for an extreme. Any policy that targets “the worst N” and evaluates itself by before/after is measuring regression plus its effect, tangled together.
Case 3: feeling better after seeing the doctor
When do you finally book the appointment, try the supplement, or start the new regimen? When symptoms are at their worst — the back pain peaks, the cold is miserable, the mood bottoms out. Many conditions are self-limiting: they wax and wane and ease on their own. So you seek help at the trough, the condition regresses toward its average (better) state, and whatever you took the credit rolls in.
This is the engine behind glowing testimonials for remedies that do nothing. “I was in agony, I tried it, and within days I was fine!” is exactly what regression predicts even for a completely inert treatment, because you started at a low point selected for being extreme. It’s why “it worked for me” is such weak evidence, and why medicine cannot trust anecdotes — it needs controlled trials. (Lesson 04 pushes deeper into how this breeds full-blown superstition and quackery.)
A friend swears an herbal remedy cured their week-long migraine within a day of taking it. Why is this weak evidence the remedy works?
The fix: a control group
If regression fools before/after comparisons, what beats it? A control group — a comparable set of units that you measure but do not treat. Here’s the magic: both groups regress. If you select the worst sites and split them into treated and untreated, both halves will drift back toward the mean next year by roughly the same amount, because both were selected for the same extreme. Whatever’s left over — the difference between treated and untreated — is the part regression can’t explain. That difference is your real effect.
Why controlled trials exist
The whole apparatus of the randomized controlled trial is, in large part, a machine for subtracting regression to the mean (and other confounds). Give the drug to one randomly chosen half of sick patients and a placebo to the other. Both halves were sick, both regress; only the gap between them measures the drug. “Before vs after” on the treated group alone is nearly worthless when the “before” was an extreme.
Watch it work on numbers. Suppose we take crash-prone intersections (selected for a bad year) and randomly camera half of them, leaving the other half untreated as controls:
| Group | Crashes last year (selected extreme) | Crashes this year | Change |
|---|---|---|---|
| Treated (cameras) | 20 | 12 | −8 |
| Control (nothing) | 20 | 14 | −6 |
| Real camera effect | — | — | −2 |
The naïve read brags “cameras cut crashes by 8 (a 40% drop)!” But the untreated control sites fell by 6 all on their own — that’s the regression. The honest camera effect is the difference, just −2. Without the control group, you’d have credited the cameras with four times their real benefit and possibly rolled out an expensive program on a mirage.
Second-best defenses when a true control is impossible: (1) use a long baseline instead of one extreme year, so you’re not selecting on a single lucky/unlucky reading; (2) compare against similar untreated sites/regions that weren’t chosen for being extreme; (3) explicitly forecast the regression you’d expect with no intervention, and only claim the effect beyond that. The worst option — sadly the most common — is a bare before/after on the group you selected precisely because it was extreme.
A hospital gives a new therapy to its most severely depressed patients. Three months later they've improved markedly. What single design change would best reveal whether the therapy actually helped?
When to use it
Raise the regression-fallacy alarm on any before/after or “we did X and things improved” claim where the “before” was an extreme, or where the units were chosen for being extreme (the worst schools, the sickest patients, the crashiest roads, the coldest quarter). Ask two questions: Were these units selected for an extreme? and Is there a control group that regressed too? If the answer is “yes” and “no,” treat the reported effect as regression until proven otherwise.
Recap
What defines the regression fallacy?
Check your answer to continue.
Where this goes next
The regression fallacy is expensive because the improvement it hands you is genuine — it just wasn’t yours. Lesson 04 takes this to its most unsettling conclusion: when the feedback we give (praise, punishment) lands right after extremes, regression teaches us the exact opposite of the truth about what works. Kahneman’s flight instructors learned that screaming works and praise backfires — and they were dead wrong for a beautifully mathematical reason.