In the 1880s, a wealthy and relentlessly curious Victorian named Francis Galton did something almost nobody had bothered to do: he measured a lot of families. Hundreds of sets of parents and their grown children, heights recorded and plotted. He was hunting for how traits pass down the generations — and he found something so odd it looked, at first, like a mistake in the data.
The very tall parents did tend to have tall children. But not as tall as themselves. And the very short parents had short children — who were, on average, a bit taller than they were. Both extremes had children who leaned back toward the middle of the population. Galton called it “regression towards mediocrity in hereditary stature,” and the word he reached for — regression — is why, to this day, the entire statistical toolkit of fitting lines to data is called regression analysis. The name of a whole field is a fossil of this one afternoon of measuring families.
Before you read — take a guess
Galton found that unusually tall parents tend to have children who are tall but shorter than the parents, and unusually short parents have children taller than themselves. Over generations, what does this do to the overall spread of human height?
What Galton actually saw
Plot every child’s height against the average of their two parents (Galton’s “mid-parent” height) and you get a cloud of dots with a clear upward tilt: taller parents, taller children. If height were passed on perfectly, the dots would sit on the 45° line — child exactly matches parents. They don’t. The best-fit line through the cloud is flatter than 45°. It tilts up, but gently.
That flatter slope is regression made visible. Read it at the tall end: parents who are, say, 15 cm above average have children only about 10 cm above average — the prediction slides back toward the mean. Read it at the short end and the same line lifts the children of short parents up toward the middle. The slope of that line is the reliability from lesson 01: the fraction of the parents’ deviation that carries forward.
Where the word comes from
Galton’s flatter-than-45° line was the first “regression line.” When statisticians later generalized the technique of fitting a line (or any model) to a cloud of data, they kept his word. So every time you hear “run a regression,” you’re hearing an echo of tall parents and their slightly-less-tall children.
Why it happens — same signal + noise, now in biology
Nothing here is special to genetics. A child’s height is a signal (the heritable, shared component passed from the parents) plus noise (the other parent’s contribution, developmental luck, childhood nutrition, the roll of which genes actually got expressed). Very tall parents are usually tall for two reasons at once: genuinely tall-making genes and a favorable roll of all the other factors. Only the first part reliably passes to the child. The lucky roll doesn’t come along for the ride — so the child, on average, lands closer to the population mean than the parents did.
It’s the exam-winner from lesson 01, wearing a different costume. Select an extreme (very tall parents), and you’ve selected partly for a favorable noise draw that the next generation won’t inherit.
Why do the children of exceptionally tall parents tend to be shorter than their parents (on average), in signal-plus-noise terms?
The paradox resolved: it works in both directions
The thing that makes people’s heads hurt: if tall parents have shorter kids and short parents have taller kids, shouldn’t everyone funnel into one identical middle over the centuries? No — and seeing why cements the whole model.
Regression is a statement about prediction from a selected extreme, not a physical compression of the population. Two facts keep the spread alive:
- It runs both ways. Just as the children of the tallest regress down, the children of the shortest regress up. Neither tail is being pushed inward preferentially; they’re both pulled toward the center — which is symmetric, not shrinking.
- The extremes get refilled. Perfectly average parents don’t have perfectly average children — they have a scatter, and some of that scatter lands way out in the tall or short tail (their own favorable/unfavorable noise draw). Those newly-minted extremes replace the ones that regressed inward. The tails are constantly restocked.
Run both effects at once and the population’s distribution is in a steady state: generation after generation, the same bell-shaped spread of heights. Individuals’ predictions regress; the population does not shrink.
Don't confuse 'this individual, predicted' with 'the whole group, over time'
Regression to the mean predicts that a randomly re-measured extreme will be less extreme. It does not predict that variety collapses. The most common misuse of this model is inferring a doomed convergence — “everyone’s becoming average” — from a fact that only ever described the follow-up to a selected extreme.
Worked example: the top decile that drops on its own
Let’s put numbers on it with a sales team instead of heights, so the “nothing changed” point is unmistakable. Model each rep’s monthly result as true ability + luck, with a population mean of 100. Suppose the reliability works out to 0.6 — that is, 60% of the spread between reps is real ability and 40% is luck (deals happening to close, a hot lead, a rival stumbling). The prediction rule from lesson 01:
Take a few individual reps from this month, and predict next month:
| Rep | This month | Above mean | Expected next month |
|---|---|---|---|
| Ana | 150 | +50 | 100 + 0.6×50 = 130 |
| Ben | 140 | +40 | 100 + 0.6×40 = 124 |
| Cai | 130 | +30 | 100 + 0.6×30 = 118 |
| Dee | 100 | 0 | 100 + 0.6×0 = 100 |
| Eve | 60 | −40 | 100 + 0.6×(−40) = 76 |
Now suppose we hand out a “Top Performers” bonus to this month’s top decile, whose average was 130. What does the model predict for that same group next month? Their average deviation was +30, so their predicted next-month average is . Watch what happens to the two numbers side by side:
| Top decile (selected) | Whole population | |
|---|---|---|
| This month’s average | 130 | 100 |
| Predicted next month | 118 | 100 |
| Change | −12 | 0 |
The top decile’s average is expected to fall from 130 to 118 — a 12-point drop — while the population mean sits perfectly still at 100. Nobody got lazy, nobody burned out, no bonus “backfired.” The group was scooped up at a lucky peak, and the peak deflated on its own. If a manager had done anything to that group — a pep talk, a new quota, a stern email — the 118 would have arrived anyway and the manager would have “learned” its effect. (That trap is lesson 03’s whole subject.)
Perfectly symmetric. A bottom decile averaging 70 (that is, −30 below the mean) is predicted to rise to 100 + 0.6×(−30) = 82 next month — up 12 points — with no coaching, no intervention, nothing. This is why the “worst” units so often improve after any action taken against them: they were caught in an unlucky trough, and troughs fill back in. Punish them and the punishment gets the credit.
Using the sales model above (mean 100, reliability 0.6), a rep posts 200 this month. What's the model's best guess for next month — and what's the trap in reading it?
When to use it
Any time you rank people, teams, schools, hospitals, or investments and then focus on the top or the bottom of the ranking, Galton is looking over your shoulder. Before you re-measure that selected group and attribute the change to anything, subtract the regression you should have expected. The steeper the luck component, the bigger the built-in drift you’d otherwise mistake for a real effect.
Recap
Why is the entire statistical method of fitting lines to data called 'regression'?
Check your answer to continue.
Where this goes next
Galton showed us regression is real, ancient, and named a field. Lesson 01 gave us the machine; this lesson worked the numbers and killed the “everyone converges” myth. Lesson 03 turns the model loose on the real world — and on your wallet. When you act right after an extreme, the reversion we just computed shows up on schedule and takes a bow for your intervention. That’s the regression fallacy, and it is one of the most expensive thinking errors there is.