Skip to content
Mental Models

The Evolution of Cooperation

The Shadow of the Future

One-shot dilemmas have no tomorrow, so defection wins. Add a future and the payoffs invert: when the continuation probability w is high enough, cooperating becomes the selfish move. The master key of the whole course, with real numbers.

13 min Updated Jul 5, 2026

Last lesson left the paradox airtight: in a single, anonymous prisoner’s dilemma, defection is the dominant move, and no amount of good character rescues you. If that were the whole story, the living world would be a desert of cheats. It isn’t. So something the one-shot matrix leaves out must be doing the heavy lifting — and this lesson names it, measures it, and shows exactly how much of it you need.

The missing thing is almost embarrassingly simple. It’s a tomorrow. This is the master key of the entire course: get it in your hands here and the next four lessons are just variations on the theme.

The missing ingredient: a future

Picture a market stall in a town you’ll never visit again. The vendor can hand you a bruised peach hidden under the good ones, take your money, and never face you again. Now picture the greengrocer at the end of your own street, whom you’ll buy from every week for a decade. Same transaction, wildly different incentive — and the only thing that changed is whether there’s a next time.

The one-shot dilemma has a fatal blind spot: there is no next time in which your betrayal can be punished. Defect, grab the temptation payoff, walk away — the victim never gets to retaliate, because the game is over. But almost no relationship that matters is genuinely one-shot. Your supplier, your colleague, your neighbour, your own gut bacteria: you’ll deal with them again next week and next year.

A repeated game (or iterated game) is simply the same game played over and over by the same players, who remember what happened last time. That memory is the whole trick. Once a defection today can be answered by a defection tomorrow, your present choice stops being free — it buys a future. The prospect that today’s move shapes the other player’s future moves is called the shadow of the future: the looming prospect of more rounds, reaching back to discipline what you do right now.

Info:

Define it once, precisely

The shadow of the future is the disciplining pressure that the prospect of future rounds exerts on a present choice. A long shadow (many rounds likely, each one valued) makes betrayal expensive; a short shadow (the end is near, or future rounds barely count) lets betrayal off the hook. Everything in this lesson is a way of measuring how long that shadow is.

Why repetition rewrites the payoffs

In the Nash Equilibrium course there’s a lesson with a blunt title: Change the Game, Not the Players. Its whole point is that a bad equilibrium isn’t a moral failing of the people in it — it’s the resting point of the payoffs they face — so if you want different behaviour, you rewire the payoffs, not the souls. Repetition is that move in its purest, most natural form. Nobody legislates it, nobody signs a contract; time itself edits the matrix.

Here’s the edit, stated as a comparison of what you’re actually weighing:

  • One-shot: you weigh the temptation to defect now against nothing. There is no tomorrow on the other side of the scale, so temptation wins by default.
  • Repeated: you weigh the temptation to defect now against all the future cooperation you would destroy by defecting. Betrayal today doesn’t just cost you the victim’s trust this round — it poisons every round after, because a player who remembers will stop cooperating with you.

When that second term is heavy enough, cooperation wins — and notice why it wins. Not because you’ve become kind. Because defecting stopped paying. Cooperating is now the selfish move. This is the sentence to tattoo on the inside of your skull: in a repeated game with a long enough future, self-interest and cooperation point the same way. It is not morality. It’s arithmetic. The rest of this lesson is that arithmetic, made explicit.

The continuation probability w

To do the arithmetic we need a number for “how much future is there?” In the real world the game rarely ends on a scheduled bell; instead, after each round, there’s some chance you’ll interact again and some chance you won’t. Call that chance w, the continuation probability: after any given round, the game continues to another round with probability w, and ends with probability 1 − w. (Economists sometimes call the same knob the discount factor — a future round is worth a fraction w of a present one, either because it might not happen or because you value it less. Same lever, either interpretation.)

ww runs from 0 to 1. At w=0w = 0 there’s certainly no next round — you’re back in the one-shot dilemma, and defection dominates. At ww near 1 the relationship stretches on almost surely — a long shadow. The magic is that somewhere in between there’s a threshold: below it the future is too flimsy to restrain you and defection pays; above it the future is heavy enough that cooperation pays. Let’s find it with real numbers.

Worked example: does cooperation pay at w = 0.8?

Use the standard payoffs from the machine below — the four numbers that define a prisoner’s dilemma. By convention they’re labelled T, R, P, S:

SymbolMeaningValue
T — TemptationYou defect, they cooperate (you exploit the sucker)5
R — RewardBoth cooperate3
P — PunishmentBoth defect1
S — SuckerYou cooperate, they defect (you get exploited)0

Now suppose you’re both playing tit-for-tat — cooperate first, then copy the opponent’s last move — so any defection will be answered next round. Compare two lives you could lead against such a partner.

Life A — cooperate forever. You collect the reward R = 3 this round, then again next round (worth 3w3w because that round only happens with probability ww), then again (3w23w^2), and so on. Summing that infinite stream gives the tidy formula for a geometric series:

Vcoop=3+3w+3w2+=31wV_{\text{coop}} = 3 + 3w + 3w^2 + \cdots = \frac{3}{1-w}

Life B — defect once, then eat the punishment. You give in to temptation this round and grab T = 5 against a partner who’s still cooperating. But your partner plays tit-for-tat, so from next round on they retaliate; the best you can then sustain is mutual defection, worth P = 1 every round forever. That future stream of 1s, starting next round, is w1w×1\frac{w}{1-w} \times 1. So:

Vdefect=5+w1wV_{\text{defect}} = 5 + \frac{w}{1-w}

Cooperation is the selfish choice exactly when Life A beats Life B — when 31w5+w1w\frac{3}{1-w} \geq 5 + \frac{w}{1-w}. Let’s just plug in w=0.8w = 0.8 and see:

StrategyValue at w = 0.8Arithmetic
Cooperate forever153/(10.8)=3/0.2=153 / (1 - 0.8) = 3 / 0.2 = 15
Defect once, then punished95+(0.8/0.2)×1=5+4=95 + (0.8 / 0.2) \times 1 = 5 + 4 = 9

Cooperating is worth 15; the one-time betrayal is worth only 9. The peach you’d steal today (a one-off gain of T − R = 2) is dwarfed by the lifetime of 3-a-round cooperation you’d be torching. A selfish player cooperates. Now shorten the future to w=0.2w = 0.2: cooperating is worth 3/0.8=3.753/0.8 = 3.75, while defecting is worth 5+(0.2/0.8)=5.255 + (0.2/0.8) = 5.25. Betrayal wins. Same players, same payoffs — only the length of the shadow changed, and it flipped the answer.

The intuitive threshold: cooperation pays when the future is valued highly enough that the ongoing losses from being punished outweigh the one-time gain from cheating. For these exact numbers the tipping point sits at w=0.5w = 0.5 — above half, cooperate; below half, defect. Fatten the temptation T or thin the reward R and the required w climbs; the greedier the short-term prize, the longer a shadow you need to resist it.

Tip:

The whole lesson in one line

Cooperation is rational whenever the future is heavy enough — when the continuation probability w clears the threshold at which tomorrow’s lost cooperation outweighs today’s one-time temptation. Lengthen the shadow and self-interest itself starts cooperating.

Play with the shadow yourself

Below is the repeated dilemma as a live machine, with those same T, R, P, S payoffs (5, 3, 1, 0). It’s loaded with tit-for-tat against a pure cheat over 14 rounds. Run the experiments in the caption and watch the shadow lengthen as you crank the round count — because more rounds is exactly what a higher w buys you.

Repeated game

The shadow of the future, made live

Pick a strategy for each side, choose how many rounds they play, then step or play the match. Watch who pulls ahead — and notice that the defector’s edge fades the longer the game runs.

14

Player A

Player B

C = cooperateD = defect
Player A · Tit-for-Tat0
Player B · Always Defect0
Total: 0/14
Experiment 1 (as loaded): Tit-for-Tat vs Always Defect over 14 rounds. The cheat steals round one, but retaliation caps its lead — and notice the defector's ADVANTAGE per round shrinks as the game runs on. Experiment 2: crank rounds up to the maximum; the longer the game (the longer the shadow), the more the steady cooperator would have out-earned the cheat. Experiment 3: set BOTH sides to Tit-for-Tat — two 'selfish' strategies cooperate every round with nobody policing them, and rack up the high score. Experiment 4: drop rounds to just 2 or 3 (a short shadow) and see how the cheat's early theft dominates. Short game rewards the cheat; long game rewards the cooperator.

The lever to feel is Experiment 4 against Experiment 2. A short game is a short shadow — a low w — and the cheat’s one-time theft is never repaid. A long game is a long shadow — a high w — and every round of stolen advantage is answered until steady cooperation pulls ahead. You are watching the threshold with your own eyes: somewhere between “2 rounds” and “many rounds,” the winner flips.

A supplier deal pays you R = 3 per year for cooperating (delivering honestly), but you could defect once — ship shoddy goods for a one-time gain of T = 5 — after which the buyer, who plays tit-for-tat, never trusts you again (P = 1 forever). Which single factor most determines whether cheating is worth it?

The evolutionary reading: why selection favours a long shadow

So far this is game theory. Now make it evolutionary, because that’s what this course is really about. Selection doesn’t run the algebra consciously — but it doesn’t need to. Strategies that reap higher payoffs leave more copies of themselves, generation after generation, and “payoff” here just means survival and offspring. So the arithmetic above isn’t merely what a clever player would do; it’s what selection builds whenever the conditions hold.

And the key condition is precisely a long shadow of the future: selection favours reciprocators when encounters repeat with the same partner. If a helpful gene keeps meeting the same individuals — a stable social group, a home reef, a lifelong pair-bond — then any help it gives can be repaid, and the reciprocity arithmetic pays for itself in offspring. This is why cooperation blooms in exactly the settings where partners recur: territorial animals with fixed neighbours, long-lived social species, tight-knit groups. A long shadow isn’t a nice-to-have; it’s a condition cooperation needs to evolve at all.

The sharp, testable prediction is what happens when you shorten the shadow — and it’s everywhere:

  • Repeat customers vs. the tourist trap. The café by the train station that will never see you again has every incentive to overcharge and underdeliver; the neighbourhood bakery living on regulars cannot afford to. Same food, opposite behaviour — a difference in w, nothing more.
  • Long supplier relationships. Firms that expect to trade for years quietly cooperate — honouring handshake deals, absorbing the odd bad break — because the relationship’s future dwarfs any single squeeze. A one-off vendor has no such restraint.
  • “Endgame” behaviour. People behave worse when the shadow collapses to nothing: the last week before someone quits a job, a dying industry cashing out, a soldier’s last patrol, a landlord who knows a tenant is leaving. When there’s no tomorrow to protect, the one-shot logic reasserts itself and defection returns — exactly as the model predicts.
Warning:

The prediction, stated as a lever

Want more cooperation from someone? Lengthen their shadow of the future. Make the relationship feel ongoing, make defections visible and remembered, raise the odds you’ll meet again. Want to know where cooperation will break? Look for a collapsing shadow — a last day, a final round, a partner you’ll never see again. The model doesn’t just describe cooperation; it tells you where to expect its absence.

The backward-induction crack — and the evolutionary escape

There’s one famous crack in all this, and honesty requires naming it. Suppose the game repeats a known, finite number of times — say exactly 100 rounds, and both players know it. Reason about the last round, 100: there’s no future after it, so it’s effectively a one-shot dilemma, and defection dominates. So you’ll both defect on 100. But then round 99 has no future worth protecting either (you already know 100 is a mutual defection no matter what), so you defect on 99 — and by the same logic on 98, 97, all the way back to round 1. Cooperation unravels from the end. This chain is called backward induction, and it predicts defection from the very first move of a known-length game. (The Game Theory Basics course works through this in full; here it’s just the caveat to our master key.)

The evolutionary escape is beautiful in its simplicity: real horizons are almost never known and fixed. No bat knows which night is its last; no trading partnership has a bell that rings on round 100; no friendship comes with a printed expiry date. When the endpoint is uncertain — when after every round there’s just some probability w of another — the backward-induction chain has no final round to start unravelling from, and cooperation stays rational. This is why evolution reaches for the continuation probability framing rather than a fixed round count: an open-ended, probabilistic horizon keeps w high and the shadow long, and that’s the world organisms actually live in.

Warning:

The pitfall that flips every prediction

The single most expensive modelling error here is mistaking a repeated game for a one-shot, or a one-shot for a repeated game. Analyse a decade-long supplier relationship as a cold one-shot dilemma and you’ll wrongly predict betrayal — you threw away the very future that sustains the cooperation. Run the error backwards — treat a genuine one-off, no-tomorrow encounter as if reputation and retaliation were on the line — and you’ll wrongly expect cooperation that has nothing holding it up. Before you predict anyone’s behaviour, ask one question first: how long is the shadow of the future? Get that wrong and every downstream prediction is wrong.

Summarise the master key.

Pick the right option for each blank, then check.

In a game the same players meet again and remember, so a defection today costs tomorrow's cooperation — a pressure called the . Its length is measured by the continuation probability : when w clears the threshold, cooperating becomes the move, so selection favours reciprocity wherever encounters with the same partner.

Match each idea to what it actually is.

Pick a term, then click its definition.

When to use it

Reach for the shadow of the future any time you’re trying to explain, predict, or produce cooperation among self-interested agents — people, firms, animals, or algorithms. First, ask the diagnostic question: is this really one-shot, or does it repeat with memory? Get that classification right before anything else, because it flips every prediction. Then, if you want cooperation, don’t preach it — lengthen the shadow: make the relationship ongoing, make actions visible and remembered, raise the odds of meeting again, and signal that you reciprocate. And if you’re bracing for betrayal, hunt for the collapsing shadow — the endgame, the last round, the partner you’ll never see again — because that’s exactly where the one-shot logic comes roaring back.

You now hold the master key. Repetition rewrites the payoffs; a long enough shadow makes cooperation the selfish choice; and selection builds reciprocators wherever partners recur. But we’ve been hand-waving about which strategy a reciprocator should actually play. Cooperate first? Punish how hard? Forgive when? Next lesson answers it with one of the most famous results in all of social science: How Tit-for-Tat Won — Axelrod’s computer tournaments, where the simplest program in the room beat every scheming rival, and the four traits that made it unbeatable.

Mark lesson as complete