You have spent four lessons learning to distrust a comforting story. A behaviour is not stable because it looks clever, or because it “helps the species,” or because a documentary narrator said so in a hushed voice. It is stable only if it survives the invasion test: a strategy is an ESS when, for every mutant , either , or with .
This final teaching lesson turns that lens on the most emotionally loaded question in all of evolutionary game theory — can cooperation be stable? — and then, having earned the right, catalogues honestly everywhere the model quietly lies to you.
Before you read — take a guess
Before we start: in the one-shot prisoner's dilemma, cooperation is not an ESS because a defector mutant always out-earns a cooperator against a cooperative population. What single feature of the game do you think could flip that — making cooperation defensible?
Repetition changes the game: the iterated prisoner’s dilemma
In the one-shot prisoner’s dilemma, defection dominates. Whatever your partner does, you earn more by defecting, so a population of cooperators is trivially invaded — cooperation is not an ESS, full stop.
Now play it repeatedly. After each round, the same two players meet again with probability — the continuation probability, or as biologists love to call it, the shadow of the future. If is high, “the same partner again” is likely, and a strategy can remember: it can reward yesterday’s cooperation and punish yesterday’s defection.
The famous conditional strategy is Tit-for-Tat (TFT): cooperate on the first move, then copy whatever your opponent did last round. Nice (never defects first), retaliatory (punishes defection immediately), and forgiving (returns to cooperation the moment the opponent does).
Think of it as a thermostat for trust: it starts warm, cools instantly when chilled, and reheats the moment warmth returns. The question that made Robert Axelrod famous is whether that thermostat is evolutionarily stable — not just a good idea, but uninvadable.
The payoffs we'll reuse
We use the standard dilemma per round: mutual cooperation pays 3 each, mutual defection 1 each, and the tempted defector gets 5 while the exploited cooperator (the “sucker”) gets 0. So the ranking is Temptation (5) > Reward (3) > Punishment (1) > Sucker (0). Over a repeated game these per-round payoffs accumulate, weighted by how long the relationship lasts.
What does a high continuation probability w actually buy a conditional strategy like Tit-for-Tat?
Tit-for-Tat is only almost an ESS
Here is the result that separates people who have actually done the algebra from people who have read a pop-science summary. When the shadow of the future is long enough, TFT resists Always Defect (AllD) — but it is not a strict ESS. It is only neutrally stable, and that “only” matters enormously.
Step 1 — TFT beats the AllD invader. Put one AllD mutant in a world of TFT players. On round one AllD defects against a cooperating TFT: it scores 5, the sucker scores 0. Sweet — for exactly one round. From round two on, TFT retaliates and both defect forever, scoring 1 each. Meanwhile two TFT players cooperate every round, scoring 3 each. Let’s add it up over, say, a long relationship of many rounds:
| Matchup | Round 1 | Every later round | Long-run per-round average |
|---|---|---|---|
| TFT vs TFT | 3 | 3, 3, 3, … | ≈ 3 |
| AllD vs TFT | 5 | 1, 1, 1, … | ≈ 1 (the 5 washes out) |
So while . Because , TFT is a strict best reply to itself against AllD — AllD cannot invade. The one-round theft is drowned by a lifetime of punishment. So far, so heroic.
Step 2 — the quiet saboteur: Always Cooperate (AllC). Now try a nice mutant, AllC, which cooperates unconditionally. Against a TFT population, AllC cooperates every round and TFT — never provoked — also cooperates every round. They both earn 3 per round, forever. That means:
A tie. TFT’s first line of defence () does not hold against AllC. So we go to the tie-breaker: is ? Against an AllC opponent, TFT cooperates the whole way (AllC never defects, so TFT never retaliates) and earns 3 per round; two AllC players also cooperate the whole way and earn 3 per round. Tie again: . The tie-breaker fails too.
Both ESS conditions fail against AllC. TFT is therefore not a strict ESS — AllC is selectively neutral against TFT. In a finite world, neutral mutants drift: with no fitness penalty to stop them, AllC copies accumulate by chance.
Step 3 — the softened population gets mugged. And here is the sting. Once AllC has drifted in and the population is a soft mix of TFT and AllC, an AllD mutant reappears — and now it thrives, because it can exploit all those unconditional cooperators for free (AllC never punishes) while only the dwindling TFT minority retaliates. Cooperation, undefended by memory, gets mugged.
Neutral is not the same as safe
The lesson people miss: TFT is not felled by a better strategy. It is undermined by an equally good one. AllC never out-earns TFT — it just ties, drifts in on that tie, and rots away the retaliation that kept AllD out. “Uninvadable by anything that does strictly worse” is a weaker shield than “uninvadable, period.” That gap between strict ESS and merely neutrally stable is the whole story here.
Precisely why does Tit-for-Tat fail the ESS test — even though Always Defect cannot invade it?
Always Defect is an ESS too — so the world is bistable
Cooperation’s defenders often stop at “reciprocity can be stable” and take a victory lap. Don’t. Run the invasion test on the other equilibrium and something sobering falls out.
Consider a population of pure AllD. Everyone defects every round; everyone earns the punishment payoff of 1 per round. Now drop in any nice mutant — TFT or AllC. On round one it cooperates (that is what “nice” means), the AllD natives defect, and the mutant eats the sucker’s payoff of 0. From then on:
- AllC keeps cooperating into a wall of defectors — 0 every round. Catastrophic.
- TFT cooperates once (0), gets burned, then retaliates and defects for the rest — 1 per round thereafter, exactly matching the natives, but it never recovers that lost first round.
Either way, (which is at most 1, and strictly less once the round-one sucker loss is counted). No nice mutant can get a toehold. AllD is a strict ESS.
So the repeated prisoner’s dilemma is bistable: it has two evolutionarily stable equilibria. “Everyone reciprocates” (a TFT-like cooperative regime) and “everyone defects” are both uninvadable. Which one a population ends up in is not decided by which is better — mutual cooperation pays 3, mutual defection pays a miserable 1 — but by history, initial conditions, and the basin of attraction it happened to start in.
Two valleys, one ball
Picture the fitness landscape as a terrain with two valleys — a deep, rich one (cooperation) and a shallow, poor one (defection) — separated by a ridge. A ball (the population) rolls to whichever valley it started nearest. It does not teleport to the better valley just because that valley is better. This is path-dependence in its purest form: the outcome is locked in by where you began, not by where it would be best to end up.
This is why cooperation is hard to restart once lost (a defecting population sits in its own stable valley and punishes the first nice mutant) and yet self-defending once established (a reciprocating population repels defectors). Trust is expensive to build and cheap to burn — and now you can see why in the arithmetic, not just the aphorism.
Pick the right option for each blank, then check.
Because both all-reciprocate and all-defect pass the invasion test, the repeated dilemma is : it has more than one ESS. Which equilibrium a population reaches is set by its , not by which outcome pays more — a textbook case of .
Sort each game by how many evolutionarily stable strategies it has.
Place each item in the right group.
- A three-morph frequency game where each type beats one and loses to another
- The repeated prisoner's dilemma — all-reciprocate and all-defect are both stable
- A coordination game where both "everyone drives left" and "everyone drives right" resist invasion
- The sex-ratio game — one ESS at a 50:50 split
- Hawk–Dove with V < C — one mixed ESS at p* = V/C
- Rock–paper–scissors side-blotched lizards — orange/blue/yellow chase each other
Replicator dynamics: the engine that does the moving
We keep saying selection “carries” a population toward an ESS. Time to name the engine. The replicator equation is the simplest honest model of that motion, and its one-line idea is the whole thing:
A strategy’s share grows in proportion to how far above the population average its fitness is. Fitter-than-average expands; below-average shrinks; exactly-average holds.
For two strategies with shares and and fitnesses and :
Read it left to right and it almost narrates itself. The term is just “you can only change if both types are actually present” (nothing to select between if everyone is already A or already B). The term is the steering wheel: positive where A beats B, so A’s share rises; negative where A loses, so it falls; zero where they tie, which is a rest point. It’s compound interest for behaviour — winners get proportionally more of next generation’s “capital.”
Three connections tie this back to everything you’ve built:
- You already used it. The “let selection run” button in the Hawk–Dove lab from Lesson 2 is this equation. Whatever fraction of hawks you started with, it drove the population to — because at any other mix, one strategy is above average and grows until the gap closes.
- An ESS is an attractor. In dynamical language, an ESS is an asymptotically stable rest point of the replicator dynamics: a resting mix that, if you nudge it, pulls itself back. That is the precise meaning of “stable” — not a metaphor, a mathematical property.
- But not every system rests. Rock–paper–scissors — the side-blotched lizards — has no rest point that attracts; the replicator dynamics send it orbiting around the center forever, orange chasing blue chasing yellow chasing orange. No attractor, no ESS. Cycles are a legitimate long-run behaviour, not a bug.
The one sentence to keep
Replicator dynamics = “fitter-than-average grows, below-average shrinks.” An ESS is where that process comes to rest and stays put when disturbed. Some games rest at one point, some at several, and some never rest at all.
Match each dynamical behaviour to what it means for the existence of an ESS.
Pick a term, then click its definition.
Where the model lies: the honest catalogue
You now have a powerful tool. Powerful tools earn respect precisely because their owners know where they break. Here is the unflattering, load-bearing list of what ESS does not tell you.
1. Stability is not optimality
This is the big one, and it’s the one that trips up smart people. An ESS is what selection settles into — not what would be best for the group. The Hawk–Dove game proves it with numbers.
Recall from Lesson 2, with value and cost , the mixed ESS is hawks. At that equilibrium, every individual — hawk or dove — earns the same expected payoff, and that payoff works out to :
Now imagine a population of all Doves. They share resources peacefully, and each earns . That is more than double the ESS payoff of 1.2. Everyone would be richer in the all-Dove world.
| World | Composition | Payoff per individual |
|---|---|---|
| All-Dove utopia | 0% hawks | 3.0 |
| The ESS | 60% hawks | 1.2 |
So why doesn’t the population just go to the lovely all-Dove world? Because all-Dove is invadable: drop one Hawk into a room of Doves and it wins every contest unopposed, earning the full while doves split 3. Hawks flood in until . The peaceful optimum is not stable, so selection can’t hold the population there. Selection optimises the individual gene’s spread, not the group’s welfare — and those are different goals. ESS predicts the wasteful 1.2, and reality obliges.
A population sits at the Hawk–Dove ESS earning 1.2 each, when everyone would earn 3.0 in an all-Dove world. Why can't natural selection just move the population to the better outcome?
2. Not every game has an ESS — and some have several
“Find the ESS” quietly assumes there is one to find. There need not be. Rock–paper–scissors (the lizards) has none — the dynamics cycle. Some games have only a mixed ESS (Hawk–Dove). Some have multiple ESSs (the bistable dilemma; any coordination game). The honest answer to “what is the ESS here?” can be zero, one, or several — and you don’t know which until you’ve actually run the test.
3. The classic model assumes an infinite, well-mixed population of unrelated strangers
The tidy algebra of rides on three assumptions that real populations routinely break:
- Infinite population. The math assumes mutants are vanishingly rare and averages are exact. Real populations are finite, so random drift can push a strategy out even when selection favours it (exactly the drift that let AllC creep into TFT), or fix a worse one by luck.
- Unrelated players. ESS payoffs assume you meet strangers. Add kin structure and the effective payoffs change: helping a relative propagates shared genes, so cooperation can pay when Hamilton’s rule holds (relatedness times benefit to the recipient exceeds cost to the actor) — even where the stranger-only ESS says defect.
- Well-mixed encounters. ESS assumes everyone is equally likely to meet everyone. Real interactions have spatial or network structure — who meets whom is lumpy. On a lattice, cooperators can cluster and shelter each other, surviving where a well-mixed model says they’d be exploited to extinction.
Each of these is not a footnote; each rewrites the payoff matrix.
4. It’s a static snapshot of a moving world
ESS is an equilibrium concept. It asks “what mix, once reached, stays put?” — and by construction ignores that the game itself is moving. and drift as environments change. Opponents coevolve, so a strategy chasing today’s best reply aims at a moving target — a moving ESS. There are time lags, seasons, invasions. Real evolution is a chase, not a tidy resting point; ESS photographs one frame and pretends the film has stopped.
5. Loose narration is the biggest practical trap
The most common abuse of this whole field is not a math error — it’s storytelling. “That behaviour is evolutionarily stable” is trivially easy to assert and genuinely hard to earn. Just-so stories breed like rabbits: any trait can be dressed up as adaptive after the fact. The discipline that separates science from narration is mechanical — name the strategy set, write the payoffs , and actually run the invasion test. If you haven’t done those three things, you’re telling a campfire story, not making a claim.
6. The naturalistic fallacy
Finally, the ethical guardrail carried over from the cooperation course: an ESS describes what evolves, not what is good, fair, or right. Stable fighting (Hawk–Dove), stable deception, and stable defection (the all-AllD world) are all ESSs. Being evolutionarily stable confers exactly zero moral endorsement. “It’s natural” and “it’s stable” are descriptions; “it’s good” is a separate claim you have to argue on its own terms.
Practitioner's checklist — before you claim an ESS
To earn the phrase “this is an ESS,” specify all of it:
- (a) the strategy set — what alternatives are on the table;
- (b) the payoff function E(A,B) — the numbers, not a vibe;
- (c) best-reply-to-itself — check E(I,I) >= E(J,I) for every mutant J;
- (d) the tie-breaker where E(I,I) = E(J,I) — check E(I,J) > E(J,J);
- (e) confirm it’s a rare-mutant-proof attractor of the dynamics.
And keep in mind the count: zero, one, or many ESSs may exist. If you skipped (a)–(b), you never had a claim — just a story.
Which of the following are genuine limitations of the classic ESS model? (Select all that apply.)
Match each limitation of the ESS model to a concrete example of it biting.
Pick a term, then click its definition.
Spaced recall: don’t lose the core machinery
You’ve added a lot today. Let’s make sure the foundation from earlier lessons is still load-bearing.
Quick self-test — cover the answer. A colleague says, “Cooperation obviously evolved because it’s better for everyone — look how much richer a cooperative world is.” Using only what you know about the invasion test, name the two things that claim gets wrong.
Answer. (1) “Better for everyone” is a group-welfare argument, but the invasion test is about individual advantage against the current population — a strategy is stable only if no mutant out-earns it, regardless of group welfare (stability ≠ optimality; see the all-Dove world paying 3 yet being invadable). (2) The claim skips the mechanism entirely: it never names the strategy set or checks — it’s a just-so story. Cooperation can be stable, but only via repetition/kinship/structure that actually passes the test, and even then the defect equilibrium is stable too (bistability).
Back to basics: Hawk–Dove with V = 6, C = 10. What is the mixed ESS fraction of Hawks, and what does the invasion test say about a population sitting there?
Recap
You’ve now seen the model at its most powerful — pinning down when cooperation can and can’t be stable — and at its most honest, cataloguing where it stops being trustworthy. The two halves are the same skill: run the test, then respect its boundaries.
Big picture
Cooperation, waste, and the limits of ESS
- ESS: cooperation & limits
- Iterated prisoner's dilemma
- Shadow of the future w makes memory pay
- TFT: nice, retaliatory, forgiving
- TFT is only almost an ESS
- Beats AllD: E(TFT,TFT)≈3 > E(AllD,TFT)≈1
- Ties AllC on both clauses → neutral, drifts in
- Softened population then invaded by AllD
- AllD is a strict ESS → bistable world
- Nice mutant eats sucker payoff, never recovers
- Two basins; history/path-dependence decides
- Replicator dynamics = the engine
- dp/dt = p(1−p)(W_A − W_B): fitter-than-average grows
- ESS = attractor (asymptotically stable rest point)
- RPS lizards orbit → no attractor, no ESS
- Where the model lies
- Stability ≠ optimality (Hawk–Dove 1.2 vs 3)
- Zero, one, or many ESSs
- Infinite / well-mixed / unrelated assumptions break
- Static snapshot of a coevolving, drifting game
- Loose narration + naturalistic fallacy
- Iterated prisoner's dilemma
Why is Tit-for-Tat NOT a strict ESS, even when the shadow of the future is high?
Check your answer to continue.
That’s the arc: from “is uninvadable?” to a full working knowledge of when cooperation can be stable, what engine drives populations there, and every place the model’s tidy story frays. Use it the way a good mechanic uses a torque wrench — precisely, and knowing exactly what it can’t measure.