You’ve spent three lessons watching metrics rot: cobra farms, clickbait, teachers drilling exam papers, a factory stamping out seventy-tonne nails. It’s tempting to file Goodhart’s Law as its own private curiosity — a quirky rule about numbers going bad.
It isn’t. Goodhart’s Law is not a fundamental law at all. It’s an emergent symptom — what you get when three models you already carry in your latticework fire at the same target. Pull them apart and Goodhart stops being a mystery and becomes a machine with three named parts: incentives, the principal–agent problem, and feedback loops. This is the lesson where the whole thing clicks into the rest of your toolkit.
Before you read — take a guess
Before we dissect it — take a guess. Goodhart's Law is best understood as…
Incentives are the engine
Analogy. A thermometer clipped to the wall is a neutral instrument — nobody bribes a thermometer. But bolt a bonus to it (“everyone gets $500 if the reading hits 22°C”) and watch what happens: someone holds a lighter under it. The instrument didn’t change. The reason to touch it did.
Mechanism. The incentives model says people optimise the reward they actually face, not the goal you meant. A metric with no stake attached is a passive read-out — a thermometer. Attach a reward and the metric flips from a thermometer (something you read) into a lever (something you pull). Goodhart cannot even begin without this step: with no reward riding on the proxy, nobody bothers to game it. The corruption needs fuel, and the incentive is the fuel.
Worked example. A support team’s dashboard tracks “average ticket resolution time.” For a year nobody’s pay depends on it, and it’s a fair read on how fast customers get helped. Then management announces bonuses for the fastest resolvers. Within a month, agents are closing tickets the instant they reply (before the customer confirms anything is fixed), splitting one problem into five “resolved” tickets, and pushing hard cases to a colleague’s queue. Resolution time plummets; actual help gets worse. The clock was always ticking — but only once money rode on it did anyone start racing it.
Thermometer → lever
The single sharpest test for whether Goodhart is about to bite: has this measurement become a lever someone is rewarded for pulling? A number you merely observe stays honest far longer than a number someone is paid to move.
So the first parent, stated cleanly:
Goodhart = incentives pointed at a proxy instead of the goal.
The KPI stops being a neutral thermometer the instant a reward is bolted to it.
The principal–agent gap is the opening
Analogy. You hire a mechanic (the agent) to fix your car; you (the principal) can’t see under the hood the way they can. They know which repairs are real and which are padding; you only know what the invoice says. That gap in knowledge is exactly the room they have to charge you for a part you never needed.
Mechanism. The principal–agent problem describes any setup where a principal delegates to an agent whose actions and information the principal can’t fully observe — hidden action (you can’t watch what they do) and hidden information (they know things you don’t). Whoever is measured almost always understands how the number is produced better than whoever is measuring. That asymmetry is the opening Goodhart crawls through: it lets the agent satisfy the letter of the metric while the spirit quietly rots, and — crucially — do it undetected. This is textbook moral hazard: the agent games the gap precisely because the gaming is hidden.
Worked examples.
- A fund manager is judged against a benchmark index. They know exactly how the number is built; you don’t. So they “hug the index” — quietly buying almost the same holdings to avoid ever underperforming it — collecting active-management fees for near-passive work. The tracking number looks fine; the value you were paying for isn’t there.
- A surgeon is ranked on a public mortality league table. Nobody can observe the judgment call in each case, only the aggregate outcome. So the rational move is to refuse the riskiest patients — the ones most in need of surgery. Mortality stats improve; care for the sickest collapses. The metric can’t see the patients who were turned away.
If the measurer could see everything, Goodhart would starve
Goodhart needs the fog. A principal who could perfectly observe the agent’s actions and knowledge would just reward the real goal directly — no proxy, no gap, no gaming. Every real case of Goodhart lives inside an information asymmetry. Shrink the asymmetry and you shrink the law’s grip.
A hospital's surgeons quietly start declining high-risk patients once a public mortality league table launches. Which parent model is doing the heavy lifting here?
Feedback loops make it self-reinforcing and self-erasing
Analogy. Point a microphone at the speaker it’s feeding and you get that ear-splitting screech: the output loops back into the input and amplifies itself until the signal is pure noise. Goodhart has the same architecture. The proxy’s output loops back into the decisions that produce it — and cranks itself until the signal is gone.
Mechanism. The feedback loop model describes a system whose output is wired back into its own input. Wire a stake into a proxy and you close a reinforcing loop: the proxy feeds back into decisions → effort floods toward the proxy → that flood corrupts the very signal the proxy used to carry → so the metric’s validity decays over time. It’s a loop that literally eats its own information content: the harder it’s optimised, the less it means. The contrast matters — a well-designed proxy has no such loop at all until you wire a stake into it. The measurement sits there, honest and stable, until the incentive closes the circuit and the screech begins.
Worked example. A social platform optimises for an engagement metric (time-on-site, clicks). Early on, engagement genuinely tracks “people find this valuable.” Then the loop closes: engagement rewards → creators learn clickbait wins → the feed fills with clickbait → clickbait gets rewarded even harder → real quality drains out while the engagement number keeps climbing. The metric looks better every quarter precisely as the thing it was supposed to measure — genuine value — gets worse. The loop didn’t just detach the signal; it inverted it.
Reinforcing, not balancing
Not every feedback loop is destructive — balancing loops (like a thermostat) are self-correcting. Goodhart rides a reinforcing one: each turn makes the next turn stronger, so the corruption compounds rather than settles. That’s why a gamed metric doesn’t drift a little and stop — it runs away.
Put the three together
Now assemble the machine. Goodhart’s Law is what emerges when all three parents fire on one measured proxy under stakes:
Read left to right, that’s the whole life cycle of a metric going bad:
- Incentives attach a reward to the proxy — turning a thermometer into a lever, giving anyone a reason to move it.
- The principal–agent gap gives them the room — hidden action and information mean they can move the letter while the spirit rots, undetected.
- The feedback loop supplies the engine of decay — optimisation floods back in and erodes the metric’s own validity, faster the harder you push.
None of the three alone is Goodhart. A rewarded metric with a perfectly observable agent and no loop stays honest. A measured-but-unrewarded proxy just sits there. Goodhart is the emergent failure that appears only when the three overlap on a single number that has a stake on it.
Match each parent model / term to the exact role it plays in producing Goodhart's Law.
Pick a term, then click its definition.
Why naming the parents is powerful, not pedantic
This isn’t taxonomy for its own sake. Because Goodhart is made of three models, each fix in the next lesson attacks one parent: change what you reward (incentives), close the observation gap (principal–agent), or break/slow the loop (feedback). If you only know “Goodhart” as one undifferentiated blob, you can’t aim your intervention. Decompose it and every fix has an address.
Diagnosis: which parent is loudest?
Here’s the practical payoff. When you’re staring at a metric that’s clearly being gamed, don’t just sigh “Goodhart’s Law.” Diagnose which parent is doing the most damage — because that’s the one to attack.
Ask three questions:
- Incentives — is the reward too concentrated? Does a large, sharp stake (a bonus, a ranking, a promotion) ride on this one number? The more of someone’s outcome depends on a single proxy, the harder they’ll pull the lever.
- Principal–agent — is the gap unobservable? Can the measured party act on knowledge the measurer can’t see? The foggier the setup, the more room to game the letter undetected.
- Feedback — is the loop tight and fast? Does moving the metric immediately change the decisions that produce it, with a short delay? A fast reinforcing loop runs away before anyone notices.
Whichever answer is loudest points you at the fix that will actually move the needle — which is exactly where the next lesson goes.
| Parent model | What it contributes to Goodhart | Where to intervene |
|---|---|---|
| Incentives | The motive — a reward turns the proxy from a thermometer you read into a lever people pull. | Spread the stake across several measures, soften the reward, or reward the true goal (or judgment) rather than the proxy. |
| Principal–agent problem | The opportunity — hidden action/information lets the agent satisfy the letter while the spirit rots, undetected. | Shrink the asymmetry: better observation, audits, aligned interests, or rotating/randomised checks the agent can’t anticipate. |
| Feedback loop | The engine of decay — optimisation floods back and erodes the metric’s own validity, faster the harder you push. | Break or slow the loop: retire/rotate metrics before they saturate, add lag, or hold some measures back as unrewarded honest reads. |
You run a call centre. A single, large monthly bonus rides entirely on 'calls handled per hour', and agents can mark a call 'resolved' with zero verification. Which parent is loudest — and so where should you aim first?
Recap
Big picture
Goodhart, decomposed into its three parents
- Goodhart = incentives + principal–agent + feedback
- INCENTIVES — the engine
- Reward on the proxy → thermometer becomes a lever
- No stake, no gaming
- PRINCIPAL–AGENT — the opening
- Agent knows more than the measurer
- Games the letter, undetected (moral hazard)
- FEEDBACK — self-reinforcing decay
- Optimisation floods back and corrupts the signal
- Metric validity erodes over time
- DIAGNOSIS — which parent is loudest?
- Reward too concentrated? Gap unobservable? Loop too fast?
- The loudest parent is where you intervene
- INCENTIVES — the engine
Goodhart’s Law was never a standalone rule to memorise — it’s three familiar models overlapping on one measured number under stakes. That’s what “special case” means: not rare, but derived. And the decomposition isn’t just satisfying — it’s operational, because it hands you three separate places to intervene. Next lesson: the fixes that actually work — each one aimed squarely at one of the three parents — and how quoting “Goodhart!” can itself curdle into a lazy excuse.