The intro lesson sold you the vibe: keyboards designed to slow you down, worse formats winning, markets locking onto accidents. Fun. But a vibe is not a model, and “history matters” is the kind of phrase that feels deep at 2 a.m. and means nothing by breakfast. This lesson does the unglamorous, load-bearing work: it states exactly what path dependence claims, and — more importantly — what it does not claim. Get this distinction wrong and the rest of the course collapses into just-so storytelling. Get it right and you’ll have a scalpel where most people have a slogan.
Before we begin, feel the shape one more time — commit to an answer before you reveal it.
Before you read — take a guess
Two towns each need a public transit system. Town A studies every option and builds the one an engineer would pick from scratch. Town B builds whatever was cheapest to lay down first, and every later extension has to connect to that, so 40 years on its network is a tangle nobody would design today — but ripping it out is unthinkable. Which town's outcome is 'path dependent' in the strong sense we care about?
If your instinct rebelled at “both equally” — good. That rebellion is the entire lesson. Let’s earn it.
The precise claim: outcomes depend on the path, not just the destination
Start with an analogy. Imagine two hikers who both end up standing on the same mountain ridge at sunset. A path-independent account of how they got there would say: forget the walk, just tell me the terrain and the weather and I can predict where any sensible hiker ends up. A path-dependent account says: no — which ridge they’re on depends on the fork they took three hours ago, and had they turned the other way at that one fork they’d be on a completely different ridge tonight, with no easy way to cross over. Same hikers, same map, same weather — different early turn, permanently different endpoint.
That’s the whole idea, so let’s make it precise:
Definition — path dependence
A process is path dependent when the outcome it settles on depends on the sequence of past states it moved through — its history — and not only on its present conditions. Small events early in the sequence can have large, persistent, and effectively irreversible consequences. Change the early history and you can change the final outcome, even with every “present” fact held identical.
The technical way to say it: in a path-independent system, if you know the current conditions you can predict the outcome, and the way you arrived at those conditions is irrelevant. In a path-dependent system, the current conditions under-determine the outcome — you also have to know the road that got you here, because different roads that pass through the same point lead to different destinations.
A quick worked example. Drop a marble into a smooth bowl and it always rolls to the same single low point, no matter where you release it — path-independent, one destination, history irrelevant. Now drop it onto an egg-carton surface with many little dimples: where it settles depends entirely on where you dropped it and which way it first rolled. Two dimples might sit right next to each other, one of them deeper (better), yet the marble can get trapped in the shallower one purely because that’s the direction it happened to start. That trap is path dependence with a physical face, and — spoiler — it’s exactly the picture behind lock-in.
When to reach for it
Reach for path dependence when three things are true at once: (1) the system had more than one possible stable outcome, (2) early, small, partly-chance events helped select among them, and (3) once selected, the outcome reinforces itself so it’s costly to reverse. Miss any one and you probably have something simpler — ordinary cause and effect, or a system with one inevitable answer. Which brings us to the most abused word in this whole subject.
The trivial version vs. the real one — learn to tell them apart or nothing else works
Here is the single most important distinction in the course, and the one that separates people who understand path dependence from people who just enjoy saying it.
There are two claims that sound almost identical and are worlds apart:
| The weak (trivial) claim | The strong (real) claim | |
|---|---|---|
| What it says | ”The past influenced the present." | "A small, contingent early difference got amplified into a large, persistent, hard-to-reverse outcome that a from-scratch chooser wouldn’t have picked.” |
| Is it true? | Always — of literally everything. | Sometimes — and it’s an empirical question each time. |
| Can it be wrong? | No. It’s unfalsifiable. | Yes. If the early event was tiny but the outcome didn’t get amplified, or the outcome was actually the best one, the claim fails. |
| Does it tell you anything? | No. It’s a description of time. | Yes. It predicts lock-in, contingency, and inefficiency. |
The weak claim is true of your breakfast, the weather, the French Revolution, and the position of your left shoe. “History mattered” applied to everything is a way of saying nothing, because there is no possible world it rules out. A claim that can’t be wrong can’t be informative — that’s not pedantry, it’s the whole reason the strong version has teeth.
The strong claim is a bet with a way to lose. It says four specific things had to happen: the triggering difference was small (not a big obvious advantage), it was contingent (it could easily have gone the other way), it got amplified (some feedback took a molehill and grew a mountain), and the result is persistent and hard to reverse (it’s stuck, even if better options now exist). Knock out any one of those and the strong claim is false for that case — which is exactly what makes it worth asserting.
Let’s watch the two readings pull apart on one case.
Worked example — two readings of the same city
A tech industry clusters in one particular valley. Two explanations:
Weak reading: “It’s there because of its history — a few firms started there, talent followed, and here we are.” All true, and all useless. It rules out nothing; you could say the identical sentence about any industry in any place. It’s just a restatement of ‘events happened in order.’
Strong reading: “A couple of essentially accidental early anchors — a university, one defence contract, two founders who happened to live nearby — gave the region a tiny early lead. That lead got amplified: talent moves where the jobs are, jobs form where the talent is, capital chases both. The loop ran until the cluster became nearly impossible to relocate — and a dozen other regions could have played this role equally well had the early dice fallen differently.”
Only the strong reading is path dependence. It’s riskier — it commits to ‘small trigger,’ ‘amplification,’ and ‘other regions were viable’ — and it’s because it’s riskier that it’s worth anything. The weak reading is safe, unfalsifiable, and empty.
Now catch the trap in the wild.
Four people explain an outcome by appealing to 'history.' Which one is making the STRONG, falsifiable path-dependence claim — the one that could actually be wrong?
Contingency vs. efficiency: why the best option can lose
Here’s the mechanism that makes the strong claim interesting. Any outcome you observe got where it is by one of two very different routes:
- By efficiency — it won because it was the best available option, and it would win again in any re-run. The market (or evolution, or engineering) did its job and selected the good thing.
- By contingency — it won because it got there first and then got amplified, and in a re-run with slightly different early luck a different option could have won just as easily. Being best had little to do with it.
Path dependence is precisely the claim that contingency, not efficiency, picked the winner — and its most provocative consequence is that the best option can lose. In an ordinary market for, say, sandwiches, an inferior sandwich is punished: customers taste better and switch, so quality wins. But when returns increase with adoption — when an option gets more valuable the more it’s used, as with standards, formats, and networks — an early lead can compound faster than a quality gap can be corrected. The good-enough-but-first option pulls ahead, everyone piles on because everyone else did, and the genuinely-better latecomer never gets the critical mass it needs. Efficiency was on the loser’s side; contingency won anyway.
An analogy that makes it click: think of a party. The best party in town isn’t the one with the best music — it’s the one everyone else is going to, and which one that is often gets decided by a few early RSVPs, a rumour, who texted whom first. A better party across the street can stay empty all night, not because it’s worse, but because emptiness begets emptiness and a crowd begets a crowd. That’s increasing returns, and it’s why “best” and “winner” quietly divorce.
The load-bearing twist
Increasing returns are what let contingency beat efficiency. When every extra adopter makes an option more attractive to the next (not less, and not neutral), quality stops being self-correcting and order of arrival starts deciding the race. Remember this from the intro: it’s the same reinforcing loop your feedback-loops prerequisite taught, now running on history — and it only tips once it crosses the critical-mass threshold. Below the threshold, quality still rules; above it, the accident that got amplified rules.
”First,” “best,” and “what won” are three different things
Sloppy path-dependence talk smears three ideas into one. Pull them apart and keep them apart:
- First — which option arrived earliest (or grabbed the earliest lead). A fact about timing.
- Best — which option is intrinsically superior on the merits, ignoring who adopted it. A fact about quality.
- What won — which option the system actually settled on and locked in. A fact about the outcome.
In a world of constant or decreasing returns — ordinary goods — these tend to coincide: the best product wins and being first buys you little, because quality keeps correcting. That coincidence is exactly why the “best wins” story feels like common sense; in most of everyday economic life it’s roughly right.
But under increasing returns, all three can diverge, and every combination is possible:
| First? | Best? | Won? | What happened |
|---|---|---|---|
| Yes | Yes | Yes | The happy case — first and best, amplified into victory. No puzzle. |
| Yes | No | Yes | Path-dependent inefficiency — the worse option got there first and locked in. (QWERTY’s reputation; VHS’s legend.) |
| No | Yes | No | The tragedy — the better option arrived late, couldn’t get critical mass, and lost. |
| No | No | Yes | An accidental early lead (not even first-to-market, just first-to-tip) ran away with it. |
The row that should unsettle you is the second one: first, not best, and it won anyway. That’s the signature of path dependence, and spotting it in the wild is a genuine skill. The mistake to avoid runs the other way too — assuming every winner must secretly be best (“markets are efficient, so if it won it was best”) smuggles the conclusion into the premise and makes path dependence unfalsifiable from the other side. Whether “what won” equals “best” is a question you have to actually investigate, case by case — not a law you get to assume.
Under INCREASING returns to adoption, which statement is true?
The small-events intuition: how a fair coin locks onto an unfair outcome
Now the intuition pump that makes contingency feel inevitable rather than mystical. Picture an urn with one black ball and one white ball. You draw one at random, look at its colour, put it back — and add one more ball of that same colour. Repeat forever. This is the Pólya urn, and it is path dependence distilled to arithmetic.
Notice what “add one of the same colour” does: it’s increasing returns in a jar. Whatever’s ahead gets more likely to be drawn next, which makes it get further ahead. Run it and the fraction of black balls wanders around at first — the early draws swing it wildly because a couple of balls is a big share of a small pot — and then, as the pot grows, it settles down and freezes onto some ratio, say 71% black. Here’s the punchline in three parts, and it’s the whole model in a toy:
- It locks in. The ratio stops moving; the outcome is stable and self-reinforcing, just like a market that’s tipped.
- Where it locks is contingent. Re-run from the identical starting urn and you’ll get a different final ratio — 71% one time, 34% the next, 88% the time after. Same rules, same start, different history, different frozen outcome. Nothing was “best”; the early draws just happened to break one way.
- It’s unpredictable before, obvious after. You genuinely cannot forecast the resting ratio in advance — but once it’s settled it looks stable and inevitable, and hindsight will happily invent a reason it “had” to be 71%. (Keep that hindsight temptation in mind; lesson 6 calls it the just-so-story trap.)
This is exactly the machine behind the lock-in explorer you played with in the intro. “New run” is a fresh spin of the urn; the early draws are the random early lead; increasing-returns strength is how many same-colour balls you add each draw. Crank that up and the market freezes hard and early — onto whichever standard the early draws favoured, better or worse. Turn it to zero and you’re back to a fair coin every time: quality-driven, never locked.
Sensitive dependence on initial conditions — the tame, accurate version
You’ve met the cousin of this idea dressed up as “the butterfly effect” — a butterfly flaps in Brazil and, weeks later, a storm forms in Texas. The useful, un-exaggerated kernel is: in some systems, tiny differences in starting conditions get amplified into large differences in outcome. Path dependence is a mild, well-behaved case of that — the amplifier is the reinforcing loop of increasing returns, and the “tiny difference” is which option grabbed the early lead.
One honest caveat, because overclaiming here is a classic sin. Path dependence is not chaos, and it’s not “everything is random and unknowable.” Three sober differences: (1) The amplification is structured — it runs through a specific, nameable feedback mechanism, not through generalised turbulence. (2) It’s usually a one-way amplification of early events specifically: once the system locks, later shocks get damped, not amplified — the exact opposite of chaos, where sensitivity never switches off. (3) The set of possible endpoints is often small and knowable (standard A or standard B); what’s genuinely unpredictable is only which of them, and only before the early events resolve. So keep the tame reading: early contingency, amplified through a named loop, into a persistent lock. Reach for “butterfly effect” as a memory hook, not as a licence to declare everything a coin flip.
Match each term to the definition that actually teaches it. These are the load-bearing words for the whole course — get them crisp now.
Putting it together
Let’s consolidate before the recap — a short quiz that mixes the type of question, because if you can only recognise the idea in one costume you haven’t learned it.
Check yourself: what does path dependence actually claim?
A colleague says: 'Of course our software architecture is the way it is — it's the product of all our past decisions. That's path dependence.' What's the most precise critique?
Check your answer to continue.
Key takeaways
- The claim, precisely. A process is path dependent when its outcome depends on the sequence of past states — its history — not just present conditions. Small early events can have large, permanent effects; change the early history and you can change the endpoint.
- Trivial vs. real is the whole game. “The past influenced the present” is true of everything and says nothing. The real, falsifiable claim: a small, contingent early difference got amplified into a large, persistent, hard-to-reverse outcome a from-scratch chooser wouldn’t have picked. If it can’t be wrong, it isn’t the real claim.
- Contingency vs. efficiency. An outcome can win because it’s best (efficiency) or because it got there first and got amplified (contingency). Path dependence says contingency picked the winner — so the best option can lose.
- Three different words. First (timing), best (quality), and what won (outcome) coincide under constant/decreasing returns but can all diverge under increasing returns. “First, not best, and won anyway” is the signature.
- The urn intuition. A Pólya urn (add a ball of the colour you drew) locks onto a stable ratio that’s different every run — contingent, unpredictable before, obvious-in-hindsight after. That’s the lock-in explorer’s engine.
- Tame, not chaotic. Sensitive dependence on early conditions runs through a named feedback loop; once locked, later shocks are damped, not amplified. Use “butterfly effect” as a hook, never as a licence to call everything random.
Where this goes next
You now have the scalpel: you can state the strong claim, separate it from the empty one, and tell “first” from “best” from “won.” Time to put it against the most famous case in the entire field — and then watch it get attacked.
Lesson 2, The QWERTY Story — and the Honest Debate, tells the canonical keyboard example straight, and then stages the academic brawl over it (Paul David’s classic account versus Liebowitz and Margolis’s revisionist takedown). Because a model you only ever see confirming itself is a superstition, not a tool. You’ll learn precisely where QWERTY is strong evidence for path dependence — and where the story is thinner than it’s usually told. Bring the scalpel; you’ll need it on both sides of the fight.