Last lesson we cut confirmation bias into three moving parts: biased search, biased interpretation, and biased memory. This lesson zooms all the way in on the first one — search — and it does it with a single, almost insultingly simple puzzle that has been quietly humiliating clever people since 1960. You met it in passing in the introduction. Now you’re going to run it on yourself, watch your own mind reach for yes, and feel exactly why a confident pile of confirmations can be worth precisely nothing.
The puzzle is Peter Wason’s 2-4-6 task, named for the London psychologist who built it — the same man who, that year, coined the phrase confirmation bias itself. Wason wasn’t trying to catch people being stupid. He was trying to see how people search for evidence when they want to test an idea, and what he found is the whole subject of this course in miniature: people don’t test their ideas. They pet them.
Read this lesson with your hands
There’s a live experiment halfway down this page. If you scroll past it nodding along, you’ll learn nothing — which would be a deliciously on-brand way to fail a lesson about confirmation bias. When you reach the island, actually propose triples before you guess. The sting of the result is the lesson.
Positive testing vs negative testing
Suppose you have a hunch — call it a hypothesis — and you want to find out if it’s right. You’re going to look for evidence. But there are two completely different kinds of evidence-hunt available to you, and which one you reach for decides whether you ever learn anything.
A positive test is a case you run because you expect it to confirm your hypothesis. You believe the rule is “even numbers,” so you try 8, 10, 12 — a triple you fully expect to come back “yes.” A negative test is a case you run because you expect it to disconfirm your hypothesis — a case you predict will come back “no” if your rule is right. You believe the rule is “even numbers,” so you deliberately try 1, 2, 3, expecting it to fail, precisely to see whether it actually does.
The names are about your expectation going in, not the result that comes out. A positive test is one you’re betting will succeed; a negative test is one you’re betting will fail. Most people, faced with a hypothesis they like, run positive test after positive test after positive test — and never once run the negative one. That lopsided habit has a name: positive test strategy, the deep tendency to probe a hypothesis by looking for cases that should fit it rather than cases that should break it.
The 2-4-6 puzzle, set up
Someone has a secret rule that some sequences of three numbers obey. They tell you only this: the triple 2, 4, 6 obeys it. Your job is to discover the rule. You may propose as many triples of your own as you like, and for each one you’ll be told fits or does not fit. When you’re sure, you announce the rule. Pause and ask yourself, right now: what triples would you try first?
Here’s the trap in slow motion. You see 2, 4, 6, and your mind instantly offers a tidy hypothesis: even numbers going up by two. Feels obvious. So you test it the natural way — positively. You try 8, 10, 12: fits. You try 20, 22, 24: fits. You try 100, 102, 104: fits. Three for three. The little click of see, I’m right fires each time, and after a handful of these you announce, with real confidence, “the rule is even numbers ascending by two.”
And you’re wrong. The actual rule is far broader, and not one of your three confident tests could ever have revealed that — because each was chosen to pass. Every “fits” you collected was consistent with your guess and with a hundred other rules you never considered. You didn’t test your theory. You took it out for a victory lap.
Before you read — take a guess
You think the hidden rule behind 2, 4, 6 is 'even numbers going up by two.' Which triple is a NEGATIVE test — the kind that could actually expose your guess as wrong?
Why positive tests pile up false confidence. Imagine the real rule were genuinely “even numbers ascending by two.” Then positive tests would correctly confirm it. The problem is that positive tests give you the same reassuring “fits” whether or not your guess is the real rule — as long as your guess is a special case of something broader. “Even numbers up by two” is a tiny island inside the giant continent of “any three increasing numbers.” Every triple on your little island also sits on the continent, so it passes either way. You can stand on your island collecting “fits” forever and never discover you’re on a continent. Confirmations don’t distinguish between the narrow rule you guessed and the broad rule that’s actually true. Only a triple that steps off the island — one you expect to fail — can tell the two apart.
Run it yourself
Enough theory. Time to be the experimental subject. Below is the live task: you’re told 2, 4, 6 fits a hidden rule, and you discover the rule by proposing your own triples and watching each come back fits or does not fit.
A few rules of engagement, because how you play is the lesson:
- Don’t peek at the answer choices first. Form your own hunch from
2, 4, 6alone, then test it. - Test as many triples as you want before you commit to a guess — there’s no penalty for curiosity.
- And the one move almost nobody makes: at some point, try a triple you genuinely expect to come back “does not fit.” If every test you run comes back “fits,” ask yourself why you never tried to break your own theory. Then go break it.
Propose your triples now. Notice what your hand reaches for.
Discover the rule
Wason's 2-4-6 task — find the hidden rule
The triple 2, 4, 6 fits a rule I have in mind. Propose your own triples to discover the rule — each comes back 'fits' or 'does not fit'. Test as many as you like, then guess. Bonus challenge: try at least one triple you expect to FAIL.
2, 4, 6fits ✓
What do you think the rule is?
The debrief. Two things probably happened. First, you almost certainly arrived at a tidy guess fast — “even, plus two” or “constant step” — because 2, 4, 6 suggests one so strongly. Second, if you’re like the majority of Wason’s subjects, most or all of your tests came back “fits,” and you felt more and more confident with each one. That confidence was a trap. The rule was simply any three numbers in strictly increasing order — 1, 2, 3 fits, 5, 50, 500 fits, even 7, 8, 9 fits. Your narrow guess was a tiny special case of it, which is exactly why every positive test passed and told you nothing.
Now the score that matters: did you ever run a triple you expected to fail? If yes, you did the rare and powerful thing — you ran a negative test, and a single “does not fit” (try 3, 2, 1) instantly teaches you more than a hundred “fits” ever could. In Wason’s original studies, only about one in five people solved it on their first announced rule. The other four-fifths didn’t lack intelligence. They lacked the instinct to look for the answer they didn’t want.
Here’s a subtlety that catches careful people. Some subjects, after guessing “even, +2,” would test 8, 10, 12 and 14, 16, 18 — and then, feeling rigorous, test 3, 5, 7 (“odd, +2”) and 10, 20, 30 (“even, +10”). When those also came back “fits,” they didn’t conclude “my rule is too narrow.” They quietly broadened their guess to “ascending by a constant step” and kept testing that positively. They were testing variation after variation, all positive, never once proposing a triple they expected to fail. The fix isn’t to test more; it’s to test differently — to deliberately propose a triple that your current best guess says should come back “does not fit,” like 2, 4, 5 or 1, 2, 100. Effort isn’t the missing ingredient. Direction is.
Match each triple to what it actually does for someone whose current guess is 'even numbers ascending by two.'
Pick a term, then click its definition.
Why “yes” teaches you nothing
The 2-4-6 task isn’t just a party trick about numbers. It’s a hands-on demonstration of one of the deepest ideas in the philosophy of science, made famous by the philosopher Karl Popper: falsification.
Popper was wrestling with a question — what makes a theory scientific? His answer flipped the intuitive picture upside down. We tend to think a theory earns its stripes by accumulating confirmations: the more facts that fit it, the more scientific and trustworthy it is. Popper said no. A theory’s strength comes not from what it predicts will happen but from what it forbids from happening. A scientific theory sticks its neck out: it rules things out, makes risky predictions that could turn out false. A theory that forbids nothing tells you nothing. If your theory is compatible with every conceivable observation, then no observation can ever count as evidence for it, because none could ever count against it.
This is exactly why “fits” teaches you nothing in the Wason task. Consider the asymmetry, because it’s the whole game:
- A confirming result (a “fits”) is consistent with many hypotheses at once. When
8, 10, 12fits, that’s compatible with “even +2,” “constant step,” “increasing,” “all positive,” “sum is even,” and dozens more. It can’t pick the winner out of the crowd. It eliminates nothing. - A disconfirming result (a “does not fit”) eliminates a hypothesis outright. If your rule is “even numbers” and you test
1, 2, 3and it comes back “fits,” that one result kills “even numbers” dead. No amount of further argument can save it. One negative test did what a thousand positive tests couldn’t.
That’s the logical engine under the whole thing. Confirmation is cheap — many theories share it. Disconfirmation is decisive — it belongs to one theory and destroys it. The informative test is always the one that could fail, because only a test that could fail can teach you anything when it doesn’t.
The unfalsifiable theory: comfortable and useless
Beware any belief you can’t imagine evidence against. “The market will go up, unless conditions change.” “He’s a genius — and when he fails, it’s because he was sabotaged.” “My diet works, and if you didn’t lose weight you did it wrong.” These feel unshakeable, but that’s the problem: a claim that survives every possible outcome has been tested by none of them. Unfalsifiable isn’t strong. It’s empty. The moment you notice your belief has an excuse ready for every contrary result, you’ve found a belief that’s never actually been on trial.
Worked example, outside the number puzzle. A teacher is convinced that a particular student, Mara, is lazy. Watch the positive-test strategy run: the teacher notices Mara doodling in class (fits — lazy), notices a late assignment (fits — lazy), notices her zoning out (fits — lazy). Pile of confirmations, growing confidence. But every one of those is a positive test — observations the teacher went looking for because they expected “lazy” to explain them, each equally consistent with “bored,” “struggling,” “anxious,” “exhausted,” or “bad eyesight at the back of the room.” The negative test — the falsifying move — is to ask: what would I see if Mara were NOT lazy, that I’d be unlikely to see if she were? Maybe: does she work hard on anything she’s allowed to choose? Does she stay after class to ask questions? If the teacher looks and finds Mara grinding for hours on a project she cares about, that single observation does to “lazy” what 1, 2, 3 does to “even numbers” — it kills it. Until the teacher seeks that disconfirming case, “lazy” is an unfalsifiable label collecting confirmations forever.
Fill in the logic of why disconfirmation beats confirmation.
Pick the right option for each blank, then check.
A confirming result is consistent with hypotheses at once, so it can't pick the winner and eliminates nothing. A disconfirming result a hypothesis outright — one such result can kill a guess that a thousand confirmations left standing. Popper's point: a theory that tells you nothing, because no observation could ever count against it.
Karl Popper argued that what makes a theory scientific is its ability to be falsified. Translated to the Wason task, what does that imply about the BEST move when testing a hypothesis?
Seeking the disconfirming case in real life
The 2-4-6 task is a sealed little world, but the failure it exposes is everywhere people form a hypothesis and then go gather evidence — which is to say, everywhere. The fix transfers cleanly: for any belief, stop asking “what would confirm this?” and start asking “what would disconfirm this — and have I actually gone and looked?”
- Medicine. A doctor who decides early it’s the flu orders only the tests that would confirm flu, reads ambiguous symptoms as flu, and discharges the patient — never running the one test that would have caught the meningitis the favored diagnosis conveniently explained away. This is premature closure, and it’s a leading cause of diagnostic error. The disconfirming move: what test would come back abnormal if this were NOT the flu, and have I ordered it?
- Hiring. You like a candidate in the first five minutes, then spend the interview asking questions that let them shine — collecting “fits.” The disconfirming question: what would a weak hire say to this hard question that a strong one wouldn’t? Ask the question designed to expose the flaw, not the one designed to confirm the glow.
- Relationships. Convinced your partner is distant lately, you notice every short reply and missed text (fits) and skip the warm moments. The disconfirming move: actively look for the evidence that you’re wrong — and ask, rather than infer.
- Investing. You’re bullish on a stock, so you read the bull case, follow the optimists, and dismiss the bears as clueless. The disconfirming move famously belongs to Charlie Munger and Warren Buffett: you don’t really understand a position until you can argue the other side better than the people who hold it.
- Debugging code. Your favorite theory is that the bug is in module A. You add logging that would prove it’s in A — and find some. The disconfirming move: run the test that should pass if A is fine. A green result there clears A and saves you hours of staring at innocent code.
Here is the same flip laid out as a tool you can steal. For each domain, the left column is the question confirmation bias wants you to ask; the right column is the one that actually finds the truth.
| Domain | The confirming question (feels productive, isn’t) | The disconfirming question (the one that finds truth) |
|---|---|---|
| Medicine | ”What symptoms fit my favored diagnosis?" | "What test would come back abnormal if this diagnosis were wrong — have I run it?” |
| Hiring | ”What can this candidate tell me that confirms they’re great?" | "What question would a weak hire stumble on that a strong one wouldn’t?” |
| Investing | ”What’s the strongest case for this position?" | "What would have to be true for this to be a terrible idea — and is any of it true?” |
| Debugging | ”What logging would show the bug is where I think it is?" | "What test should pass if my suspected cause is innocent?” |
Sort each move by whether it's a CONFIRMING search (piles up false confidence) or a DISCONFIRMING search (the one that finds truth).
Place each item in the right group.
- Noticing every short reply from your partner while skipping the warm moments
- Asking the interview question a weak candidate would visibly fumble
- Reading only the analysts who are bullish on the stock you already own
- Listing what would have to be true for your investment thesis to be wrong, then checking it
- Ordering one more lab test that would back up the diagnosis you already favor
- Running the test that should pass if your suspected bug location is actually innocent
The link to inversion
If this lesson is starting to feel familiar, that’s because you’ve met its logic before under a different name. Earlier on this path you learned inversion — the model of attacking a problem backward, from the failure state. Instead of asking “how do I succeed?”, you ask “how would I guarantee failure?” and then avoid all of that. Confirmation bias has an inversion-shaped cure baked right into it.
The lawyer-mind’s default question is “what evidence supports my belief?” That’s the forward question, and as you’ve now felt firsthand, it only ever collects “fits.” The inverted question — the negative test, the falsifying move, the Wason 1, 2, 3 — is: “What would have to be true for my belief to be FALSE — and is it?” You deliberately stand in the failure state, imagine the world where you’re wrong, and then go check whether you’re already living in it.
That’s not a coincidence; it’s the same machine. Inversion is the general thinking move. The negative test is inversion applied to evidence-gathering. Popper’s falsification is inversion applied to whole theories. When you seek the disconfirming case, you are inverting: looking for the route to “I’m wrong” instead of the route to “I’m right,” because the route to “I’m wrong,” if it stays empty after an honest search, is the only thing that ever earns you the right to believe you’re right.
How does seeking the disconfirming case relate to the inversion model you learned earlier?
Recap
You took the Wason task, ran it on yourself, and watched biased search happen in real time. The shape to carry forward:
Big picture
- Hunting for Yes
- Positive vs negative testing
- Positive test: a case you expect to confirm — piles up false confidence
- Negative test: a case you expect to disconfirm — the informative one
- The positive-test strategy is the default; it tests nothing
- The 2-4-6 task (Wason, 1960)
- Real rule: any three numbers strictly increasing
- “Even +2” is a tiny special case — every positive test passes
- Only ~1 in 5 solve it first try
- Falsification (Popper)
- A theory that forbids nothing tells you nothing
- “Fits” is consistent with many hypotheses — eliminates none
- “Does not fit” eliminates one outright — decisive
- The real-life move
- Medicine, hiring, relationships, investing, debugging
- Swap the confirming question for the disconfirming one
- Link to inversion
- Ask: what would make my belief FALSE — and is it?
- Negative test = inversion applied to evidence
- Positive vs negative testing
Check yourself: hunting for yes
In the Wason 2-4-6 task you just ran, why does collecting a long string of “fits” results fail to confirm a narrow guess like “even numbers ascending by two”?
Check your answer to continue.
You’ve now seen confirmation bias in its first habitat — the search for evidence — and felt why the only test worth running is the one you’d rather not. Next, in Same Evidence, Opposite Conclusions, we move from search to interpretation: what happens after the evidence is already in front of you, identical for everyone, and two opposed camps somehow each walk away more convinced they were right all along. The lawyer doesn’t just choose which evidence to gather. It rewrites what the evidence means.