Skip to content
Mental Models

Calibration: Knowing What You Know

Overconfidence & Overprecision

The two faces of overconfidence — overestimation vs. overprecision, the 90%-interval test that catches the truth half the time, the hard-easy effect, and why it costs you.

10 min Updated Jun 29, 2026

In lesson 1 you drew the reliability curve and watched almost everyone’s points sag below the diagonal — confidence outrunning accuracy. That sag has a name, overconfidence, and it is the single most reliable finding in the whole science of judgment. It shows up in students, doctors, engineers, CEOs, and CIA analysts. It barely budges with intelligence, education, or stakes.

But “overconfidence” is really a family of distinct failures wearing the same coat. This lesson takes the coat off. You’ll meet its two main faces — your numbers run too high, and your error bars run too narrow — see exactly how to catch each in the act, and learn why the failure gets worse precisely when you can least afford it: on the hard questions.

Before you read — take a guess

You ask 100 people for a range they're 90% sure contains the height of Mount Everest. On average, what fraction of those ranges will actually contain the true height?

The two faces of overconfidence

Imagine two ways an archer can be bad. The first aims at the wrong spot — every arrow lands a foot to the right of the bullseye. The second aims fine but swears the whole quiver will land inside a tiny circle, then watches arrows scatter all over the target. Same word — “overconfident” — two completely different problems, and a different cure for each.

That’s the core distinction of this lesson. Overconfidence splits into two faces:

  • Overestimation — overconfidence in the magnitude or likelihood. Your point guesses run too high and your probabilities run too hot: you say “90% sure” and you’re right 75% of the time. This is the sag below the diagonal on the reliability curve. The first archer.
  • Overprecision — overconfidence in the precision of your beliefs. Your error bars, ranges, and confidence intervals are drawn too narrow: your “90% interval” contains the truth half the time. You might be aimed at the right answer on average, but you radically understate how uncertain you are. The second archer.

(There’s a third, quieter face — overplacement, or the “better-than-average” effect: thinking you’re above the median driver, investor, or comedian, which can’t be true for most people at once. Worth knowing it exists, but this lesson stays on the two that wreck forecasts: overestimation and overprecision.)

Tip:

Why the split matters

The two faces need different fixes. Overestimation is cured by lowering your numbers — pull “90%” down toward “75%.” Overprecision is cured by widening your ranges — making your error bars bigger, not your guesses smaller. Confuse them and you apply the wrong medicine: shrinking a too-high probability does nothing for an error bar that’s too narrow.

A geologist says “I'm 95% sure this well will produce oil.” Across 100 such wells, exactly 70 produce. Then, separately, her 95% reserve-size intervals contain the true reserve only 55% of the time. Which faces of overconfidence are these?

Overprecision and the 90%-interval test

The cleanest way to catch overprecision is the experiment from the intro: ask for 90% confidence intervals on trivia and count how many contain the truth. If you’re honest, nine in ten should. People reliably land around four or five in ten — their ranges are roughly half as wide as honesty requires.

Here’s a worked run. Suppose someone gives these 90% intervals, then we check reality:

QuantityTheir 90% intervalTruthInside?
Length of the Nile (km)4,000 – 6,0006,650
Year Beethoven was born1740 – 17601770
Height of Mount Everest (m)8,000 – 9,0008,849
Boiling point of mercury (°C)200 – 300357
Number of bones in the human body200 – 240206
Distance Earth–Moon (thousand km)300 – 360384

Six “90% sure” ranges; two contained the truth. That’s a 33% hit rate against a claimed 90% — a 57-point calibration gap. And notice the pattern of the misses: the truth keeps landing just past the edge they drew. They weren’t aimed wildly wrong; they simply refused to make the ranges wide enough.

The cure is almost embarrassingly mechanical: deliberately widen every range. Take the interval that feels right, then push the low end lower and the high end higher until the range feels almost too generous — slightly uncomfortable, even a little silly. That discomfort is the feeling of an honest 90%. Lesson 4, Thinking in Ranges, gives this the full treatment.

Warning:

The narrow-range reflex

Overprecision is sneaky because a wide range feels like ignorance. Saying “the Nile is somewhere between 3,000 and 9,000 km” feels like admitting you don’t know — so you tighten it to sound competent, and shoot yourself in the foot. A calibrated 90% interval is supposed to feel too wide. If your ranges never feel slightly embarrassing, they’re too narrow.

Your last twenty 90%-confidence intervals contained the true answer only 9 times. What's the right correction?

The hard-easy effect

Overconfidence isn’t a flat tax. It scales with difficulty: the harder and more unfamiliar the question, the wider the gap between how sure you feel and how often you’re right. On genuinely easy questions the gap can even flip — people get a touch underconfident, right more often than they claim. This is the hard-easy effect.

Watch it in numbers. Imagine sorting your true/false answers into three buckets by how hard the questions were, and you say “80% sure” on all of them:

Question difficultyStated confidenceActual % correctGap
Easy80%87%+7 (slightly underconfident)
Medium80%72%−8 (overconfident)
Hard80%58%−22 (badly overconfident)

Same “80% sure” coming out of your mouth each time; three completely different truths behind it. On the easy bucket you actually undersold yourself. On the hard bucket your 80% was worth barely better than a coin flip. The takeaway isn’t “lower all your confidence” — that would wreck your easy-question calibration. It’s that your confidence needs to bend with difficulty far more than it naturally does.

On hard, unfamiliar questions you have less real signal — but your feeling of confidence doesn’t shrink to match, because it’s anchored to how fluently an answer comes to mind, not to how much you actually know. So feeling stays high while accuracy collapses, and the gap yawns open. On easy questions you have plenty of signal and tend to hedge needlessly, so accuracy quietly outruns confidence. Lesson: treat “this is a hard one” as a direct instruction to spread your probabilities toward 50% and widen your ranges.

According to the hard-easy effect, on which kind of question is a person's confidence MOST likely to exceed their accuracy?

Why we run hot (the mechanisms)

So why is the needle stuck on “too confident”? A handful of vivid mechanisms, each one a thread back to a model you’ve already met:

  • We don’t hunt for disconfirming evidence. Asked “will this launch succeed?” we list reasons it will and never run the reasons it won’t. That’s confirmation bias quietly inflating the probability from the inside.
  • We remember hits and forget misses. Your memory is a highlight reel of the times you called it right; the misses fade. So your felt track record is rosier than your real one, and tomorrow’s confidence borrows from a doctored past.
  • We’re socially rewarded for sounding certain. “I’m 60%, maybe?” sounds weak next to a confident bang on the table. Hedging gets punished in meetings even when it’s the more accurate answer, so we learn to perform certainty we don’t have.
  • The inside view ignores base rates. Estimating this project, this startup, this commute, we build up a rosy story from the specific details and skip the boring outside-view question: how did the last hundred projects like this go? That’s the engine of the planning fallacy — and base-rate neglect dressed up as optimism.
Info:

The common thread

Every mechanism pushes the same direction — up. That’s why overconfidence is nearly one-directional: it isn’t random noise, it’s a stack of biases all leaning the same way. Which is also good news. A systematic, predictable error is exactly the kind you can train against — that’s lesson 5.

The stakes: when running hot gets expensive

In a trivia quiz, overprecision costs you a few points. In the wild, it bankrupts funds and blows deadlines. Two flavors of expensive:

The estimate that slips. A team is “90% sure we ship in 6 weeks.” Their range was 6–6 weeks — zero width, the ultimate overprecision. The honest 90% interval, drawn from how similar projects actually went, was more like 6–14 weeks. They staffed, marketed, and signed contracts against the narrow number, and the four-week slip detonated all three. The dangerous part wasn’t the optimistic point guess; it was the missing width. A calibrated error bar would have told them to promise the client week 14, not week 6.

The move that “couldn’t happen.” A risk model pronounces a market crash a “10-sigma event” — so unlikely it shouldn’t occur once in the universe’s lifetime. Then it happens, twice in a week. The model wasn’t merely unlucky; it was overprecise about the tails, drawing its probability distribution far too narrow and declaring catastrophes impossible that were merely rare. (You met these underestimated extremes head-on in the fat tails course.) The too-confident model didn’t just miss — it told everyone there was nothing to hedge.

That’s the deep link to margin of safety. A margin of safety is the buffer you build because you might be wrong — and how much buffer you need is set entirely by how wide your honest error bar is. Overprecision shrinks the error bar to nothing, which whispers “no margin needed,” which is precisely when the rare thing arrives and there’s no cushion to absorb it. A calibrated range is the input that tells you how much safety margin to buy. Get the width wrong and every downstream safety decision is wrong with it.

A team commits to a deadline they're “95% sure” they'll hit, then misses it badly. Which diagnosis best fits — and what's the deeper cost?

Recap

The two faces of overconfidence, the test that catches each, the difficulty twist, and why width is the thing that protects you:

Question 1 of 40 correct

What's the difference between overestimation and overprecision?

Check your answer to continue.

Where this goes next

You can now name the disease and even eyeball it on a reliability curve or an interval hit rate. But “about half” and “a big gap” are still hand-wavy. Forecasting needs a single number that grades a whole batch of predictions at once — and that rewards honest probabilities while punishing the bluffer who shouts “99%.”

That number is the Brier score, and lesson 3, Scoring Your Beliefs, builds it from scratch with a worked example you can copy. Once you can score overconfidence, you can watch it shrink — which is the whole point of getting calibrated.

Mark lesson as complete