Skip to content
Mental Models

Calibration: Knowing What You Know

Calibration: Knowing What You Know

A probability is a promise: say 70% and you're claiming that, over many such calls, you'll be right about 70% of the time. Calibration is whether you keep that promise — and almost everyone breaks it in the same overconfident direction.

8 min Updated Jun 29, 2026

Here is a small, uncomfortable experiment you can run on yourself. Write down ten hard trivia questions — the year a country was founded, the length of a river, the population of a city — and for each, instead of a single guess, give a range you are 90% sure contains the true answer. As wide as you like. Now look up the answers. If you were honestly 90% confident, about nine of your ten ranges should contain the truth.

Almost nobody gets nine. Most people get five or six. Their “90% sure” ranges are so narrow that the truth slips out of them half the time. The confidence they felt — ninety percent — and the accuracy they had — sixty percent — were two completely different numbers, and they never noticed the gap. That gap has a name, it is measurable, it is nearly universal, and closing it is the entire point of this course.

The skill is called calibration: the match between how confident you are and how often you’re right. A calibrated thinker is one whose 70% predictions come true about 70% of the time, whose 90% predictions come true about 90% of the time, all the way up. It is the capstone of everything you learned in Thinking in Probabilities and Bayesian Updating — because a probability is only worth something if the number is honest, and calibration is the test of whether it is.

Before you read — take a guess

A weather forecaster says “70% chance of rain” on 100 different days. On how many of those days should it actually rain, if the forecaster is perfectly calibrated?

Calibration is not the same as being right

It is tempting to think calibration just means “being accurate,” but it is a subtler, stranger thing. Calibration is about the honesty of your confidence, not the frequency of your correctness.

Picture two people predicting coin flips of a fair coin. The first says “heads, 50%” every time. The second says “heads, 99%” every time. Both are right about half the time — identical accuracy. But the first is perfectly calibrated (they claimed 50% and were right 50% of the time) and the second is wildly overconfident (they claimed 99% and were right 50% of the time). Same hit rate; completely different quality of judgment. The second person is the one who will blow up a portfolio, miss a deadline, and walk into the storm they were “99% sure” wouldn’t come.

Tip:

The one-sentence version

Calibration is the match between your stated confidence and your actual hit rate: of all the times you say X%, you should be right about X% of the time. It is not about being smarter or more certain — it is about your confidence telling the truth. The first question to ask of any forecast, including your own, is: when this person says 80%, how often are they actually right?

Why almost everyone is overconfident

If calibration errors were random — sometimes too confident, sometimes too humble — they’d be annoying but not dangerous. The trouble is they’re not random. Across decades of studies, on tasks from trivia to medical diagnosis to engineering estimates, people lean overwhelmingly one way: overconfident. Our 90% answers come true around 75% of the time. The “I’m certain” answers fail far more than certainty should allow.

There are good reasons for it. We don’t naturally search for reasons we might be wrong (you met that as confirmation bias). We remember our hits and quietly forget our misses. We confuse the vividness of a story with its probability. And we’re rewarded socially for sounding sure — “it depends” and “about 60%, maybe” feel weak next to a confident bang on the table, even when the hedged answer is the better one. The result is a thinking tool that runs hot by default, and a course built to cool it down.

Which person is BETTER calibrated, even though both are wrong this time?

A number you can be graded on

The quiet superpower of calibration is that it makes a belief scorable. Once you attach a probability to a claim — “70% this ships on time,” “85% this diagnosis is right,” “30% this startup survives five years” — you’ve written down something the future can grade. Say 70% and it happens, you earned some points; say 99% and it fails, you lose a lot. Do this across hundreds of predictions and a true picture of your judgment emerges, one that flattering memory can’t fake. Later in the course you’ll meet the actual scoring rule — the Brier score — that does this grading, and you’ll see why it’s rigged to reward honesty and punish bluffing.

That’s what lifts calibration from a personality trait (“she’s so sure of herself”) to a measurable, trainable skill. You can’t easily train “be smarter.” You can train “make your 80% mean 80%,” because you can measure it, see the gap, and adjust — which is exactly what the best forecasters do.

What's the main practical advantage of stating a belief as a probability (“70% likely”) instead of a vague phrase (“probably”)?

The map of the course

Five teaching lessons, then a final exam you can’t undo. The climb:

  1. What Calibration Means — the idea made precise: the reliability curve that plots your stated confidence against your real hit rate, over- vs. under-confidence read straight off the diagonal, and a hands-on calibration lab that scores your own confidence on a batch of true/false claims.
  2. Overconfidence & Overprecision — the two faces of running hot: overestimation (your guesses are too high) and overprecision (your error bars are too narrow), the 90% intervals that catch the truth half the time, and the hard-easy effect that makes hard questions the most treacherous.
  3. Scoring Your Beliefs — proper scoring rules: the Brier score and the log score, why they’re mathematically rigged so that honesty maximizes your score, and a fully worked numeric example you can copy.
  4. Thinking in Ranges — turning point estimates into honest 90% intervals wide enough to actually contain the answer, why your first instinct is always too narrow, and how ranges tie back to margin of safety.
  5. Getting Calibrated — the practice: track predictions, give ranges, run calibration training, post-mortem your scores; what the superforecasters do; and the deep tension between calibration (humility) and resolution (decisiveness) — being honest about uncertainty and still useful. Ends with a whole-course recap.

Then a Final Exam — graded, one question at a time, one-way: once you answer, it locks. No back button, no retries.

How to use this course

One habit does most of the work: before you reveal any answer, commit to a number. When an exercise asks what you think, don’t just nod along — actually pick “about 75%” in your head first. The small sting of watching your private number miss is what trains the gap shut. The exercises are the lesson; the prose just sets them up.

Next up: lesson 1, where we make calibration precise and you score your very own confidence on the reliability curve.

Mark lesson as complete