Skip to content
Mental Models
🎯

Calibration: Knowing What You Know

Be right about how right you are.

A calibrated thinker's 70% guesses come true about 70% of the time. Measure your overconfidence, widen your honest error bars, score your beliefs with a Brier score, and turn vague hunches into probabilities you can actually be graded on.

You already learned to think in probabilities — to trade the yes/no switch for a dial, anchor on base rates, weigh bets by expected value, and update your beliefs by degrees with Bayes. This is the capstone of that path. It asks the question that sits on top of all the others: when you put a number on a belief, is the number any good?

Here is the whole idea in one line. A thinker is calibrated when their confidence matches their accuracy — when, of all the times they say “70% sure,” about 70% turn out to be right. Calibration is not about being smart, or even about being right more often. It is about being right about how right you are. A forecaster who says 90% and is correct 90% of the time is calibrated. So is one who says 60% and is correct 60% of the time. The failure isn’t being uncertain; the failure is claiming a certainty you haven’t earned.

And almost everyone fails it in the same direction. We are systematically overconfident: our 90% answers come true maybe 70% of the time, and the “90% confidence intervals” we draw around our estimates catch the real answer far less than 90% of the time. That gap — between the certainty we feel and the accuracy we actually have — is the single most common, most measurable, most fixable flaw in human judgment. This course measures it, names where it comes from, and hands you the tools to close it.

The route is hands-on. You’ll meet calibration precisely (the reliability curve that plots what you said against what came true) and drive an interactive calibration lab that scores your own confidence on a batch of true/false claims. You’ll separate the two faces of overconfidence — overestimation (your guesses run too high) and overprecision (your error bars run too narrow) — and feel the hard-easy effect that makes hard questions the most dangerous. You’ll learn to score a belief with proper scoring rules — the Brier score and the log score — that reward honesty and punish bluffing, so a probability becomes something you can be graded on like an exam. You’ll turn point estimates into honest ranges — 90% intervals wide enough to actually contain the truth. And finally you’ll learn how to get calibrated: track your predictions, give ranges, run calibration training, post-mortem your scores, and hold the two virtues that pull against each other — calibration (humility) and resolution (decisiveness) — at the same time, the way the best forecasters do.

By the end you’ll own the rarest meta-skill in all of judgment: not just estimating the odds, but knowing exactly how much to trust your own estimate — the honest error bar that every other model in this latticework quietly depends on.

In this topic

  1. 1 Calibration: Knowing What You Know A probability is a promise: say 70% and you're claiming that, over many such calls, you'll be right about 70% of the time. Calibration is whether you keep that promise — and almost everyone breaks it in the same overconfident direction. 8 min
  2. 2 What Calibration Means Plot your stated confidence against how often you're right and you get a reliability curve. Below the 45° line is overconfidence, above it is underconfidence — and weather forecasters live almost exactly on the line. 9 min
  3. 3 Overconfidence & Overprecision The two faces of overconfidence — overestimation vs. overprecision, the 90%-interval test that catches the truth half the time, the hard-easy effect, and why it costs you. 10 min
  4. 4 Scoring Your Beliefs Proper scoring rules turn a probability into a number the future can grade. Meet the Brier score and the log score — rigged so honesty wins and bluffing loses. 11 min
  5. 5 Thinking in Ranges Turn false-precision point estimates into honest confidence intervals: why your first error bar is always too narrow, how to widen it, and how to size a margin of safety to the range. 10 min
  6. 6 Getting Calibrated The habits that actually build calibration — tracking predictions, scoring with Brier, what superforecasters do — and the deep tension between calibration and resolution. 11 min
  7. 7 Final Exam: Calibration A graded, one-way final exam on calibration — the reliability curve, overconfidence and overprecision, the Brier and log scores, thinking in ranges, and getting calibrated. Pass mark 70%. 20 min

Mark course as finished

Done with every lesson? Lock it in — your progress is saved on this device.