Skip to content
Mental Models

Calibration: Knowing What You Know

Getting Calibrated

The habits that actually build calibration — tracking predictions, scoring with Brier, what superforecasters do — and the deep tension between calibration and resolution.

11 min Updated Jun 29, 2026

You’ve spent four lessons learning what calibration is: confidence that matches hit rate, read off a reliability curve, undermined by overconfidence and overprecision, scored by the Brier score, and made honest with ranges instead of points. All of that was diagnosis. This lesson is the cure. It answers the only question that finally matters: how do you actually become calibrated — and why that turns out to be a learnable skill rather than a personality you were born with (or without).

There’s a twist waiting at the end, too. Becoming humble about uncertainty is only half the job. A weather forecaster who says “50% chance of rain” every single day is perfectly honest and completely useless. The expert move is to be calibrated and decisive at the same time — and that tension is where this whole course has been heading.

Before you read — take a guess

Decades of research show that weather forecasters are among the best-calibrated people on Earth. What's the single biggest reason WHY?

Calibration is trainable

Here is the genuinely hopeful claim at the heart of this course: calibration is a skill you can drill, not a trait you’re stuck with. You can’t easily train “be smarter” or “have better taste.” But you can train “make your 80% mean 80%,” because — uniquely among thinking virtues — it’s measurable. You can see the gap between your confidence and your accuracy as a number, and anything you can measure, you can practice closing.

The evidence is strong and comes from the most-measured predictors alive. Weather forecasters are famously well calibrated: when they say “70% chance of rain,” it rains on close to 70% of those days. So are expert bridge players and seasoned bookmakers. What do these groups share? Not humility, not genius — a tight feedback loop: many predictions, each with an explicit probability, each promptly scored by reality. People who get that loop get calibrated. People who don’t, stay overconfident their whole lives, because nothing ever forces the gap into the open.

That gives us the core training loop, the spine of everything below:

  1. Predict — commit to an explicit probability before the outcome.
  2. Record — write it down so memory can’t quietly edit it later.
  3. Score — when the outcome lands, grade it (Brier score, reliability curve).
  4. Review — look at the pattern, especially your confident misses.
  5. Adjust — nudge your future confidence toward your real hit rate, and repeat.
Tip:

The whole lesson in one line

Calibration improves the same way a free throw does: many reps, honest scoring, and feedback you actually look at. The rest of this lesson is just five concrete ways to build that loop into how you think.

Practice 1 — Track your predictions

Memory is a liar, and it lies in a specific direction. After you learn an outcome, your brain quietly rewrites how sure you were — “I knew it all along.” That’s hindsight bias, and it’s poison for calibration: if you misremember your fuzzy 60% hunches as confident 90% calls (or your wrong 90% calls as cautious maybes), you can never see your real track record, so you can never fix it.

The fix is almost insultingly simple: write your predictions down, with a probability and a date, before the outcome is known. Keep a prediction journal (or forecasting log) — a running list the future can’t edit. Once it’s on paper, hindsight has nothing to grab. The number you wrote is the number you wrote.

A few entries from a real-looking log:

DatePredictionConfidenceResolvedOutcome
Jan 3Project ships by end of Q180%Mar 31Missed (shipped Apr 9)
Jan 12Candidate A accepts our offer65%Jan 20Hit
Feb 1This feature lifts signups55%Mar 1Missed
Feb 4Rent on our lease goes up at renewal90%May 1Hit

A worked read of that log: of the predictions made at “high confidence” (80–90%), one of two resolved correctly so far. Too few to judge anything yet — and that’s the point of the journal. You don’t trust four predictions; you accumulate dozens, then the pattern (your 80%s only come true 65% of the time, say) becomes undeniable in a way no amount of “I’m usually right” self-talk can fake.

Warning:

The pitfall: predictions you can't grade

A prediction journal only works if each entry is resolvable — a clear outcome and a clear date. “The economy will struggle” can never be scored (struggle how? by when?). “GDP growth is below 1% this calendar year” can. Vague predictions feel safe precisely because they dodge the grade — which is exactly why they teach you nothing.

When to use it

Start a journal for any domain where you make repeated, checkable calls: project deadlines, hiring, sales forecasts, sports, your own “this’ll take an hour” estimates. The mundane, high-volume predictions are the best training data, because they resolve quickly and often — fast feedback beats important-but-rare.

Practice 2 — Give ranges and explicit probabilities

Two reflexes do enormous work, and you already met both. From Thinking in Ranges: give ranges, not points. A point estimate (“revenue will be $2M”) is untrackable and almost always overprecise. A 90% range (“$1.4M to $2.7M”) states your real uncertainty and can be honestly scored — did the truth land inside it or not? Across many such ranges, about 90% should contain the answer. If far fewer do, your bars are too narrow (overprecision); if nearly all do, you’re sandbagging.

From Scoring Your Beliefs: attach an explicit probability to every yes/no call, so it becomes scorable. “Probably” is not a forecast — one person’s “probably” is 55%, another’s is 90%. A number is unambiguous and gradeable. Vague words are comfortable exactly because they can never be proven wrong; that comfort is the enemy of getting better.

A colleague says “there's a good chance this launch slips.” You're trying to build a calibration habit. What's the single most useful thing to do?

When to use it

Default to ranges for any quantity (time, money, counts) and to explicit percentages for any yes/no event. The friction you feel writing “90% sure it’s between X and Y” instead of a confident single number is the calibration happening — it forces you to confront how much you actually don’t know.

Practice 3 — Calibration training and feedback

A journal of unscored predictions is a diary, not a gym. The rep that actually builds the muscle is scoring a batch and looking at the result. This is calibration training: answer a set of questions with explicit probabilities, grade them, and plot where your confidence and your accuracy diverge.

The drill, concretely:

  1. Take a batch of true/false claims (or resolved predictions from your journal).
  2. For each, state your probability that it’s true.
  3. Score the batch with the Brier score from Scoring Your Beliefs — the mean squared distance between your probabilities and what actually happened. Lower is better; honest probabilities minimize it.
  4. Plot your reliability curve from What Calibration Means: group your answers by stated confidence, and for each bucket compare “how sure you said you were” to “how often you were actually right.” Points above the diagonal mean you were underconfident in that bucket; points below mean overconfident.
  5. Read the gap and adjust next time.

If your “90% sure” bucket only comes true 70% of the time, the curve shows it instantly — your high-confidence points sag below the diagonal — and the fix writes itself: when you feel “90% sure,” say 75% until the curve straightens out.

Info:

You already have the gym

The calibration lab from Lesson 1 is exactly this drill. It hands you a batch of true/false claims, takes your confidence on each, and scores you against the diagonal. That wasn’t a one-time demo — it’s the rep you’re meant to repeat. Going back and running it again, watching the gap shrink across attempts, is calibration training in its purest form.

You run a 50-question calibration batch. Your “80% confident” answers turn out correct only 55% of the time, but your “60% confident” answers are right almost exactly 60% of the time. What does this tell you?

When to use it

Run a scored batch periodically — monthly is plenty — rather than obsessing over single predictions. Single outcomes are noisy (a 90% call that fails proves nothing on its own); the pattern across a batch is the signal. Treat it like weighing yourself: the trend over many readings, not any one day.

Practice 4 — Post-mortem your scores

Scoring tells you that you’re off. A post-mortem tells you why — and it’s where the deepest learning lives. Keep a decision journal: alongside each prediction, jot the reasoning behind it. Then, when outcomes land, review them — hits and misses both, but especially your confident misses, the times you said 90% and reality said no. Those are where your mental model is most wrong and has the most to teach.

The trap lurking here is outcome bias — judging a decision by how it turned out rather than by whether it was sound given what you knew. As you learned back in Thinking in Probabilities: a good decision can have a bad outcome, and a bad decision can get lucky. If you bet on a 90% favorite and it loses, that was a good call with an unlucky result — punishing yourself rewires you toward timidity, not accuracy. The post-mortem’s real job is to separate decision quality from outcome luck: was the probability honest given the information at the time, or did you genuinely misjudge?

Worked example. You said “85% this hire works out,” and six months later it didn’t. Two very different post-mortems:

  • Decision quality was fine, outcome was unlucky: the candidate looked strong on every signal you could see; an unforeseeable family emergency pulled them away. Lesson: none — 85% calls fail 15% of the time, and this was one. Don’t lower your confidence on the next strong candidate.
  • Decision quality was bad, outcome merely exposed it: you ignored two lukewarm references because you’d already fallen in love with the résumé. Lesson: 85% was overconfident; you discounted disconfirming evidence. That you fix.

Only the second is a calibration error. Conflate them and you’ll “learn” the wrong lesson from noise.

Warning:

The pitfall: only reviewing the misses

It’s tempting to post-mortem disasters and ignore wins. But a lucky win from a bad decision is just as misleading as an unlucky loss from a good one — it teaches you a sloppy process “works.” Review confident hits too, and ask the same question: was I right for the right reasons, or right by luck?

When to use it

Post-mortem after any prediction resolves where the stakes or the surprise were high — and always after a confident miss. The single most valuable habit is the question “was my probability honest given what I knew then?”, asked with the outcome deliberately set aside.

Practice 5 — Start from base rates (the outside view)

Before you reach for an inside-view story about why this case is special, anchor on the base rate: how often does this kind of thing happen in general? This is the outside view (or reference-class forecasting) — find the reference class your situation belongs to, start from its base rate, then adjust for specifics.

A worked example. “How likely is our software project to ship on its 3-month deadline?” The inside view spins an optimistic tale: the team is sharp, the plan is tight, 85%. The outside view asks a cooler question: of comparable projects, how many hit their original deadline? Historically, most software projects ship late — so the reference class might say more like 30%. The calibrated forecaster starts near 30% and nudges up only for genuine reasons this team is exceptional, landing somewhere honest like 45% — not the wishful 85%.

This is the same machinery as Bayesian updating: the base rate is your prior, and starting from it gives you a calibrated prior before any specific evidence moves you. Skip the base rate and you’re updating from a number you made up — garbage in, garbage out. Anchor on it, and your evidence does honest work on an honest starting point.

A founder pitches you: “90% chance we'll be profitable within two years — our team is world-class.” From a calibration standpoint, what's the most likely problem?

When to use it

Reach for the outside view first whenever you’re forecasting a kind of event that has happened many times before (projects, hires, launches, treatments). It’s most powerful exactly when you feel most special — “this time is different” is usually the inside view talking, and the base rate is usually right.

What the superforecasters do

You don’t have to take all this on faith. The most rigorous evidence comes from political scientist Philip Tetlock and the Good Judgment Project — a multi-year forecasting tournament in which thousands of ordinary volunteers predicted real-world geopolitical and economic events, with every forecast scored against what actually happened. A subset of these volunteers — dubbed superforecasters — were consistently, measurably better than the crowd, and in some comparisons outperformed subject-matter experts with access to classified information.

The striking finding, and the reason it belongs in this lesson: they weren’t geniuses or insiders. They were ordinary people with better habits — the exact habits you’ve just read. Drawn from that research, superforecasters tend to:

  • Break big questions into smaller, answerable pieces instead of reacting to the whole.
  • Start from base rates — the outside view — before layering in specifics.
  • Update in small increments as evidence trickles in, rather than lurching from certain to certain (pure Bayesian updating, done gently).
  • Think in fine-grained probabilities — they meaningfully distinguish 63% from 70%, where most of us round everything to “likely,” “maybe,” or “no chance.”
  • Keep score and post-mortem — they revisit their misses and ask what their model got wrong.
Info:

No magic, no exact numbers

The honest version of this story is qualitative. Superforecasting is well-supported research showing that trainable habits, not innate brilliance, drive good forecasting — and that calibration improves with practice and feedback. It is not a claim that anyone can predict anything, or a specific guaranteed accuracy. The defensible takeaway is the hopeful one: the loop works, for ordinary people.

Because rounding throws away real information. If you collapse everything into “likely / maybe / unlikely,” you can’t tell a strong 70% from a marginal 55% — and a scoring rule like the Brier score rewards you for the difference. Superforecasters treat the gap between 63% and 70% as meaningful because, across hundreds of forecasts, those small distinctions add up to a measurably better track record. It’s the difference between a thermometer with one mark and one with a hundred.

Calibration vs. resolution — the central tension

Now the expert payoff, the idea this whole course has been climbing toward. There are two virtues in a good forecast, and they pull in different directions.

Calibration is honesty about uncertainty — the humility virtue. Of all the times you say 70%, you’re right about 70% of the time.

Resolution (also called discrimination or decisiveness) is the willingness to move off 50% — to make sharp, confident calls that actually separate the things that will happen from the things that won’t. It’s the courage virtue.

Here’s the catch that makes this an expert skill: you can be perfectly calibrated and still useless. The forecaster who says “50% chance of rain” every single day is flawlessly honest — over a year, it rains about half the time, so their 50%s come true 50% of the time. Perfect calibration. Zero resolution. They’ve never once told you to bring an umbrella. A broken clock and “always say 50%” share this curse: honest, and worthless. Compare two forecasters:

Forecaster A: “always 50%“Forecaster B: confident calls
Typical forecast50% rain, every day90% on rainy days, 10% on clear days
CalibrationPerfect — 50% claims hit 50%Good — 90%s hit ~90%, 10%s hit ~10%
ResolutionZero — never separates outcomesHigh — sharply sorts rain from shine
Useful?No — tells you nothingYes — you know when to bring an umbrella

Forecaster B is the goal: calibrated and decisive. They commit to 90% when the evidence warrants it, 10% when it doesn’t — and they’re right about that often. That’s a forecast you can act on.

This is exactly why the Brier score from Scoring Your Beliefs decomposes into (among other terms) a calibration component and a resolution component. A good Brier score requires both: being honest (calibration) and being sharp (resolution). “Always say 50%” gets you perfect calibration but terrible resolution, and the Brier score sees right through it. The decomposition is the mathematical proof that humility alone isn’t enough — you have to be brave with your probabilities too.

Two forecasters are scored over a rainy year. Aiko says “50% chance of rain” every day. Ben says “90%” on days it rains and “10%” on days it doesn't, and he's right about that often. Who's the better forecaster, and why?

When to use it

Watch for both failure modes in yourself. If you’re the type who hedges everything to “could go either way,” your problem is resolution — push yourself to commit to sharper numbers when the evidence supports them. If you’re the type who’s loudly certain and often wrong, your problem is calibration — rein the confident calls back toward your real hit rate. The master forecaster is honest and brave.

Putting it together: where calibration sits in the latticework

Step back and notice what a calibrated error bar really is: the honest input that the rest of your thinking quietly depends on. Margin of safety is only as good as the worst case you’re willing to take seriously — set your range too narrow (overprecision) and the margin is a comforting fiction. Expected value is a weighted average of probabilities — feed it overconfident numbers and the whole calculation is garbage, however clean the arithmetic. Bayesian updating starts from a prior and moves with evidence — and a calibrated prior, anchored on the base rate, is what makes the update trustworthy. Calibration isn’t one more model beside these; it’s the quality control on the inputs they all share. Get your probabilities honest and sharp, and every other quantitative model you own starts telling you the truth.

Recap

You’ve reached the end of the teaching. The arc, in one breath: calibration is confidence that matches reality (Lesson 1), wrecked by overconfidence and overprecision (Lesson 2), measured by proper scoring (Lesson 3), expressed as honest ranges (Lesson 4), and built through deliberate practice (this one) — track your predictions, give ranges and explicit probabilities, run scored calibration batches, post-mortem your hits and misses, and start from base rates. The superforecasters prove the loop works for ordinary people. And the expert insight is that humility isn’t enough: aim to be calibrated and decisive.

Question 1 of 40 correct

What's the core reason calibration counts as a trainable skill rather than a fixed trait?

Check your answer to continue.

The exam awaits

That’s every idea in the course. There’s no Lesson 6 — what comes next is the Final Exam, and it plays by stricter rules than anything you’ve seen so far: it’s graded, served one question at a time, and it locks each answer the moment you submit it. No back button, no retries, no restart — just like a real forecast, where you commit and the future grades you. You’ll need 70% to pass, and you won’t see your score until the end.

Before you start, one last rep of the habit that runs through this whole course: commit to a number before you reveal anything. Calibration was never about sounding sure — it was about your confidence telling the truth. Go prove yours does.

Mark lesson as complete