Skip to content

Calibration

Calibration is the agreement between stated confidence and observed frequency. A calibrated forecaster who says seventy percent is right about seventy percent of the time across many such calls. It is a property of a record rather than of any single judgement, so it can only be measured over a series.

Calibration is measured with scoring rules applied to probabilistic predictions, which requires that the predictions were recorded before the outcome.

Why it matters when the plan changes

Confidence is the part of a judgement a decision-maker acts on, and it is the part that is almost never checked. A forecaster who is consistently right and consistently overconfident is dangerous: their correct calls carry the same weight as their wrong ones, so a reader cannot tell how much to rely on any of them. Scoring the probability itself, not just the direction of the call, is what a proper scoring rule such as the Brier score is built to do, which is why a single win proves nothing.

The tension is that calibration demands a record and a wait. It cannot be asserted at the start of a relationship, because there is nothing yet to compute it from. Any supplier claiming it before outcomes exist is describing an intention; the honest version says what is being recorded and when the number will be available. Scored feedback drives the improvement: Good Judgment Project forecasters improved once their calls were scored, which is the argument for waiting.

In practice

A judgement made with stated high confidence turns out wrong. Read alone, it is a failure. Read against a record of thirty judgements, it is one expected miss among calls at that confidence, and the record shows whether stated confidence has been tracking observed outcomes. Only the second reading is usable.

Evidence

What it cannot tell you

Calibration tells you whether stated confidence matches observed frequency across many calls; it says nothing about a single judgement, which cannot be scored in isolation. It is also silent on the value or relevance of what was predicted: a forecaster can be perfectly calibrated on questions that carry no decision weight at all.

Questions

Accuracy asks how often the call was right. Calibration asks whether the stated confidence matched what happened: things called seventy percent likely occurring about seventy percent of the time. A forecaster can be accurate and badly calibrated, or well calibrated and rarely useful.

Enough that the frequencies mean something, which is tens of calls rather than a handful. The Good Judgment Project scored forecasters over many rounds before calibration figures were meaningful, because a handful of calls cannot separate a well-calibrated forecaster from a lucky one.

A way of scoring probabilistic predictions that gives the best expected score to a forecaster who reports what they actually believe. That property matters because it removes the incentive to hedge toward the middle, or to overstate confidence for effect, in order to score better.

Yes, and it is one of the better established findings in the field. The Good Judgment Project scored forecasters' calls using Brier scores, and those who received that scored feedback improved measurably, while forecasters who never saw their record did not improve at all.

Not recording the forecast before the outcome, changing what was predicted after the fact, and predicting things too vague to be scored. Each removes the comparison the measure depends on, and each is common enough that most professional judgement is never calibrated at all.