What is a calibrated forecast? Why 70% should happen 7 times out of 10
A calibrated forecaster's 70% calls come true about 70% of the time. What calibration means, how to check it, and why people and AI are often overconfident.
Read the articleThe Brier score measures how close a probability was to what happened. Lower is better; always saying 50% scores 0.25. Formula, examples and good scores.

The Brier score measures how good a probability forecast was. It is the squared difference between the probability you gave and what actually happened, averaged over all your forecasts. Lower is better: 0 is perfect, 0.25 is what you get by always saying 50%, and 1 is the worst possible.
It is also how you check any claim that AI can predict the future. It was introduced by the meteorologist Glenn W. Brier in 1950 to grade weather forecasts (Brier, 1950), and it is now the standard way to score forecasters, from weather services to AI superforecasters.
For a yes-or-no question, the Brier score of one forecast is:
Brier score = (forecast − outcome)²
For many forecasts, add up the scores and divide by the number of forecasts.
| You said | Did it happen? | Calculation | Score |
|---|---|---|---|
| 90% | Yes | (0.9 − 1)² | 0.01 |
| 90% | No | (0.9 − 0)² | 0.81 |
| 70% | Yes | (0.7 − 1)² | 0.09 |
| 70% | No | (0.7 − 0)² | 0.49 |
| 50% | Either | (0.5 − 1)² or (0.5 − 0)² | 0.25 |
| 10% | No | (0.1 − 0)² | 0.01 |
Two patterns stand out:
Because the error is squared, a few overconfident misses can wreck an otherwise good record.
Say you made four forecasts:
Your Brier score is the average: (0.04 + 0.09 + 0.36 + 0.0025) ÷ 4 = 0.123. That is clearly better than the 0.25 you would get by saying 50% every time.
It depends heavily on how hard the questions are, but these rules of thumb help:
Never compare Brier scores across different question sets. Easy questions ("Will the sun rise tomorrow?") make everyone look brilliant. Only compare forecasters on the same questions.
The Brier score is a proper scoring rule: your expected score is best when you report exactly what you believe. If you think something is 70% likely, saying 90% to sound bold makes your expected score worse, and so does hedging down to 50%.
This is what makes it fair. It rewards honesty, not confidence and not caution. It is also why the score pairs so naturally with calibration: a forecaster who wants a good Brier score has every reason to make their probabilities mean what they say.
In 1973 Allan Murphy showed that the Brier score splits into three parts (Murphy, 1973):
A forecaster who always says the average rate can be perfectly calibrated and still useless, because they have no resolution. Good forecasters are calibrated and decisive when the evidence allows.
To compare forecasters fairly, use the Brier skill score:
BSS = 1 − (your Brier score ÷ baseline Brier score)
The baseline is usually a simple strategy, such as always forecasting the historical base rate. A skill score above 0 means you beat the baseline. It is how benchmarks compare AI forecasters with humans and prediction markets in AI vs. superforecasters.
Log loss is another proper scoring rule. It punishes confident misses even more severely: a 100% forecast that turns out wrong gets an infinite penalty. The Brier score is bounded, easier to explain, and more forgiving of a single disaster, which is why most forecasting tournaments use it.
Brier's 1950 paper summed the error over both possible outcomes, which gives a range of 0 to 2 for a yes-or-no question. Most forecasting work today uses the simpler version on this page, from 0 to 1. Rankings are the same; only the scale differs.
Lower. A Brier score of 0 means every forecast was perfect. Always forecasting 50% scores 0.25.
0.25. Saying 50% scores (0.5 − 1)² = 0.25 if the event happens and (0.5 − 0)² = 0.25 if it does not.
Calibration asks whether your 70% forecasts come true about 70% of the time. The Brier score measures overall accuracy and includes calibration as one of its three parts, together with resolution and uncertainty.