# The Brier score, explained with examples

> The Brier score measures how close a probability was to what happened. Lower is better; always saying 50% scores 0.25. Formula, examples and good scores.

Sep 28, 2026 · Methods · Sikt Intelligence · https://www.siktintelligence.com/blog/brier-score-explained

The **Brier score** measures how good a probability forecast was. It is the squared difference between the probability you gave and what actually happened, averaged over all your forecasts. **Lower is better: 0 is perfect, 0.25 is what you get by always saying 50%, and 1 is the worst possible.**

It is also how you check any claim that [AI can predict the future](/blog/can-ai-predict-the-future). It was introduced by the meteorologist Glenn W. Brier in 1950 to grade weather forecasts ([Brier, 1950](https://journals.ametsoc.org/view/journals/mwre/78/1/1520-0493_1950_078_0001_vofeit_2_0_co_2.xml)), and it is now the standard way to score forecasters, from weather services to [AI superforecasters](/blog/what-is-an-ai-superforecaster).

## The formula

For a yes-or-no question, the Brier score of one forecast is:

> **Brier score = (forecast − outcome)²**

- The **forecast** is the probability you gave, written from 0 to 1. So 70% is 0.7.
- The **outcome** is 1 if the event happened and 0 if it did not.

For many forecasts, add up the scores and divide by the number of forecasts.

## Worked examples

| You said | Did it happen? | Calculation | Score |
|---|---|---|---|
| 90% | Yes | (0.9 − 1)² | 0.01 |
| 90% | No | (0.9 − 0)² | 0.81 |
| 70% | Yes | (0.7 − 1)² | 0.09 |
| 70% | No | (0.7 − 0)² | 0.49 |
| 50% | Either | (0.5 − 1)² or (0.5 − 0)² | 0.25 |
| 10% | No | (0.1 − 0)² | 0.01 |

Two patterns stand out:

- **Confident and right is rewarded.** A 90% call that happens scores 0.01.
- **Confident and wrong is punished hard.** A 90% call that misses scores 0.81, more than three times a coin flip.

Because the error is squared, a few overconfident misses can wreck an otherwise good record.

## A four-forecast example

Say you made four forecasts:

1. 80% that a central bank cuts rates. It did. Score (0.8 − 1)² = **0.04**
2. 30% that a launch happens this year. It did not. Score (0.3 − 0)² = **0.09**
3. 60% that a bill passes. It did not. Score (0.6 − 0)² = **0.36**
4. 95% that a company reports on time. It did. Score (0.95 − 1)² = **0.0025**

Your Brier score is the average: (0.04 + 0.09 + 0.36 + 0.0025) ÷ 4 = **0.123**. That is clearly better than the 0.25 you would get by saying 50% every time.

## What is a good Brier score?

It depends heavily on how hard the questions are, but these rules of thumb help:

- **0.25** is the coin-flip line: saying 50% on everything. Anything above it is worse than not trying.
- **Around 0.2** means you are adding some information, but not much.
- **Below 0.15** is solid on genuinely uncertain questions.
- **Below 0.10** on hard, real-world questions is excellent, the territory of top forecasters.

Never compare Brier scores across different question sets. Easy questions ("Will the sun rise tomorrow?") make everyone look brilliant. Only compare forecasters on the same questions.

## Why you cannot game it

The Brier score is a **proper scoring rule**: your expected score is best when you report exactly what you believe. If you think something is 70% likely, saying 90% to sound bold makes your expected score worse, and so does hedging down to 50%.

This is what makes it fair. It rewards honesty, not confidence and not caution. It is also why the score pairs so naturally with [calibration](/blog/forecast-calibration): a forecaster who wants a good Brier score has every reason to make their probabilities mean what they say.

## What the score is made of

In 1973 Allan Murphy showed that the Brier score splits into three parts ([Murphy, 1973](https://www.researchgate.net/publication/234395762_A_New_Vector_Partition_of_the_Probability_Score)):

- **Reliability:** how well your probabilities match reality. This is calibration. Lower is better.
- **Resolution:** how well your forecasts separate the events that happen from those that don't. Higher is better.
- **Uncertainty:** how unpredictable the questions were in the first place. You cannot control this.

A forecaster who always says the average rate can be perfectly calibrated and still useless, because they have no resolution. Good forecasters are calibrated *and* decisive when the evidence allows.

## Brier skill score: comparing against a baseline

To compare forecasters fairly, use the **Brier skill score**:

> **BSS = 1 − (your Brier score ÷ baseline Brier score)**

The baseline is usually a simple strategy, such as always forecasting the historical [base rate](/blog/base-rates-forecasting). A skill score above 0 means you beat the baseline. It is how benchmarks compare AI forecasters with humans and [prediction markets](/blog/how-prediction-markets-work) in [AI vs. superforecasters](/blog/ai-vs-superforecasters).

## Brier score versus log loss

**Log loss** is another proper scoring rule. It punishes confident misses even more severely: a 100% forecast that turns out wrong gets an infinite penalty. The Brier score is bounded, easier to explain, and more forgiving of a single disaster, which is why most forecasting tournaments use it.

## A note on the original formula

Brier's 1950 paper summed the error over both possible outcomes, which gives a range of 0 to 2 for a yes-or-no question. Most forecasting work today uses the simpler version on this page, from 0 to 1. Rankings are the same; only the scale differs.

## Key takeaways

- Brier score = (forecast − outcome)², averaged over many forecasts. Lower is better.
- 0 is perfect, 0.25 is always saying 50%, 1 is certain and wrong.
- Confident misses are punished hard, because the error is squared.
- It is a proper scoring rule: honesty gives the best expected score.
- Only compare Brier scores on the same questions, ideally with a skill score.

## FAQ

### Is a lower or higher Brier score better?

Lower. A Brier score of 0 means every forecast was perfect. Always forecasting 50% scores 0.25.

### What is the Brier score of a coin flip?

0.25. Saying 50% scores (0.5 − 1)² = 0.25 if the event happens and (0.5 − 0)² = 0.25 if it does not.

### What is the difference between the Brier score and calibration?

Calibration asks whether your 70% forecasts come true about 70% of the time. The Brier score measures overall accuracy and includes calibration as one of its three parts, together with resolution and uncertainty.

## Sources

- Brier, G. W. (1950): [Verification of forecasts expressed in terms of probability](https://journals.ametsoc.org/view/journals/mwre/78/1/1520-0493_1950_078_0001_vofeit_2_0_co_2.xml), Monthly Weather Review
- Murphy, A. H. (1973): [A new vector partition of the probability score](https://www.researchgate.net/publication/234395762_A_New_Vector_Partition_of_the_Probability_Score), Journal of Applied Meteorology
- Wikipedia: [Brier score](https://en.wikipedia.org/wiki/Brier_score)
- UVA Library: [A Brief on Brier Scores](https://library.virginia.edu/data/articles/a-brief-on-brier-scores)
