AI forecasting benchmarks explained: ForecastBench, Prophet Arena and FutureEval
Which AI forecasts best? How ForecastBench, Prophet Arena and Metaculus FutureEval test AI on real future events, and what their 2026 results show.
Read the articleWhat AI forecasting is, how AI systems predict real-world events, how accurate they are in 2026, how they are scored, and where they still fail.

AI forecasting means using AI to put numbers on the future: how much, how many, or how likely. In 2026 it works remarkably well in some areas, such as the weather, and is close to the best humans in others, such as real-world events. It cannot see the future. What it can do is turn the evidence that exists today into a probability that can be checked when the answer arrives. This guide covers how it works, how accurate it is, how it is judged and where it still fails.
The term covers two quite different jobs:
Both are judged the same way in the end: by comparing forecasts with what actually happened.
Weather is where AI forecasting has made its most convincing progress. In 2023, Google DeepMind's GraphCast was more accurate than the leading conventional medium-range forecast on 90% of 1,380 test targets (Lam et al., Science). In 2024, its probabilistic successor, GenCast, was significantly more skilful than the world's top operational ensemble forecast on 97.2% of 1,320 targets (Price et al., Nature).
Weather is also where the most common scoring rule for probability forecasts comes from: the Brier score, designed in 1950 to grade weather forecasters.
Events are harder than weather: there is no physics to simulate, and each question is new. Most AI forecasting systems follow the same broad loop, much like a careful human analyst:
A system built this way is often called an AI superforecaster.
On real-world events, close to the best humans, and improving fast:
We compare the benchmarks in AI forecasting benchmarks explained and the human-versus-machine debate in AI vs. superforecasters.
A single forecast of 70% can never be right or wrong on its own. AI forecasters are judged across many questions:
The only test that really counts is forecasting events before they happen, in public, and then being scored. Sikt Intelligence, which builds an AI superforecaster, is doing that on the 2026 US midterms:
A live AI forecast, scored in public. In the Sikt Midterm Bench, Sikt forecasts the 2026 midterms next to Kalshi and Polymarket. Latest round, Oct 3, 2026: Senate: Democrats 56% (markets 64% to 66%), every round; House: Democrats 88% (markets 91% to 92%), every round. Every forecast will be graded after election day, misses included.
Using AI to estimate future quantities, such as demand or temperature, or the probability of future events, such as an election result. Event forecasts are probabilities that are checked against what actually happens.
In weather, AI models now beat the leading conventional forecasts on most targets. On real-world events in 2026, the best AI systems are statistically tied with superforecasters on ForecastBench and narrowly behind professional forecasters on Metaculus.
Not with certainty, but it can estimate the odds of real events close to the level of the best human forecasters. We cover the evidence in can AI predict the future?
With a proper scoring rule such as the Brier score across many questions, a calibration check, and a comparison with a baseline such as the market's price, using only questions that were in the future when the forecasts were made.