What is an AI superforecaster? How AI puts honest odds on the future
An AI superforecaster researches a question about the future and gives a calibrated probability. How it works, how it's scored, and where it stands in 2026.
Read the articleThe best AI is now statistically tied with superforecasters on ForecastBench, while top humans still lead on Metaculus. What each test shows, and what's fair.

The honest answer in 2026: it is close, and it depends on the test. On ForecastBench, the best AI systems are now statistically indistinguishable from superforecasters. On Metaculus, top human forecasters still beat the best bots head to head. Both results can be true at once, and the reasons why tell you a lot about how forecasting should be measured.
Superforecasters are the best human forecasters ever measured. They emerged from a forecasting tournament run from 2011 by IARPA, the research arm of the US intelligence community: four years, around 500 questions and over a million forecasts. The top 2% of participants in the winning team, the Good Judgment Project, were named superforecasters (Good Judgment).
They became the gold standard that every AI superforecaster is now measured against.
ForecastBench, run by the Forecasting Research Institute (FRI), asks AI systems about events that have not happened yet, every two weeks, then scores them as questions resolve. Because the answers do not exist when the forecasts are made, the test cannot be passed by memory.
In July 2026, FRI reported that several AI systems are now statistically indistinguishable from superforecasters, with the leading system's accuracy not significantly different from theirs (FRI). FRI also noted the caveats: the superforecaster forecasts were collected in 2024, and confidence intervals overlap for many systems.
In the Spring 2026 Metaculus Cup, a bot ranked 33rd of 1,130 human forecasters, around the top 3% (EA Forum summary).
Metaculus runs a stricter comparison: its top "Pro" forecasters and the best bots answer the same questions at the same time. In every season reported so far, the Pros have led the bots by a large margin (Metaculus).
Good Judgment, the company built on the original superforecasters, argues that the ForecastBench comparison is not settled. The superforecaster numbers are two years old and come from different questions. As its analysis puts it, comparing across question sets "is not the same as having humans and AI forecast the same questions, at the same time, under the same rules" (Good Judgment).
| ForecastBench | Metaculus head-to-head | |
|---|---|---|
| Same questions for humans and AI? | No: human baseline from 2024 | Yes |
| Same time window? | No | Yes |
| Humans compared | Superforecasters | Metaculus Pro forecasters |
| 2026 result | AI statistically tied | Pros ahead |
The difference is mostly about fairness of comparison, not about which side is lying. Forecasting accuracy depends heavily on which questions are asked and when. A benchmark that holds everything equal is harder for AI to win, and more convincing when it does.
The strongest comparison would put humans and AI on:
Until a test like that is run at scale, "AI has surpassed superforecasters" is a hypothesis, not a fact.
The most practical finding may be that the choice is not binary. In a study of 991 people, forecasters who could consult an AI assistant were 24–28% more accurate than a control group (Schoenegger et al.). AI is already a strong forecasting partner, even where it is not yet the best forecaster.
That last point is how Sikt Intelligence approaches it: every forecast is a probability, every probability gets scored, and performance claims come only from a published track record.
Not conclusively. AI systems are statistically tied with superforecasters on ForecastBench, but that comparison uses different questions and older human data. In same-question comparisons on Metaculus, top human forecasters still lead.
A benchmark run by the Forecasting Research Institute that asks AI systems about future events every two weeks and scores them when the questions resolve, so the answers cannot come from training data.
Many researchers expect it, based on the current rate of improvement. Whether and when it happens on a fair, same-question test is still open.