AI forecasting benchmarks explained: ForecastBench, Prophet Arena and FutureEval

How the three main AI forecasting benchmarks work, what they measure and what they show in 2026: ForecastBench, Prophet Arena and Metaculus FutureEval.

Three benchmarks decide most claims about how well AI forecasts: ForecastBench, Prophet Arena and Metaculus FutureEval. All three share one idea that makes them hard to fool: the AI must forecast events before they happen, and is scored once reality answers. They differ in where their questions come from, who the AI is compared with, and how the score is counted. Here is how each one works, what they show in 2026, and how to read their results.

Why AI forecasting needs special benchmarks

Most AI benchmarks test questions with known answers. For forecasting, that does not work: a model trained on data from after an event can simply remember the result. This is lookahead bias, and it makes many impressive-looking backtests worthless (we show how in can AI predict the stock market?).

The fix is to ask about the future. Every benchmark below poses questions whose answers do not exist yet, collects probabilities, and scores them after the events resolve, usually with the Brier score and a check of calibration.

ForecastBench

Run by: the Forecasting Research Institute, a nonprofit research group.

How it works: every two weeks, ForecastBench generates a new round of 500 questions: 250 drawn from prediction markets (Kalshi, Metaculus and Polymarket) and 250 generated automatically from real-world data series (ACLED for conflict, DBnomics and FRED for economic data). AI systems forecast them, and the answers are scored with the Brier score, converted into an easier-to-read Brier Index, with market and dataset questions weighted equally. Because every question is about the future, the benchmark is, in its own words, "contamination-free" (ForecastBench).

What it shows: in July 2026, the institute reported that several AI systems were "statistically indistinguishable from superforecaster-level accuracy." It also added important caveats: the results are "more consistent with superforecaster parity than with outperformance," and the comparison relies on superforecaster forecasts last collected in 2024 (Forecasting Research Institute).

Prophet Arena

What it is: a live leaderboard that ranks AI models and agents on real prediction markets. "Every day, frontier models assign probabilities to newly opened prediction markets," across sports, elections, prices and science, and the standings update as events resolve (Prophet Arena).

How it scores: with Brier scores and calibration, and directly against the market's own price, so the question is not only "was the AI accurate?" but "was it more accurate than the traders?"

What it shows: the research paper behind the benchmark found that AI models had "small calibration errors, consistent prediction confidence and promising market returns," but also weaknesses: inaccurate recall of events, misunderstanding of data sources, and "slower information aggregation compared to markets when resolution nears" (Yang et al., 2025). In other words, markets still react faster to breaking news at the end.

Metaculus FutureEval

Run by: Metaculus, the forecasting platform.

How it works: FutureEval has three parts (Metaculus):

  • a model leaderboard that runs major AI models with a fixed prompt on most open Metaculus questions
  • bot tournaments, where developers enter their own forecasting bots for a share of $175,000 in yearly prizes, including a $50,000 tournament every four months
  • a human baseline of hand-picked Metaculus Pro Forecasters, who forecast a subset of the same questions

What it shows: for most of its history, the pros led the bots "every season by a large margin." But the gap is closing. In the spring 2026 tournament (January to April), ten pros beat the top ten bots by just 1.25 points per question, a difference too small to be statistically significant (Metaculus, September 2026).

The three side by side

ForecastBenchProphet ArenaMetaculus FutureEval
Run byForecasting Research InstituteAcademic researchersMetaculus
Questions fromPrediction markets and real-world data seriesLive prediction marketsMetaculus questions
Compared withSuperforecasters (2024 data)The market priceMetaculus Pro Forecasters
ScoringBrier score, Brier IndexBrier score, calibration, accuracy vs. the marketHead-to-head scores vs. pros
Cadence500 new questions every two weeksDailySeasonal tournaments, plus a continuous leaderboard

What the benchmarks agree on

  • AI is close to strong human forecasters, but not clearly ahead. ForecastBench suggests parity with superforecasters; Metaculus shows a lead for pros that has shrunk to statistical noise.
  • Markets are still fast. Near resolution, when news arrives quickly, markets absorb it faster than AI systems do.
  • The baseline matters. Beating a market, beating superforecasters and beating a crowd are different achievements, so always check who the AI was compared with. See AI vs. superforecasters for the full debate.

How to read an AI forecasting result

  1. Which benchmark, and which leaderboard? Results differ by question type and by human baseline.
  2. Were the forecasts made before the outcomes were known? If not, ignore the result.
  3. Is the difference statistically significant? Many top results sit within each other's margins of error.
  4. Who was the comparison? A market, a crowd and superforecasters set very different bars.

Where Sikt Intelligence fits

Sikt Intelligence is building an AI superforecaster, and we hold it to the same three standards these benchmarks use:

  • Only the future counts. Sikt is judged only on questions whose outcomes the AI could not have known, the same rule that makes these benchmarks trustworthy.
  • Proper scores, misses included. Every forecast is scored when the answer arrives, with the Brier score and a calibration check.
  • The market as the baseline. Every Sikt forecast sits next to the prediction-market price, like Prophet Arena's comparison, because the useful question is not just "is the AI accurate?" but "does the evidence support the market's price?"

Benchmarks tell you which system is best on average. Sikt is built to answer the question you actually face: on this event, is the market right? Our first public test is the Sikt Midterm Bench, where Sikt is forecasting the key 2026 Senate races ahead of election day; the forecasts will be published on our 2026 midterms page, next to Kalshi and Polymarket, and scored against the results.

Key takeaways

  • ForecastBench, Prophet Arena and Metaculus FutureEval all score AI forecasts made before events happen, which rules out memorized answers.
  • ForecastBench reports AI at roughly superforecaster level, with caveats about its 2024 human data.
  • Prophet Arena compares AI with live market prices and finds markets still faster near resolution.
  • Metaculus pros still lead the best bots, but by a margin that is no longer statistically significant.
  • Always check the baseline, the timing and the margin of error before trusting an AI forecasting claim.

FAQ

What is ForecastBench?

A benchmark from the Forecasting Research Institute that asks AI systems 500 new forecasting questions every two weeks, from prediction markets and real-world data, and scores them with the Brier score once they resolve.

What is Prophet Arena?

A live leaderboard that ranks AI models on real prediction markets, scoring them with Brier scores, calibration and accuracy against the market price.

Are AI forecasters better than humans?

Not clearly. On ForecastBench, top AI systems are statistically tied with superforecasters; on Metaculus, pro forecasters still lead the best bots, though by a margin that is no longer statistically significant.

Sources