# AI forecasting benchmarks explained: ForecastBench, Prophet Arena and FutureEval

> How the three main AI forecasting benchmarks work, what they measure and what they show in 2026: ForecastBench, Prophet Arena and Metaculus FutureEval.

Oct 2, 2026 · AI forecasting · Sikt Intelligence · https://www.siktintelligence.com/blog/ai-forecasting-benchmarks

**Three benchmarks decide most claims about how well AI forecasts: ForecastBench, Prophet Arena and Metaculus FutureEval.** All three share one idea that makes them hard to fool: the AI must forecast events *before* they happen, and is scored once reality answers. They differ in where their questions come from, who the AI is compared with, and how the score is counted. Here is how each one works, what they show in 2026, and how to read their results.

## Why AI forecasting needs special benchmarks

Most AI benchmarks test questions with known answers. For forecasting, that does not work: a model trained on data from after an event can simply remember the result. This is **lookahead bias**, and it makes many impressive-looking backtests worthless (we show how in [can AI predict the stock market?](/blog/can-ai-predict-the-stock-market)).

The fix is to ask about the future. Every benchmark below poses questions whose answers do not exist yet, collects probabilities, and scores them after the events resolve, usually with the [Brier score](/blog/brier-score-explained) and a check of [calibration](/blog/forecast-calibration).

## ForecastBench

**Run by:** the Forecasting Research Institute, a nonprofit research group.

**How it works:** every two weeks, ForecastBench generates a new round of 500 questions: 250 drawn from prediction markets (Kalshi, Metaculus and Polymarket) and 250 generated automatically from real-world data series (ACLED for conflict, DBnomics and FRED for economic data). AI systems forecast them, and the answers are scored with the Brier score, converted into an easier-to-read Brier Index, with market and dataset questions weighted equally. Because every question is about the future, the benchmark is, in its own words, "contamination-free" ([ForecastBench](https://www.forecastbench.org/about/)).

**What it shows:** in July 2026, the institute reported that several AI systems were "statistically indistinguishable from superforecaster-level accuracy." It also added important caveats: the results are "more consistent with superforecaster parity than with outperformance," and the comparison relies on superforecaster forecasts last collected in 2024 ([Forecasting Research Institute](https://forecastingresearch.substack.com/p/ai-models-have-likely-reached-parity)).

## Prophet Arena

**What it is:** a live leaderboard that ranks AI models and agents on real prediction markets. "Every day, frontier models assign probabilities to newly opened prediction markets," across sports, elections, prices and science, and the standings update as events resolve ([Prophet Arena](https://www.prophetarena.co/)).

**How it scores:** with Brier scores and calibration, and directly against the market's own price, so the question is not only "was the AI accurate?" but "was it more accurate than the traders?"

**What it shows:** the research paper behind the benchmark found that AI models had "small calibration errors, consistent prediction confidence and promising market returns," but also weaknesses: inaccurate recall of events, misunderstanding of data sources, and "slower information aggregation compared to markets when resolution nears" ([Yang et al., 2025](https://arxiv.org/abs/2510.17638)). In other words, markets still react faster to breaking news at the end.

## Metaculus FutureEval

**Run by:** Metaculus, the forecasting platform.

**How it works:** FutureEval has three parts ([Metaculus](https://metaculus.substack.com/p/metaculus-futureeval-ai-forecasting-benchmark)):

- a **model leaderboard** that runs major AI models with a fixed prompt on most open Metaculus questions
- **bot tournaments**, where developers enter their own forecasting bots for a share of $175,000 in yearly prizes, including a $50,000 tournament every four months
- a **human baseline** of hand-picked Metaculus Pro Forecasters, who forecast a subset of the same questions

**What it shows:** for most of its history, the pros led the bots "every season by a large margin." But the gap is closing. In the spring 2026 tournament (January to April), ten pros beat the top ten bots by just 1.25 points per question, a difference too small to be statistically significant ([Metaculus, September 2026](https://www.lesswrong.com/posts/wZBbDqzfBjYG58CxK/futureeval-spring-results-pros-beat-bots-but-the-gap-is)).

## The three side by side

| | ForecastBench | Prophet Arena | Metaculus FutureEval |
|---|---|---|---|
| Run by | Forecasting Research Institute | Academic researchers | Metaculus |
| Questions from | Prediction markets and real-world data series | Live prediction markets | Metaculus questions |
| Compared with | Superforecasters (2024 data) | The market price | Metaculus Pro Forecasters |
| Scoring | Brier score, Brier Index | Brier score, calibration, accuracy vs. the market | Head-to-head scores vs. pros |
| Cadence | 500 new questions every two weeks | Daily | Seasonal tournaments, plus a continuous leaderboard |

## What the benchmarks agree on

- **AI is close to strong human forecasters, but not clearly ahead.** ForecastBench suggests parity with superforecasters; Metaculus shows a lead for pros that has shrunk to statistical noise.
- **Markets are still fast.** Near resolution, when news arrives quickly, markets absorb it faster than AI systems do.
- **The baseline matters.** Beating a market, beating superforecasters and beating a crowd are different achievements, so always check who the AI was compared with. See [AI vs. superforecasters](/blog/ai-vs-superforecasters) for the full debate.

## How to read an AI forecasting result

1. **Which benchmark, and which leaderboard?** Results differ by question type and by human baseline.
2. **Were the forecasts made before the outcomes were known?** If not, ignore the result.
3. **Is the difference statistically significant?** Many top results sit within each other's margins of error.
4. **Who was the comparison?** A market, a crowd and superforecasters set very different bars.

## Where Sikt Intelligence fits

Sikt Intelligence is building an [AI superforecaster](/blog/what-is-an-ai-superforecaster), and we hold it to the same three standards these benchmarks use:

- **Only the future counts.** Sikt is judged only on questions whose outcomes the AI could not have known, the same rule that makes these benchmarks trustworthy.
- **Proper scores, misses included.** Every forecast is scored when the answer arrives, with the Brier score and a calibration check.
- **The market as the baseline.** Every Sikt forecast sits next to the prediction-market price, like Prophet Arena's comparison, because the useful question is not just "is the AI accurate?" but "does the evidence support the market's price?"

Benchmarks tell you which system is best on average. Sikt is built to answer the question you actually face: on *this* event, is the market right? Our first public test is the **Sikt Midterm Bench**, where Sikt is forecasting the key 2026 Senate races ahead of election day; the forecasts will be published on our [2026 midterms page](/midterms), next to Kalshi and Polymarket, and scored against the results.

## Key takeaways

- ForecastBench, Prophet Arena and Metaculus FutureEval all score AI forecasts made before events happen, which rules out memorized answers.
- ForecastBench reports AI at roughly superforecaster level, with caveats about its 2024 human data.
- Prophet Arena compares AI with live market prices and finds markets still faster near resolution.
- Metaculus pros still lead the best bots, but by a margin that is no longer statistically significant.
- Always check the baseline, the timing and the margin of error before trusting an AI forecasting claim.

## FAQ

### What is ForecastBench?

A benchmark from the Forecasting Research Institute that asks AI systems 500 new forecasting questions every two weeks, from prediction markets and real-world data, and scores them with the Brier score once they resolve.

### What is Prophet Arena?

A live leaderboard that ranks AI models on real prediction markets, scoring them with Brier scores, calibration and accuracy against the market price.

### Are AI forecasters better than humans?

Not clearly. On ForecastBench, top AI systems are statistically tied with superforecasters; on Metaculus, pro forecasters still lead the best bots, though by a margin that is no longer statistically significant.

## Sources

- ForecastBench: [About ForecastBench](https://www.forecastbench.org/about/)
- Forecasting Research Institute: [AI models have likely reached parity with superforecasters](https://forecastingresearch.substack.com/p/ai-models-have-likely-reached-parity) (July 2026)
- Prophet Arena: [AI forecasting benchmark on real prediction markets](https://www.prophetarena.co/)
- Yang et al.: [LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena](https://arxiv.org/abs/2510.17638) (arXiv, 2025)
- Metaculus: [FutureEval, Metaculus's AI forecasting benchmark](https://metaculus.substack.com/p/metaculus-futureeval-ai-forecasting-benchmark)
- Metaculus: [FutureEval spring 2026 tournament results](https://www.lesswrong.com/posts/wZBbDqzfBjYG58CxK/futureeval-spring-results-pros-beat-bots-but-the-gap-is) (September 2026)
