# AI vs. superforecasters: who predicts better in 2026?

> The best AI is now statistically tied with superforecasters on ForecastBench, while top humans still lead on Metaculus. What each test shows, and what's fair.

Sep 28, 2026 · AI forecasting · Sikt Intelligence · https://www.siktintelligence.com/blog/ai-vs-superforecasters

**The honest answer in 2026: it is close, and it depends on the test.** On ForecastBench, the best AI systems are now statistically indistinguishable from superforecasters. On Metaculus, top human forecasters still beat the best bots head to head. Both results can be true at once, and the reasons why tell you a lot about how forecasting should be measured.

## Who are the superforecasters?

Superforecasters are the best human forecasters ever measured. They emerged from a forecasting tournament run from 2011 by IARPA, the research arm of the US intelligence community: four years, around 500 questions and over a million forecasts. The top 2% of participants in the winning team, the Good Judgment Project, were named superforecasters ([Good Judgment](https://goodjudgment.com/resources/the-superforecasters-track-record/)).

They became the gold standard that every [AI superforecaster](/blog/what-is-an-ai-superforecaster) is now measured against.

## The case that AI has caught up

### ForecastBench: statistically tied

[ForecastBench](https://www.forecastbench.org/about/), run by the Forecasting Research Institute (FRI), asks AI systems about events that have not happened yet, every two weeks, then scores them as questions resolve. Because the answers do not exist when the forecasts are made, the test cannot be passed by memory.

In July 2026, FRI reported that several AI systems are now **statistically indistinguishable from superforecasters**, with the leading system's accuracy not significantly different from theirs ([FRI](https://forecastingresearch.substack.com/p/ai-models-have-likely-reached-parity)). FRI also noted the caveats: the superforecaster forecasts were collected in 2024, and confidence intervals overlap for many systems.

### Tournaments: bots among the top humans

In the Spring 2026 Metaculus Cup, a bot ranked **33rd of 1,130 human forecasters**, around the top 3% ([EA Forum summary](https://forum.effectivealtruism.org/posts/Spyz3wESZu2eeqhDj/ai-forecasting-in-2026-what-11-analyses-say)).

## The case that humans still lead

### Metaculus head-to-head: Pros ahead

Metaculus runs a stricter comparison: its top "Pro" forecasters and the best bots answer the **same questions at the same time**. In every season reported so far, the Pros have led the bots by a large margin ([Metaculus](https://metaculus.substack.com/p/metaculus-futureeval-ai-forecasting-benchmark)).

### Good Judgment: not a fair comparison yet

Good Judgment, the company built on the original superforecasters, argues that the ForecastBench comparison is not settled. The superforecaster numbers are two years old and come from different questions. As its analysis puts it, comparing across question sets "is not the same as having humans and AI forecast the same questions, at the same time, under the same rules" ([Good Judgment](https://goodjudgment.substack.com/p/the-hype-and-the-evidence-about-ai)).

## Why the benchmarks disagree

| | ForecastBench | Metaculus head-to-head |
|---|---|---|
| Same questions for humans and AI? | No: human baseline from 2024 | Yes |
| Same time window? | No | Yes |
| Humans compared | Superforecasters | Metaculus Pro forecasters |
| 2026 result | AI statistically tied | Pros ahead |

The difference is mostly about **fairness of comparison**, not about which side is lying. Forecasting accuracy depends heavily on which questions are asked and when. A benchmark that holds everything equal is harder for AI to win, and more convincing when it does.

## What a fair test would look like

The strongest comparison would put humans and AI on:

1. **The same questions**, chosen before anyone forecasts.
2. **The same time window**, with the same information available.
3. **The same rules** for updating forecasts.
4. **The same scoring**, with a proper rule such as the [Brier score](/blog/brier-score-explained) and a check of [calibration](/blog/forecast-calibration).
5. **Enough questions** that luck cannot decide the result.

Until a test like that is run at scale, "AI has surpassed superforecasters" is a hypothesis, not a fact.

## Humans + AI beats either alone

The most practical finding may be that the choice is not binary. In a study of 991 people, forecasters who could consult an AI assistant were **24–28% more accurate** than a control group ([Schoenegger et al.](https://arxiv.org/abs/2402.07862)). AI is already a strong forecasting partner, even where it is not yet the best forecaster.

## What this means for AI forecasting

- **Progress is real and fast.** AI went from near chance to near the top in about two years (see [can AI predict the future?](/blog/can-ai-predict-the-future)).
- **The last step is the hardest.** Matching the very best humans on equal terms is not yet proven.
- **Markets are a benchmark too.** Some benchmarks draw questions from [prediction markets](/blog/how-prediction-markets-work) and compare AI with market prices, which have their own biases.
- **Track records decide.** The only claim that should convince anyone is a published record of scored forecasts on real, future questions, misses included.

That last point is how Sikt Intelligence approaches it: every forecast is a probability, every probability gets scored, and performance claims come only from a published track record.

## Key takeaways

- On ForecastBench, the best AI systems are now statistically tied with superforecasters.
- On Metaculus, top human Pros still lead the best bots on the same questions.
- The benchmarks disagree mainly because only one of them compares humans and AI on identical questions at the same time.
- A fair test needs the same questions, timing, rules and proper scoring.
- Humans with AI assistants forecast better than either alone.

## FAQ

### Has AI beaten superforecasters?

Not conclusively. AI systems are statistically tied with superforecasters on ForecastBench, but that comparison uses different questions and older human data. In same-question comparisons on Metaculus, top human forecasters still lead.

### What is ForecastBench?

A benchmark run by the Forecasting Research Institute that asks AI systems about future events every two weeks and scores them when the questions resolve, so the answers cannot come from training data.

### Will AI forecasters overtake humans?

Many researchers expect it, based on the current rate of improvement. Whether and when it happens on a fair, same-question test is still open.

## Sources

- Good Judgment: [The Superforecasters’ Track Record](https://goodjudgment.com/resources/the-superforecasters-track-record/)
- ForecastBench: [About](https://www.forecastbench.org/about/)
- Forecasting Research Institute: [AI models have likely reached parity with superforecasters on ForecastBench](https://forecastingresearch.substack.com/p/ai-models-have-likely-reached-parity) (July 2026)
- Good Judgment: [The Hype and the Evidence about AI Forecasting](https://goodjudgment.substack.com/p/the-hype-and-the-evidence-about-ai) (July 2026)
- Metaculus: [FutureEval, Metaculus’s AI forecasting benchmark](https://metaculus.substack.com/p/metaculus-futureeval-ai-forecasting-benchmark)
- EA Forum: [AI Forecasting in 2026: What 11 Analyses Say](https://forum.effectivealtruism.org/posts/Spyz3wESZu2eeqhDj/ai-forecasting-in-2026-what-11-analyses-say)
- Schoenegger et al.: [AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy](https://arxiv.org/abs/2402.07862)
