What is an AI superforecaster? How AI puts honest odds on the future

An AI superforecaster researches a question about the future and gives a calibrated probability. How it works, how it's scored, and where it stands in 2026.

An AI superforecaster is an AI system built to answer questions about the future with a probability, and to be right about as often as that probability says. It does not guess or give a hot take. It researches the question, weighs the evidence, and says something like "34%". When the answer arrives, the forecast is scored.

This guide explains where the idea comes from, how these systems work, how they are judged, and how close they are to the best human forecasters in 2026.

Where the word “superforecaster” comes from

The term comes from human forecasting research. In 2011, IARPA, the research arm of the US intelligence community, launched a four-year forecasting tournament on geopolitical questions. A team led by Philip Tetlock and Barbara Mellers, the Good Judgment Project, won it. Its best performers, roughly the top 2% of participants, became known as superforecasters (Good Judgment).

Their edge was not secret information. It was method: starting from base rates, breaking questions into parts, updating often in small steps, and stating uncertainty as precise probabilities. According to Good Judgment, the team was over 30% more accurate than intelligence analysts with access to classified information.

An AI superforecaster tries to do the same thing with software: follow the method that made the best humans accurate, at a scale no human team can match.

What an AI superforecaster actually does

Most AI forecasting systems follow the same broad loop. It looks a lot like the way a careful human analyst works.

  1. Take a clear question. "Will China land astronauts on the Moon before 2030?" works. "Will China win the space race?" does not, because nobody can check the answer later.
  2. Start from the base rate. How often have similar things happened before? This anchors the forecast before any news is read. See base rates, the first step in every good forecast.
  3. Research the evidence. AI agents gather news, data and background that exist today, as short, dated facts with sources.
  4. Check every source. Claims that cannot be traced to a real source are thrown out before they can move the number.
  5. Weigh it, more than once. Several independent forecasts are produced and combined, because averaging independent views tends to cancel individual errors.
  6. Calibrate. The combined number is adjusted using the system's own track record, so that its 70% forecasts come true about 70% of the time. More on this in what a calibrated forecast is.
  7. Get scored. When the question resolves, the forecast is graded with a proper scoring rule such as the Brier score. The misses stay on the record.

Technically, the core is usually a large language model inside a "scaffold": a program that guides it through search, reading, reasoning and aggregation, instead of asking it for an instant answer.

How it is judged: calibration and accuracy

A forecast of "70%" can never be right or wrong on its own. If the event does not happen, that is exactly what a 70% forecast expects to see three times out of ten. So AI forecasters are judged across many questions, on two things:

  • Calibration: do its 70% calls come true about 70% of the time?
  • Accuracy: how close were its probabilities to what happened, measured with a score like the Brier score (lower is better, and always saying 50% scores 0.25)?

The strongest evidence comes from benchmarks that only ask about events that have not happened yet, so the answers cannot be in the model's training data. ForecastBench, run by the Forecasting Research Institute, works this way: it asks systems about future events every two weeks and scores them as the answers arrive.

Where AI forecasting stands in 2026

Progress has been fast, and the evidence is genuinely mixed at the top:

  • In July 2026, the Forecasting Research Institute reported that the best AI systems on ForecastBench are now statistically indistinguishable from superforecasters, with the leading system's accuracy not significantly different from theirs (FRI).
  • Good Judgment disputes the framing: the superforecaster numbers on that benchmark were collected in 2024 on different questions, and only a head-to-head test with the same questions, timing and rules would settle it (Good Judgment).
  • On Metaculus, the top human "Pro" forecasters have kept a lead over the best bots in head-to-head comparisons (Metaculus), even as a bot ranked 33rd of 1,130 people in the Spring 2026 Metaculus Cup (EA Forum summary).

The fair summary: AI forecasters have moved from near chance to close to the best humans in about two years, and whether they have fully caught up is still an open question. We compare the two in detail in AI vs. superforecasters.

AI superforecasters and prediction markets

The most natural place to test an AI forecaster is next to a prediction market. A market price is roughly the probability traders are willing to bet on. An AI forecast is the probability the checked evidence supports. When the two disagree, the gap is worth a closer look.

Markets are not a perfect benchmark either. They have known distortions, such as the favorite-longshot bias, where unlikely outcomes tend to be priced too high. Showing both numbers side by side is more informative than either one alone.

What an AI superforecaster is not

  • Not a crystal ball. It gives probabilities, not certainties. A 10% event will still happen one time in ten.
  • Not a pundit. Every forecast is a number that gets scored, not an opinion that is forgotten.
  • Not always confident. A good forecaster says when the evidence is thin, instead of guessing.
  • Not investment advice. A probability about an event is information, not a recommendation to trade.

How Sikt approaches it

Sikt Intelligence is an AI forecasting startup in Oslo, building an AI superforecaster around the loop above: evidence first, every source checked, calibrated against its own record, and scored every time. We are in the research stage, so we make no performance claims until there is a published track record. You can see the kind of questions we work on in the predictions feed, and read about us.

Key takeaways

  • An AI superforecaster is an AI system that researches a question about the future and gives a calibrated probability, then gets scored.
  • The name comes from the top 2% of human forecasters in a four-year intelligence-community tournament.
  • The method: base rates, evidence, source checks, independent forecasts, calibration, scoring.
  • In 2026 the best AI systems are close to superforecaster level on some benchmarks; whether they have fully caught up is still debated.
  • Judge any forecaster by calibration and a proper score across many questions, never by one call.

FAQ

Is an AI superforecaster better than a human superforecaster?

Not clearly, yet. On ForecastBench the best systems are now statistically indistinguishable from superforecasters, but that comparison uses different questions and older human data. Top human forecasters still lead in head-to-head comparisons on Metaculus.

How is an AI forecast scored?

With a proper scoring rule, most often the Brier score: the squared difference between the probability and the outcome (1 if it happened, 0 if not), averaged over many forecasts. Lower is better.

Can an AI superforecaster predict anything?

No. It works best on clear questions with a checkable answer and a fixed date. On questions with thin or contradictory evidence, a good system should say so instead of guessing.

Sources