Can AI predict the future? What the evidence says in 2026
AI can't see the future, but it now estimates the odds of real events close to the best human forecasters. What the 2026 research shows, and its limits.
Read the articleAn AI superforecaster researches a question about the future and gives a calibrated probability. How it works, how it's scored, and where it stands in 2026.

An AI superforecaster is an AI system built to answer questions about the future with a probability, and to be right about as often as that probability says. It does not guess or give a hot take. It researches the question, weighs the evidence, and says something like "34%". When the answer arrives, the forecast is scored.
This guide explains where the idea comes from, how these systems work, how they are judged, and how close they are to the best human forecasters in 2026.
The term comes from human forecasting research. In 2011, IARPA, the research arm of the US intelligence community, launched a four-year forecasting tournament on geopolitical questions. A team led by Philip Tetlock and Barbara Mellers, the Good Judgment Project, won it. Its best performers, roughly the top 2% of participants, became known as superforecasters (Good Judgment).
Their edge was not secret information. It was method: starting from base rates, breaking questions into parts, updating often in small steps, and stating uncertainty as precise probabilities. According to Good Judgment, the team was over 30% more accurate than intelligence analysts with access to classified information.
An AI superforecaster tries to do the same thing with software: follow the method that made the best humans accurate, at a scale no human team can match.
Most AI forecasting systems follow the same broad loop. It looks a lot like the way a careful human analyst works.
Technically, the core is usually a large language model inside a "scaffold": a program that guides it through search, reading, reasoning and aggregation, instead of asking it for an instant answer.
A forecast of "70%" can never be right or wrong on its own. If the event does not happen, that is exactly what a 70% forecast expects to see three times out of ten. So AI forecasters are judged across many questions, on two things:
The strongest evidence comes from benchmarks that only ask about events that have not happened yet, so the answers cannot be in the model's training data. ForecastBench, run by the Forecasting Research Institute, works this way: it asks systems about future events every two weeks and scores them as the answers arrive.
Progress has been fast, and the evidence is genuinely mixed at the top:
The fair summary: AI forecasters have moved from near chance to close to the best humans in about two years, and whether they have fully caught up is still an open question. We compare the two in detail in AI vs. superforecasters.
The most natural place to test an AI forecaster is next to a prediction market. A market price is roughly the probability traders are willing to bet on. An AI forecast is the probability the checked evidence supports. When the two disagree, the gap is worth a closer look.
Markets are not a perfect benchmark either. They have known distortions, such as the favorite-longshot bias, where unlikely outcomes tend to be priced too high. Showing both numbers side by side is more informative than either one alone.
Sikt Intelligence is an AI forecasting startup in Oslo, building an AI superforecaster around the loop above: evidence first, every source checked, calibrated against its own record, and scored every time. We are in the research stage, so we make no performance claims until there is a published track record. You can see the kind of questions we work on in the predictions feed, and read about us.
Not clearly, yet. On ForecastBench the best systems are now statistically indistinguishable from superforecasters, but that comparison uses different questions and older human data. Top human forecasters still lead in head-to-head comparisons on Metaculus.
With a proper scoring rule, most often the Brier score: the squared difference between the probability and the outcome (1 if it happened, 0 if not), averaged over many forecasts. Lower is better.
No. It works best on clear questions with a checkable answer and a fixed date. On questions with thin or contradictory evidence, a good system should say so instead of guessing.