Can AI predict the World Cup?
We asked 25 AI models from 11 companies to forecast every match of FIFA World Cup 2026 — then measured how close they got. A scientific experiment by LumenIA.

World Cup 2026 — Final results
The tournament has ended. Here is how reality unfolded — the benchmark against which every model is judged.
The competitors
Twenty-five frontier models from eleven companies — all running on the same prompts, the same scoring rules, the same 104 matches.
Our experiments
Each protocol stresses the models in a different way — from one-shot foresight to round-by-round adaptation.
Oracle
One shot. Each model predicts the entire bracket before the opening match — group stage to final.
View full results| Model | Accuracy | Champion |
|---|---|---|
| 38.0% | ✕ No | |
| 38.0% | ✕ No | |
| 37.5% | ✕ No | |
| 35.1% | ✕ No | |
| 35.1% | ✕ No |
Progressive
Round by round. Models re-forecast after every stage, using completed matches as new context.
View full results| Model | Accuracy |
|---|---|
1 | 41.8% |
2 | 41.8% |
3 | 41.3% |
4 | 40.9% |
5 | 40.9% |
Betting
Skin in the game. Each model starts with $10,000 and stakes at every stage. Bankroll is the score.
| Model | Bankroll | Return |
|---|---|---|
1 | $30,527 | +205.3% |
2 | $29,775 | +197.7% |
3 | $17,813 | +78.1% |
4 | $17,213 | +72.1% |
5 | $13,422 | +34.2% |
GPT vs Gemini
Head to head. We run the Oracle protocol 100 times on each model and compare who calls the tournament more precisely.
Final Test
Same match, ten tries. 25 models each simulate the WC26 Final 10 times so we can measure how much a single model disagrees with itself — variance in scorelines, scorers, assists and MVPs. Run twice: with and without the tournament match history in the prompt.
| Model | Modal score | xG | Accuracy |
|---|---|---|---|
GPT-5.6 Sol OpenAI | 1–0 | 1.50–0.50 | 75.0% |
Gemini 2.5 Pro Google | 2–1 | 1.60–0.50 | 70.0% |
GPT-5.6 Terra OpenAI | 2–1 | 1.50–0.70 | 55.0% |
Claude Opus 4.8 Anthropic | 2–1 | 2.00–1.00 | 50.0% |
Claude Fable 5 Anthropic | 2–1 | 2.00–1.00 | 50.0% |
About the experiment
LumenIA designed and ran this benchmark to study how current frontier models reason about uncertain, real-world events. All prompts, outputs and scoring code are open and reproducible.
Visit lumenia.net