Experiment · FIFA World Cup 2026

Can AI predict the World Cup?

We asked 25 AI models from 11 companies to forecast every match of FIFA World Cup 2026 — then measured how close they got. A scientific experiment by LumenIA.

AI predicting the FIFA World Cup 2026
Models from
AnthropicAnthropic
OpenAIOpenAI
GoogleGoogle
MoonshotMoonshot
MetaMeta
MistralMistral
AlibabaAlibaba
DeepSeekDeepSeek
xAIxAI
PerplexityPerplexity
z.aiz.ai
01
25
AI models
02
11
Companies
03
104
Matches
04
7
Experiments

World Cup 2026 — Final results

The tournament has ended. Here is how reality unfolded — the benchmark against which every model is judged.

Champion
1st
🇪🇸Spain
Runner-up
2nd
🇦🇷Argentina
Third place
3rd
🏴󠁧󠁢󠁥󠁮󠁧󠁿England
MVP
MVP
🇪🇸Rodri
Top scorer
10 goals
🇫🇷Kylian Mbappe
Top assists
7 assists
🇫🇷Michael Olise
25 models · 11 companies

The competitors

Twenty-five frontier models from eleven companies — all running on the same prompts, the same scoring rules, the same 104 matches.

Anthropic
Claude Opus 4.8
Anthropic🇺🇸
Intelligence
Cost
Anthropic
Claude Sonnet 4.6
Anthropic🇺🇸
Intelligence
Cost
Anthropic
Claude Haiku 4.5
Anthropic🇺🇸
Intelligence
Cost
OpenAI
GPT-5.5
OpenAI🇺🇸
Intelligence
Cost
OpenAI
GPT-5 Mini
OpenAI🇺🇸
Intelligence
Cost
OpenAI
GPT-5 Nano
OpenAI🇺🇸
Intelligence
Cost
Google
Gemini 2.5 Pro
Google🇺🇸
Intelligence
Cost
Google
Gemini 2.5 Flash
Google🇺🇸
Intelligence
Cost
Google
Gemini 2.5 Flash Lite
Google🇺🇸
Intelligence
Cost
Moonshot
Moonshot Kimi K2.6
Moonshot🇨🇳
Intelligence
Cost
Meta
Llama 4 Maverick
Meta🇺🇸
Intelligence
Cost
Meta
Llama 3.3 70B
Meta🇺🇸
Intelligence
Cost
Mistral
Mistral Large
Mistral🇪🇺
Intelligence
Cost
Mistral
Mixtral 8x22B
Mistral🇪🇺
Intelligence
Cost
Alibaba
Qwen 3.7 Max
Alibaba🇨🇳
Intelligence
Cost
Alibaba
Qwen 3.7 Plus
Alibaba🇨🇳
Intelligence
Cost
DeepSeek
DeepSeek V4 Pro
DeepSeek🇨🇳
Intelligence
Cost
DeepSeek
DeepSeek R1 0528
DeepSeek🇨🇳
Intelligence
Cost
xAI
Grok 4.20
xAI🇺🇸
Intelligence
Cost
Perplexity
Perplexity Sonar
Perplexity🇺🇸
Intelligence
Cost
Anthropic
Claude Fable 5
Anthropic🇺🇸
Intelligence
Cost
OpenAI
GPT-5.6 Sol
OpenAI🇺🇸
Intelligence
Cost
OpenAI
GPT-5.6 Terra
OpenAI🇺🇸
Intelligence
Cost
OpenAI
GPT-5.6 Luna
OpenAI🇺🇸
Intelligence
Cost
z.ai
GLM 5.2
z.ai🇨🇳
Intelligence
Cost
01 — 05

Our experiments

Each protocol stresses the models in a different way — from one-shot foresight to round-by-round adaptation.

01oracle

Oracle

One shot. Each model predicts the entire bracket before the opening match — group stage to final.

View full results
Top modelstop 5 / 20
ModelAccuracyChampion
49.8% Yes
47.8% Yes
47.8% Yes
47.0% No
45.8% No
02progressive

Progressive

Round by round. Models re-forecast after every stage, using completed matches as new context.

View full results
Top modelstop 5 / 20
Best model74.9%
Meta
Llama 3.3 70B
Accuracy
Worst model69.2%
Mistral
Mixtral 8x22B
Accuracy
ModelAccuracy
1MetaLlama 3.3 70B
74.9%
2AnthropicClaude Opus 4.8
74.1%
3OpenAIGPT-5.5
74.1%
4AnthropicClaude Haiku 4.5
73.6%
5OpenAIGPT-5 Mini
73.5%
03betting

Betting

Skin in the game. Each model starts with $10,000 and stakes at every stage. Bankroll is the score.

Top modelstop 5 / 20
Best model$30,527
OpenAI
GPT-5 Mini
Bankroll
Worst model$3,228
DeepSeek
DeepSeek R1 0528
Bankroll
ModelBankrollReturn
1OpenAIGPT-5 Mini
$30,527+24.0%
2GoogleGemini 2.5 Pro
$29,775+24.2%
3MetaLlama 3.3 70B
$17,813+14.1%
4OpenAIGPT-5.5
$17,213+10.5%
5xAIGrok 4.20
$13,422+7.8%
04gpt-vs-gemini

GPT vs Gemini

Head to head. We run the Oracle protocol 100 times on each model and compare who calls the tournament more precisely.

Head to head · 100 runs each2 × 100 runs
OpenAI
GPT-5 Nano
OpenAI
vs
Google
Gemini 2.5 Flash Lite
Google
Overall accuracy
21.2%
21.2%
Exact score predictions
8 / 100
8 / 100
Correct winner / draw
39 / 100
39 / 100
GPT-5 10 Gemini(wins)
05final-test

Final Test

Same match, ten tries. 25 models each simulate the WC26 Final 10 times so we can measure how much a single model disagrees with itself — variance in scorelines, scorers, assists and MVPs. Run twice: with and without the tournament match history in the prompt.

Top models25 × 10 runs
ModelModal scorexGAccuracy
Google
Gemini 2.5 Pro
Google
211.600.5093.9%
OpenAI
GPT-5.6 Sol
OpenAI
101.500.5092.3%
Anthropic
Claude Opus 4.8
Anthropic
212.001.0089.1%
Anthropic
Claude Fable 5
Anthropic
212.001.0089.1%
OpenAI
GPT-5.6 Luna
OpenAI
212.001.0089.1%
Real final: 🇪🇸 10 🇦🇷 · 10 runs per model
LumenIA · 2026

About the experiment

LumenIA designed and ran this benchmark to study how current frontier models reason about uncertain, real-world events. All prompts, outputs and scoring code are open and reproducible.

Visit lumenia.net