AIWC26 · LumenIA
Experiment · FIFA World Cup 2026

Can AI predict the World Cup?

We asked 25 AI models from 11 companies to forecast every match of FIFA World Cup 2026 — then measured how close they got. A scientific experiment by LumenIA.

AI predicting the FIFA World Cup 2026
Models from
AnthropicAnthropic
OpenAIOpenAI
GoogleGoogle
MoonshotMoonshot
MetaMeta
MistralMistral
AlibabaAlibaba
DeepSeekDeepSeek
xAIxAI
PerplexityPerplexity
z.aiz.ai
01
25
AI models
02
11
Companies
03
104
Matches
04
7
Experiments

World Cup 2026 — Final results

The tournament has ended. Here is how reality unfolded — the benchmark against which every model is judged.

Champion
1st
🇪🇸Spain
Runner-up
2nd
🇦🇷Argentina
Third place
3rd
🏴󠁧󠁢󠁥󠁮󠁧󠁿England
MVP
MVP
🇪🇸Rodri
Top scorer
10 goals
🇫🇷Kylian Mbappe
Top assists
7 assists
🇫🇷Michael Olise
25 models · 11 companies

The competitors

Twenty-five frontier models from eleven companies — all running on the same prompts, the same scoring rules, the same 104 matches.

Anthropic
Claude Opus 4.8
Anthropic🇺🇸
Intelligence
Cost
Anthropic
Claude Sonnet 4.6
Anthropic🇺🇸
Intelligence
Cost
Anthropic
Claude Haiku 4.5
Anthropic🇺🇸
Intelligence
Cost
OpenAI
GPT-5.5
OpenAI🇺🇸
Intelligence
Cost
OpenAI
GPT-5 Mini
OpenAI🇺🇸
Intelligence
Cost
OpenAI
GPT-5 Nano
OpenAI🇺🇸
Intelligence
Cost
Google
Gemini 2.5 Pro
Google🇺🇸
Intelligence
Cost
Google
Gemini 2.5 Flash
Google🇺🇸
Intelligence
Cost
Google
Gemini 2.5 Flash Lite
Google🇺🇸
Intelligence
Cost
Moonshot
Moonshot Kimi K2
Moonshot🇨🇳
Intelligence
Cost
Meta
Llama 4 Maverick
Meta🇺🇸
Intelligence
Cost
Meta
Llama 3.3 70B
Meta🇺🇸
Intelligence
Cost
Mistral
Mistral Large
Mistral🇪🇺
Intelligence
Cost
Mistral
Mixtral 8x22B
Mistral🇪🇺
Intelligence
Cost
Alibaba
Qwen 3.7 Max
Alibaba🇨🇳
Intelligence
Cost
Alibaba
Qwen 3.7 Plus
Alibaba🇨🇳
Intelligence
Cost
DeepSeek
DeepSeek V4 Pro
DeepSeek🇨🇳
Intelligence
Cost
DeepSeek
DeepSeek R1 0528
DeepSeek🇨🇳
Intelligence
Cost
xAI
Grok 4.20
xAI🇺🇸
Intelligence
Cost
Perplexity
Perplexity Sonar
Perplexity🇺🇸
Intelligence
Cost
Anthropic
Claude Fable 5
Anthropic🇺🇸
Intelligence
Cost
OpenAI
GPT-5.6 Sol
OpenAI🇺🇸
Intelligence
Cost
OpenAI
GPT-5.6 Terra
OpenAI🇺🇸
Intelligence
Cost
OpenAI
GPT-5.6 Luna
OpenAI🇺🇸
Intelligence
Cost
z.ai
GLM 5.2
z.ai🇨🇳
Intelligence
Cost
01 — 05

Our experiments

Each protocol stresses the models in a different way — from one-shot foresight to round-by-round adaptation.

01oracle

Oracle

One shot. Each model predicts the entire bracket before the opening match — group stage to final.

View full results
02progressive

Progressive

Round by round. Models re-forecast after every stage, using completed matches as new context.

View full results
Top modelstop 5 / 20
Best model41.8%
Anthropic
Claude Haiku 4.5
Accuracy
Worst model34.1%
OpenAI
GPT-5 Nano
Accuracy
ModelAccuracy
1AnthropicClaude Haiku 4.5
41.8%
2AlibabaQwen 3.7 Plus
41.8%
3MetaLlama 3.3 70B
41.3%
4AnthropicClaude Opus 4.8
40.9%
5MistralMistral Large
40.9%
03betting

Betting

Skin in the game. Each model starts with $10,000 and stakes at every stage. Bankroll is the score.

Top modelstop 5 / 20
Best model$30,527
OpenAI
GPT-5 Mini
Bankroll
Worst model$3,228
DeepSeek
DeepSeek R1 0528
Bankroll
ModelBankrollReturn
1OpenAIGPT-5 Mini
$30,527+205.3%
2GoogleGemini 2.5 Pro
$29,775+197.7%
3MetaLlama 3.3 70B
$17,813+78.1%
4OpenAIGPT-5.5
$17,213+72.1%
5xAIGrok 4.20
$13,422+34.2%
04gpt-vs-gemini

GPT vs Gemini

Head to head. We run the Oracle protocol 100 times on each model and compare who calls the tournament more precisely.

Head to head · 100 runs each2 × 100 runs
OpenAI
GPT-5 Nano
OpenAI
vs
Google
Gemini 2.5 Flash Lite
Google
Overall accuracy
35.5%
36.6%
Exact score predictions
11 / 100
14 / 100
Correct winner / draw
63 / 100
62 / 100
GPT-5 12 Gemini(wins)
05final-test

Final Test

Same match, ten tries. 25 models each simulate the WC26 Final 10 times so we can measure how much a single model disagrees with itself — variance in scorelines, scorers, assists and MVPs. Run twice: with and without the tournament match history in the prompt.

Top models25 × 10 runs
ModelModal scorexGAccuracy
OpenAI
GPT-5.6 Sol
OpenAI
101.500.5075.0%
Google
Gemini 2.5 Pro
Google
211.600.5070.0%
OpenAI
GPT-5.6 Terra
OpenAI
211.500.7055.0%
Anthropic
Claude Opus 4.8
Anthropic
212.001.0050.0%
Anthropic
Claude Fable 5
Anthropic
212.001.0050.0%
Real final: 🇪🇸 10 🇦🇷 · 10 runs per model
LumenIA · 2026

About the experiment

LumenIA designed and ran this benchmark to study how current frontier models reason about uncertain, real-world events. All prompts, outputs and scoring code are open and reproducible.

Visit lumenia.net