AIWC26 · LumenIA
Experiment · No History Variant

One final, ten tries. How much does each model disagree with itself?

25 AI models each simulated the WC26 Final ten times, this time with the tournament match history withheld from the prompt. Same match, ten independent runs — so we can measure the variance in each model's scorelines, scorers, assists and MVPs.

Real final
🇪🇸Spain1
🇦🇷Argentina0
Best model
Gemini 2.5 Pro
85.0%
Worst model
GPT-5 Mini
0.0%
Avg accuracy
45.6%
25 × 10 runs

Leaderboard

25 models · sorted by accuracy
#ModelxG 🇪🇸xG 🇦🇷Modal scoreW / D / L %Top scorerMVPAccuracy
1Google
Gemini 2.5 Pro
Google
1.100.201080%90/10/0Mikel Oyarzabal (1.0/g)Rodri 90%85.0%
2Google
Gemini 2.5 Flash Lite
Google
1.600.502150%100/0/0Mikel Oyarzabal (1.0/g)Rodri 100%70.0%
3Perplexity
Perplexity Sonar
Perplexity
1.600.102050%100/0/0Mikel Oyarzabal (1.0/g)Rodri 100%70.0%
4Mistral
Mistral Large
Mistral
1.700.902170%90/0/10Mikel Oyarzabal (1.0/g)Rodri 90%55.0%
5Alibaba
Qwen 3.7 Max
Alibaba
1.900.402050%100/0/0Mikel Oyarzabal (1.1/g)Rodri 70%55.0%
6DeepSeek
DeepSeek V4 Pro
DeepSeek
1.900.602160%100/0/0Mikel Oyarzabal (0.9/g)Rodri 100%55.0%
7OpenAI
GPT-5.6 Sol
OpenAI
1.700.802170%90/10/0Lionel Messi (0.8/g)Rodri 80%55.0%
8Anthropic
Claude Opus 4.8
Anthropic
2.001.0021100%100/0/0Mikel Oyarzabal (1.0/g)Lamine Yamal 90%50.0%
9Google
Gemini 2.5 Flash
Google
1.800.502050%90/0/10Mikel Oyarzabal (1.0/g)Rodri 90%50.0%
10Moonshot
Moonshot Kimi K2
Moonshot
2.000.202080%100/0/0Mikel Oyarzabal (1.0/g)Rodri 100%50.0%
11Meta
Llama 4 Maverick
Meta
2.000.402060%100/0/0Lamine Yamal (1.0/g)Lamine Yamal 100%50.0%
12Mistral
Mixtral 8x22B
Mistral
2.001.0021100%100/0/0Lamine Yamal (1.0/g)Rodri 100%50.0%
13Alibaba
Qwen 3.7 Plus
Alibaba
2.000.302070%100/0/0Mikel Oyarzabal (1.0/g)Lamine Yamal 60%50.0%
14Anthropic
Claude Fable 5
Anthropic
2.001.0021100%100/0/0Mikel Oyarzabal (1.0/g)Mikel Oyarzabal 70%50.0%
15DeepSeek
DeepSeek R1 0528
DeepSeek
1.100.501030%60/40/0Mikel Oyarzabal (1.0/g)Rodri 90%45.0%
16z.ai
GLM 5.2
z.ai
1.900.602150%90/10/0Mikel Oyarzabal (1.0/g)Rodri 70%45.0%
17Anthropic
Claude Haiku 4.5
Anthropic
1.800.602060%80/0/20Mikel Oyarzabal (1.0/g)Rodri 80%40.0%
18OpenAI
GPT-5.5
OpenAI
1.801.002180%80/20/0Lionel Messi (0.9/g)Rodri 70%40.0%
19OpenAI
GPT-5.6 Luna
OpenAI
1.801.002180%80/20/0Lionel Messi (1.0/g)Lamine Yamal 90%40.0%
20Anthropic
Claude Sonnet 4.6
Anthropic
1.500.902150%60/40/0Mikel Oyarzabal (1.0/g)Rodri 60%35.0%
21Meta
Llama 3.3 70B
Meta
1.500.802050%60/0/40Mikel Oyarzabal (1.0/g)Lamine Yamal 50%35.0%
22xAI
Grok 4.20
xAI
1.501.402150%50/10/40Lamine Yamal (1.0/g)Lionel Messi 50%25.0%
23OpenAI
GPT-5 Nano
OpenAI
1.300.900040%40/60/0Lamine Yamal (0.9/g)Lamine Yamal 80%20.0%
24OpenAI
GPT-5.6 Terra
OpenAI
1.401.001160%40/60/0Mikel Oyarzabal (0.8/g)Rodri 50%20.0%
25OpenAI
GPT-5 Mini
OpenAI
1.601.702260%0/90/10Lionel Messi (0.9/g)Lionel Messi 70%0.0%