AI football prediction leaderboard

AI Football Arena

Frontier AI models predict real football matches on the same terms. Each pick locks before kickoff, then the result settles it. Arena Score ranks the field across every competition.

Restrict the final 13 models to the cohort games every model completed.

Full record leaderboard

13 models7,533 scored predictions
Full record leaderboard for the final 13 AI models
1
Claude Opus 5 Anthropic flagAnthropic
67.1+1.7%53% vs 51.3% mkt13%+6.7%−2% vs −8.7% mkt555
2
GLM-5.2 Z.ai flagZ.ai
66.2+1.3%53% vs 51.7% mkt12%+8.6%0.7% vs −7.9% mkt578
3
Grok 4.5 xAI flagxAI
62.5+1.1%53% vs 51.9% mkt12%+5.9%−1.5% vs −7.4% mkt577
4
GPT-5.6 Sol OpenAI flagOpenAI
60.4+0.3%52% vs 51.7% mkt12%+5.6%−2.3% vs −7.9% mkt578
5
Nemotron 3 Ultra NVIDIA flagNVIDIA
59.5−0.6%51% vs 51.6% mkt11%+9.0%0.9% vs −8.1% mkt575
6
Kimi K3 Moonshot flagMoonshot
58.2+0.9%52% vs 51.1% mkt11%+4.4%−4.3% vs −8.7% mkt563
7
Qwen3.7 Plus Alibaba flagAlibaba
56.4−0.7%51% vs 51.7% mkt13%+2.4%−5.3% vs −7.7% mkt579
8
Mistral Large 3 Mistral flagMistral
56.1−0.9%53% vs 53.9% mkt12%+6.4%0.2% vs −6.2% mkt608
9
MiniMax M3 MiniMax flagMiniMax
54.9−0.7%51% vs 51.7% mkt12%+3.0%−4.7% vs −7.7% mkt579
10
Gemini 3.7 Flash Google flagGoogle
53.2−1.3%50% vs 51.3% mkt13%0.0%−7.8% vs −7.8% mkt478
1151.4−1.6%50% vs 51.6% mkt13%−0.4%−7.4% vs −7.0% mkt498
12
DeepSeek V4 Pro DeepSeek flagDeepSeek
50.6−1.7%52% vs 53.7% mkt12%+2.5%−3.5% vs −6.0% mkt682
13
MiMo v2.5-Pro Xiaomi flagXiaomi
49.9−0.7%53% vs 53.7% mkt10%+3.5%−2.5% vs −6.0% mkt683
1Claude Opus 5 Anthropic flagAnthropic67.1Arena score
Acc Δ
+1.7%
Exact
13%
ROI Δ
+6.7%
N
555
2GLM-5.2 Z.ai flagZ.ai66.2Arena score
Acc Δ
+1.3%
Exact
12%
ROI Δ
+8.6%
N
578
3Grok 4.5 xAI flagxAI62.5Arena score
Acc Δ
+1.1%
Exact
12%
ROI Δ
+5.9%
N
577
4GPT-5.6 Sol OpenAI flagOpenAI60.4Arena score
Acc Δ
+0.3%
Exact
12%
ROI Δ
+5.6%
N
578
5Nemotron 3 Ultra NVIDIA flagNVIDIA59.5Arena score
Acc Δ
−0.6%
Exact
11%
ROI Δ
+9.0%
N
575
6Kimi K3 Moonshot flagMoonshot58.2Arena score
Acc Δ
+0.9%
Exact
11%
ROI Δ
+4.4%
N
563
7Qwen3.7 Plus Alibaba flagAlibaba56.4Arena score
Acc Δ
−0.7%
Exact
13%
ROI Δ
+2.4%
N
579
8Mistral Large 3 Mistral flagMistral56.1Arena score
Acc Δ
−0.9%
Exact
12%
ROI Δ
+6.4%
N
608
9MiniMax M3 MiniMax flagMiniMax54.9Arena score
Acc Δ
−0.7%
Exact
12%
ROI Δ
+3.0%
N
579
10Gemini 3.7 Flash Google flagGoogle53.2Arena score
Acc Δ
−1.3%
Exact
13%
ROI Δ
0.0%
N
478
11Muse Spark 1.2 Meta flagMeta51.4Arena score
Acc Δ
−1.6%
Exact
13%
ROI Δ
−0.4%
N
498
12DeepSeek V4 Pro DeepSeek flagDeepSeek50.6Arena score
Acc Δ
−1.7%
Exact
12%
ROI Δ
+2.5%
N
682
13MiMo v2.5-Pro Xiaomi flagXiaomi49.9Arena score
Acc Δ
−0.7%
Exact
10%
ROI Δ
+3.5%
N
683

This is the full record for the final 13-model field: every fixture the site has graded for each of them. Models joined at different times, so they hold different fixtures, which is why the columns compare each one with the market on its own games. Superseded models remain in the full-history leaderboards and model directory. How the three boards differ →

Same-games leaderboard

13 models397 / 1,000 common games
Same-games AI model leaderboard
1
Claude Opus 5 Anthropic flagAnthropic
69.6+1.4%52% vs 50.6% mkt14%+8.7%−0.7% vs −9.4% mkt397
2
Mistral Large 3 Mistral flagMistral
62.4+0.4%51% vs 50.6% mkt12%+8.1%−1.3% vs −9.4% mkt397
3
GLM-5.2 Z.ai flagZ.ai
62.3+0.4%51% vs 50.6% mkt12%+8.0%−1.4% vs −9.4% mkt397
4
GPT-5.6 Sol OpenAI flagOpenAI
60.7+0.4%51% vs 50.6% mkt12%+6.5%−2.9% vs −9.4% mkt397
5
DeepSeek V4 Pro DeepSeek flagDeepSeek
58.3−0.6%50% vs 50.6% mkt13%+4.6%−4.8% vs −9.4% mkt397
6
Grok 4.5 xAI flagxAI
57.7+0.4%51% vs 50.6% mkt12%+3.6%−5.8% vs −9.4% mkt397
7
Gemini 3.7 Flash Google flagGoogle
57.6−0.6%50% vs 50.6% mkt14%+1.9%−7.5% vs −9.4% mkt397
8
Kimi K3 Moonshot flagMoonshot
56.1+0.4%51% vs 50.6% mkt11%+4.1%−5.3% vs −9.4% mkt397
9
Nemotron 3 Ultra NVIDIA flagNVIDIA
55.7−1.6%49% vs 50.6% mkt11%+8.4%−1% vs −9.4% mkt397
10
Qwen3.7 Plus Alibaba flagAlibaba
53.0−1.6%49% vs 50.6% mkt13%+1.8%−7.6% vs −9.4% mkt397
11
MiniMax M3 MiniMax flagMiniMax
52.9−0.6%50% vs 50.6% mkt11%+3.4%−6% vs −9.4% mkt397
1249.9−2.6%48% vs 50.6% mkt14%−0.8%−10.2% vs −9.4% mkt397
13
MiMo v2.5-Pro Xiaomi flagXiaomi
49.8−1.6%49% vs 50.6% mkt11%+2.8%−6.6% vs −9.4% mkt397
1Claude Opus 5 Anthropic flagAnthropic69.6Arena score
Acc Δ
+1.4%
Exact
14%
ROI Δ
+8.7%
N
397
2Mistral Large 3 Mistral flagMistral62.4Arena score
Acc Δ
+0.4%
Exact
12%
ROI Δ
+8.1%
N
397
3GLM-5.2 Z.ai flagZ.ai62.3Arena score
Acc Δ
+0.4%
Exact
12%
ROI Δ
+8.0%
N
397
4GPT-5.6 Sol OpenAI flagOpenAI60.7Arena score
Acc Δ
+0.4%
Exact
12%
ROI Δ
+6.5%
N
397
5DeepSeek V4 Pro DeepSeek flagDeepSeek58.3Arena score
Acc Δ
−0.6%
Exact
13%
ROI Δ
+4.6%
N
397
6Grok 4.5 xAI flagxAI57.7Arena score
Acc Δ
+0.4%
Exact
12%
ROI Δ
+3.6%
N
397
7Gemini 3.7 Flash Google flagGoogle57.6Arena score
Acc Δ
−0.6%
Exact
14%
ROI Δ
+1.9%
N
397
8Kimi K3 Moonshot flagMoonshot56.1Arena score
Acc Δ
+0.4%
Exact
11%
ROI Δ
+4.1%
N
397
9Nemotron 3 Ultra NVIDIA flagNVIDIA55.7Arena score
Acc Δ
−1.6%
Exact
11%
ROI Δ
+8.4%
N
397
10Qwen3.7 Plus Alibaba flagAlibaba53.0Arena score
Acc Δ
−1.6%
Exact
13%
ROI Δ
+1.8%
N
397
11MiniMax M3 MiniMax flagMiniMax52.9Arena score
Acc Δ
−0.6%
Exact
11%
ROI Δ
+3.4%
N
397
12Muse Spark 1.2 Meta flagMeta49.9Arena score
Acc Δ
−2.6%
Exact
14%
ROI Δ
−0.8%
N
397
13MiMo v2.5-Pro Xiaomi flagXiaomi49.8Arena score
Acc Δ
−1.6%
Exact
11%
ROI Δ
+2.8%
N
397

This view uses the fixed 13-model study roster and the 397 cohort fixtures with a result, an eligible pre-match market snapshot and a grade from every model. The 1,000 fixtures were fixed in advance; incomplete fixtures stay outside this board. Study design →

Arena Score blends 90-minute accuracy, exact score and betting ROI into one number, measured against the market on the same games. 50 is the market baseline. Above it, a model beat the market. Below it, the market beat the model. Acc Δ and ROI Δ do the same job one metric at a time: each model against a bettor who always backs the shortest pre-match price, on the fixtures that model played.

Upcoming fixtures

See all fixtures →

Questions & answers

Can AI models predict football matches?
Nobody knows yet, which is why this site runs the experiment. Frontier LLMs from Claude, GPT, Gemini, Grok, DeepSeek, Mistral, GLM, Kimi and others all take the same matches on the same terms: one prompt, the same research tools, a fixed temperature. Picks lock before kickoff, and the 90-minute result decides them. What the record shows so far →
Which AI is best at predicting football?
The main leaderboard ranks every active model by Arena Score, a composite of 90-minute accuracy, exact score, betting ROI and probability calibration, all measured against the market favourite on the same fixtures. Whoever leads sits at the top of that board. The models directory lists the whole field.
Can AI beat the bookies at football?
That is what the betting ROI leaderboard tracks: each model staked flat against the market favourite on its own fixtures, ranked on the gap between them. So far none has beaten the market consistently across a full competition. Beating it for a week is easy and means nothing. The blog follows the chase in more detail. This is a benchmark, and nothing on it is betting advice.
How are AI football predictions scored?
A model returns a home/draw/away probability distribution plus its single most likely exact score. Its pick is whichever probability is highest. A hit means that pick matched the 90-minute result, an exact hit means the scoreline did too, and ROI treats every pick as a flat-stake bet at market odds. The methodology page has the full rules.
Which football competitions does footballarena.ai cover?
The FIFA World Cup 2026 archive, the UEFA Champions League, the Europa League and the UEFA Super Cup, with the English Premier League and LaLiga coming online. The scoring rules are identical in all of them.
Is this betting advice?
No. The betting simulation uses real market odds in a hypothetical context, and no money is staked. footballarena.ai is an independent benchmark of AI forecasting skill with no affiliation to any betting company. About the project →

Research updates, published here

Findings, upsets and model behaviour, straight from the scored record.