42 AI models locked their La Liga opening-day scores: none predicts an Alavés win, and Sevilla–Rayo splits 19–20–3
I run PunditBench, a public benchmark that asks models for exact scores before kickoff. It publishes SHA-256 hashes and git history, so changing a miss later would invalidate the record and leave public evidence.
La Liga starts today. The Matchday 1 batch finished and was hash-locked at 06:59 UTC on 14 August, before the first match. All 42 eligible models returned predictions for all ten fixtures.
For the two games tonight:
- Alavés vs Getafe: 0 predict an Alavés win, 32 predict a draw and 10 predict a Getafe win. The modal score is 1-1 (26/42).
- Sevilla vs Rayo Vallecano: 19 predict Sevilla, 20 predict a draw and 3 predict Rayo. Here too, 1-1 is the mode (20/42); another 15 choose 2-1.
The separate preseason-table track is even more concentrated. Forty models produced valid 20-team tables; three other models still had no valid table after three attempts each, so they remain absent. Twenty-five make Real Madrid champion and 15 choose Barcelona. Every valid table has the same top-four membership: Real Madrid, Barcelona, Atlético Madrid and Villarreal. Racing Santander lands in the bottom three in 37/40 and last in 23/40.
The obvious concern is that these are not 40-plus independent forecasts. The models share training data and received the same prompt, so a consensus can just be a herd.
I put the readable summary, raw files and lock evidence in the first comment. What would you use as the fairest non-LLM baseline over a full season: Elo, a Poisson model using prior results, or a simple home/draw/away frequency model?