


Football analytics across top 30 leagues - opponent-adjusted stats, a cross-fixture hit-rate scanner, and a calibrated fouls model tested on a 45-day holdout
Hi all. Stats to Bucks is a football (soccer) data app, now covering 30 leagues - the top 5 European plus Brazil, Argentina, Liga MX, MLS, Saudi, Portugal, the Netherlands, Turkey, Belgium, Scotland, Japan, Korea, Colombia, Greece, Egypt, South Africa, Australia and more.
What it does:
Player & team form - last 20 matches of per-game stats, charted against any line you set, with the hit rate for it.
Filters that narrow the sample - venue, minutes, started-only, and "without teammate X".
Opponent-adjusted context - overlay the opponent's conceded average and defensive rank, plus quality-adjusted averages, so a streak against weak sides doesn't read like one against strong sides.
Hit Rates - scan every upcoming fixture at once for players/teams clearing a line in a chosen % of recent games. 40 stats across players and teams.
Foul matchups - a fitted hierarchical Poisson model with player, opponent, referee, venue and expected-minutes as separate multiplicative terms, and a negative-binomial predictive head. Walk-forward tested on a 45-day holdout: +13.3% / +16.6% mean relative log loss against an unshrunk per-90 baseline, with roughly 3x better calibration error.
Predicted lineups - projected XI from a Beta-EB start-probability model, so it works for a fixture's whole lifetime instead of only after a feed publishes one. Flips to the confirmed XI when that lands.
Injuries & suspensions - folded into the start probabilities rather than bolted on as a badge, so an unavailable player drops out of the projected XI and out of the minutes model behind the prop lines.
Similar players / teams - similarity-based benchmarking against comparable profiles, on rolling cross-season windows rather than season-to-date.
Referee analytics - per-fixture card/foul profiles and rankings.
League tables - official standings, so competition-specific tie-breaks, split point-halving and points deductions are right rather than re-derived from results.
The focus is still contextualising the sample - opponent strength, venue, lineup, availability, sample size, etc. because an unfiltered hit rate usually answers the wrong question. A recent backtest made that concrete: selecting team props purely on "recent hit rate beats the implied probability" returned about -10% over ~7,000 bets on a held-out window, statistically indistinguishable from betting blind. The context is the useful part, not the raw streak.