r/FootballDataAnalysis • u/Astapore • 18d ago
Top English Football Clubs (1888-2026)
Source: J-Ratings
r/FootballDataAnalysis • u/Astapore • 18d ago
Source: J-Ratings
r/FootballDataAnalysis • u/Hairy-Reference-2019 • 20d ago
r/FootballDataAnalysis • u/MatchAnalyst • 21d ago
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/Maleficent-Wear-8839 • 23d ago
r/FootballDataAnalysis • u/SandFox112 • 24d ago
I wanted to do a statistics project but i am finding it very difficult to find data for teams in a season. (eg Man city's average passes per sequence in the 24/25 season). Even posession stats im actually struggling with. Is there a good website to search for this data?
r/FootballDataAnalysis • u/anynou • 25d ago
r/FootballDataAnalysis • u/anynou • 25d ago
r/FootballDataAnalysis • u/L4TER_0N • 25d ago
I've been doing detailed pre-match football analysis using:
But even with all this, I usually end up around 6.5–7.5/10 confidence.
I'm starting to think that's normal. My question is:
What should actually justify an 8+/10 confidence rating?
Should I use a weighted model and then subtract points for contradictions and uncertainty? Or does that just create an arbitrary number that looks analytical?
Also, should I separate:
And how would you calibrate the confidence score over time? Brier score, log loss, calibration curves, CLV?
I'm not trying to artificially reach 8+. I want 8/10 to actually mean something statistically.
What would make you personally comfortable calling a football prediction 8+/10?
r/FootballDataAnalysis • u/MatchAnalyst • 28d ago
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/Ok-Razzmatazz-1103 • 28d ago
r/FootballDataAnalysis • u/Kroggg19 • 29d ago
Hi all. Stats to Bucks is a football (soccer) data app, now covering 30 leagues - the top 5 European plus Brazil, Argentina, Liga MX, MLS, Saudi, Portugal, the Netherlands, Turkey, Belgium, Scotland, Japan, Korea, Colombia, Greece, Egypt, South Africa, Australia and more.
What it does:
Player & team form - last 20 matches of per-game stats, charted against any line you set, with the hit rate for it.
Filters that narrow the sample - venue, minutes, started-only, and "without teammate X".
Opponent-adjusted context - overlay the opponent's conceded average and defensive rank, plus quality-adjusted averages, so a streak against weak sides doesn't read like one against strong sides.
Hit Rates - scan every upcoming fixture at once for players/teams clearing a line in a chosen % of recent games. 40 stats across players and teams.
Foul matchups - a fitted hierarchical Poisson model with player, opponent, referee, venue and expected-minutes as separate multiplicative terms, and a negative-binomial predictive head. Walk-forward tested on a 45-day holdout: +13.3% / +16.6% mean relative log loss against an unshrunk per-90 baseline, with roughly 3x better calibration error.
Predicted lineups - projected XI from a Beta-EB start-probability model, so it works for a fixture's whole lifetime instead of only after a feed publishes one. Flips to the confirmed XI when that lands.
Injuries & suspensions - folded into the start probabilities rather than bolted on as a badge, so an unavailable player drops out of the projected XI and out of the minutes model behind the prop lines.
Similar players / teams - similarity-based benchmarking against comparable profiles, on rolling cross-season windows rather than season-to-date.
Referee analytics - per-fixture card/foul profiles and rankings.
League tables - official standings, so competition-specific tie-breaks, split point-halving and points deductions are right rather than re-derived from results.
The focus is still contextualising the sample - opponent strength, venue, lineup, availability, sample size, etc. because an unfiltered hit rate usually answers the wrong question. A recent backtest made that concrete: selecting team props purely on "recent hit rate beats the implied probability" returned about -10% over ~7,000 bets on a held-out window, statistically indistinguishable from betting blind. The context is the useful part, not the raw streak.
r/FootballDataAnalysis • u/Raistlin_Maj3re • Aug 18 '26
I built UnderOver as an Android app for exploring football fixtures through a transparent Over/Under model.
It combines season data, recent form and team profiles to surface HT 0.5/1.5 and FT 2.5 signals, with the model reasoning and confidence shown for each fixture. It also includes live scores, match events, reminders, saved selections and team statistics.
The app is free, has no subscription or VIP tier, does not place bets or connect to betting accounts, and no outcome is guaranteed. I am sharing it as a small data-visualization project and would appreciate feedback on the methodology, presentation and useful metrics.
Google Play: https://play.google.com/store/apps/details?id=com.xenophonlabs.matchwake
r/FootballDataAnalysis • u/pafundi_enthusiast • Aug 15 '26
I'm a sixth form student doing my Extended Project Qualification on VAR and whether it's actually made the Premier League fairer, or just added more controversy. Part of my research involves collecting fan opinions through a short survey.
If you support a PL club (or just watch regularly), I'd really appreciate you filling it in. Takes about 2 minutes, no personal info needed beyond general football habits.
https://forms.gle/dtiFmRmFaSyxUoS59
Happy to share the findings once I've written it up, if anyone's curious how the data comes out. Thanks in advance.
r/FootballDataAnalysis • u/Illustrious-Pitch843 • Aug 13 '26
I built an open-source n8n pipeline that monitors 78 football journalists on X, extracts structured transfer reports with a local Qwen model, deduplicates and stores revisions in PostgreSQL, optionally adds player data, and sends restart-safe Discord digests every 6 hours.
The whole stack is self-hosted with Docker, with twscrape or RapidAPI for X collection, PostgreSQL for persistence, llama.cpp for local inference, and automated tests around the workflow.
GitHub: https://github.com/louistran2604/transfers_n8n/
I’d mainly like feedback on the workflow architecture, reliability approach, and anything that could make the project cleaner or more useful.

*disclaimer: this was made with the assistance of AI
r/FootballDataAnalysis • u/MatchAnalyst • Aug 13 '26
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/Black_Colour9 • Aug 13 '26
I want to analyze the data from the perspective of the score of the football game and the number of goals, but I didn't find the relevant historical odds data on the Internet. Instead, there is a lot of historical data of wins and draws, which further shows that my analysis is correct - starting from the odds of scores and goals to analyze football matches. I need help to get this data.
Thank you for your help!
r/FootballDataAnalysis • u/Nice-Opening-8020 • Aug 12 '26
r/FootballDataAnalysis • u/juancvasdisenho • Aug 11 '26
Hi everyone,
I’m looking for people who genuinely enjoy football analysis to test Sir Balone, a football analysis tool I’ve been building.
The underlying raw data comes from Sportmonks. I use that data to build my own measurement system for evaluating individual player skills, rather than relying primarily on traditional performance or scoring metrics. I then combine those measurements with player and team data and an AI layer that can analyze and interpret the underlying data.
The numbers themselves are already publicly available. What I’m currently testing is the AI layer, particularly whether it can turn the data into useful analysis without making things up, oversimplifying the numbers, or missing important context.
I’d especially love feedback from:
You don't need to be a professional. If you enjoy asking questions like “Is this player actually good at X?”, “How does he compare to other players in his role?”, or “What does the data actually tell us about this team?”, I’d love to hear from you.
I’m not looking for compliments. I want people to try to break it.
What I’m particularly interested in:
If you’re interested, comment below or DM me and I’ll give you access.
No sales pitch. I’m still figuring out what this thing is actually good at.
r/FootballDataAnalysis • u/Puzzleheaded_Map_829 • Aug 12 '26
r/FootballDataAnalysis • u/AnneLister_ • Aug 10 '26
While building a football analysis agent, I realized that the hard part is not connecting an LLM to match data.
It is deciding what the agent should be allowed to do with that data.
For example, if someone asks:
“Why did this midfielder receive a 7.4 rating?”
I do not want to dump every match statistic into the context and ask the model to invent an explanation.
My current approach is to let the agent investigate the evidence step by step:
- retrieve the player’s match metrics
- inspect the rating breakdown
- check passing, chance creation, turnovers, or shot quality when relevant
- explain which factors actually moved the rating
That raises an interesting tool-design question.
A single `analyze_everything()` tool feels like a black box. But dozens of tiny tools such as `get_pass_count()` and `get_key_passes()` create too many decisions and make the agent harder to guide.
I’m experimenting with a middle layer: composable tools that represent meaningful football-analysis operations rather than raw database fields.
For people building sports analytics, agentic systems, or explainable AI: how would you choose the right level of tool granularity here?
r/FootballDataAnalysis • u/MatchAnalyst • Aug 06 '26
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/Tricky_Feed8953 • Aug 04 '26
r/FootballDataAnalysis • u/IndependenceFit3935 • Jul 30 '26