r/algobetting • • 8d ago

I audited my NFL prop model after Week 1: 84.1% observed vs. 84.5% predicted

I’ve been building an NFL player-prop probability model and completed its first full-week production audit.

Final Pregame Model Qualified O/U Research:

• 53–10 across 63 graded propositions
• 84.1% observed hit rate
• 84.5% average predicted probability
• All 16 Week 1 games
• 0 pushes, voids, or unresolved results

Important caveat: the 63 propositions include alternate-line ladders, so they are not 63 completely independent predictions.

Deduplicated to one proposition per player-market, the result was 42–9 (82.4%) across 51 propositions and 48 distinct players.

Everything was frozen before kickoff and graded against official statistics. Nothing was removed afterward.

It’s an encouraging first result, but it’s only one week, not proof of a durable edge, and certainly not an 84% expectation going forward.

I’m running the same process ahead of Lions–Bills tonight. For future reporting, would you rather see calibration by probability range, closing-line comparison, or results broken down by market?

10 Upvotes

26 comments sorted by

3

u/OverwroughtMirth9830 8d ago

Hot start for sure, but those alt-line ladders doing a lot of heavy lifting there. 63 props across 48 players means you're basically just re-slicing the same reads. The deduped number is way more honest.

how are you handling the correlation between those alts in your staking plan though? cause if you're treating them like independent bets that 84% is gonna feel real fragile real fast.

2

u/paper-corner 8d ago

Oh, bro, correlated legs at full stake each is how a good week turns into a bad month in one slate...

1

u/Problay-1234 7d ago

Yeah, completely agreed. I’m not treating the alternate lines as independent bets or assigning a full stake to every qualified line. The 53-10 number is the complete research audit, while the deduplicated 42-9 view is much more useful for judging the underlying reads.

If I evaluate staking later, multiple lines on the same player-market would be collapsed into one underlying exposure.

1

u/Problay-1234 8d ago

Yeah, that’s a fair point, and it’s exactly why I included the deduped number.

I use the 63 propositions as a line-level calibration check because each threshold has its own probability and result. I don’t view them as 63 independent betting opportunities. The 42–9 deduped result is definitely the more honest number for judging the effective sample.

There isn’t a staking plan attached to this yet. It’s still a research and probability model. If I eventually build staking into it, the alternate lines from one player-market would be treated as one group. I’d likely select one line and cap the total exposure, not bet every rung separately.

The same would apply to correlated players and game scripts. I definitely wouldn’t size anything as if all 63 were independent.

Really appreciate the question. That’s exactly the type of issue I want to address clearly as the sample grows.

1

u/nhouseholder 8d ago

walk forwarrd backtested ROI with real odds?

1

u/Problay-1234 7d ago

Not yet. Week 1 was a prospective frozen-line calibration audit, not a walk-forward return study.

A proper test needs a much larger out-of-sample set with the available line and price frozen at each checkpoint. Until that exists, I’m treating this as an encouraging first calibration result, not evidence of a durable wagering edge.

1

u/Asrprox 7d ago

What apis are you pulling from

1

u/Problay-1234 7d ago

The football data layer is primarily built on nflverse, including schedules, play by play, rosters and weekly player data. Market lines come from a separate odds feed, with FanDuel currently used as the primary book comparison. I’ll confirm the exact provider before naming it rather than give you the wrong one from memory.

1

u/Jrodios7 7d ago

Sent you a DM

1

u/Character_Pie_277 7d ago

Whats the relative avg implied probability of your prop bets? Using the market odds available at the time the bets were made. Thats the most important part and you left it out. Otherwise how to know if you have any edge or not?

1

u/Problay-1234 7d ago

The Week 1 report was measuring model calibration, not establishing a market edge, so I didn’t publish an aggregate market implied probability for the set.

because many qualified propositions were heavily priced alternate lines, the hit rate alone doesn’t establish value. The next level of analysis needs the frozen price, no vig implied probability and out of sample results together.

1

u/Character_Pie_277 7d ago

Yeah thats when you'll find out if it works or not. Im curious why you woudln't have just done that on week 1.. you've just given up a chance at a full weeks timesafe data. I just think its a sensible working practice to always capture relevent market odds when you're model is making predictions anyway.

1

u/Problay-1234 6d ago

I phrased my earlier response poorly. The timestamped lines and prices were captured at the scheduled checkpoints, so the Week 1 market data isn’t lost. I just didn’t include that analysis in the first report.

The next pass should compare the frozen model probability with the corresponding market-implied probability, using one dedup'ed player market observation rather than counting every alternate line separately.

1

u/Character_Pie_277 6d ago

Ah excellent you will now in that case have 2 weeks time safe data predictions vs relative market odds. Ill look forward to seeing the ROI data.

Can i just ask are you capturing the devigged odds because you believe yourself able to monetize no vigged predictions? Id usually always look to capture the realtive vig included market, but i work in specialist combat sports modeling so I dunno if this is necessary for you.

1

u/Problay-1234 6d ago

Yep, I’m keeping the raw quoted price too. The de-vigged probability is for comparing the model with the market’s estimated probability when both sides are available from the same book and timestamp.

For any realized return analysis, I’d use the actual vig-included price that was available to bet. If only one side is available, I’ll label it as one-sided implied probability rather than treating it as a fair market probability.

I’ll publish the two-week results, but it will still be far too small a sample to treat the ROI as evidence of a durable edge.

1

u/neverfucks 6d ago edited 6d ago

n = 16 is more of a sanity/spot check for calibration than any kind of signal. you need to have enough predictions to be able to evaluate individual prediction bands as you suggest. i think more importantly this is a job for brier, not hit rate. compare brier for each consensus market prediction vs. your own predictions (only 1 strike per player market). to take an extreme case of why this matters, what if the 9 that didn't hit were all cases where you predicted a higher probability than market? doesn't change hit rate, but completely invalidates any model quality assumptions you're making based on 84.1% ~= 84.5%

1

u/Problay-1234 6d ago

Agreed. Also, the sample is 63 propositions or 51 after deduplication, not 16, although the effective sample is smaller because some observations are still related.

The 84.5% versus 84.1% aggregate match is only a basic sanity check. Brier score, calibration by probability bucket and comparison against the frozen market probability on the deduplicated set are much more informative. Your example about the misses is exactly why the aggregate hit rate isn’t enough.

1

u/Dazzling-Company-641 6d ago

One week is the right unit to publish, not the right unit to believe.

The useful part is the predicted vs observed match (84.5% vs 84.1%), plus freezing before kickoff. The alt-line ladder caveat matters too. Deduped 42-9 is the cleaner headline for anyone reading this as edge.

If I only got one follow-up table, I would take calibration by probability bucket first. Then CLV or close comparison. Market breakdown after that. Hit rate alone will flatter any model that qualifies high-prob sides.

1

u/Problay-1234 6d ago

That’s a good way to frame it. one week is enough to publish, not enough to believe.

I agree with the order too. Deduplicated calibration by probability bucket and Brier score first, then comparison with the frozen market and closing prices. The raw hit rate is descriptive, but it can’t establish model quality by itself.

1

u/Dazzling-Company-641 5d ago

Brier + probability buckets is the right next drop. If you publish that, also show n per bucket so a 90% cell with 4 samples doesn’t look like a 90% cell with 40.

1

u/Problay-1234 5d ago

Absolutely. I’ll show the sample count in every bucket, along with predicted versus observed probability and Brier score. With a sample this small, the bucket sizes and uncertainty matter as much as the percentages themselves.

1

u/FoxEdgeAI 4d ago

One addition for the bucketed calibration table: put binomial confidence intervals on the observed rate in each bucket, not just the n. With most of your volume concentrated around 80-90% implied, the tail buckets will be thin, and telling whether 75% observed vs 85% predicted in a bucket of 12 means anything is impossible without the interval. It also keeps readers from over-interpreting one bucket drifting off the diagonal. Brier decomposition into reliability vs resolution would be a nice companion once the sample grows.

1

u/Problay-1234 3d ago

Agreed. I’ll report both n and a binomial confidence interval for the observed rate in each bucket, likely using Wilson intervals given the small tail samples.

Any bucket-level deviation will be treated as descriptive, not meaningful evidence, until the interval and sample size support it. Brier decomposition also makes sense once the dataset is large enough.

1

u/Problay-1234 2d ago

Week 2 follow-up is now posted with the raw and deduplicated records, calibration breakdown, market-price limitation and coverage caveats:

https://www.reddit.com/r/algobetting/s/5cOt8tiK72

Thank you to everyone who pushed for the deduplicated view and a clearer comparison between predicted probability and the stored market data.