How the First & Thirty Player Prop Model Works
We’ve spent a good chunk of the offseason building a player prop model for First & Thirty. With Week 1 finally here, it’s probably worth explaining what the hell it actually does.
The basic idea started with fantasy projections.
There are a lot of smart people putting out NFL projections every week. We’re currently collecting projections from FantasyPros, PFF, FTN and 4for4. Each source gives us its expectation for things like passing yards, rushing yards, receiving yards and workload.
The obvious approach would be to average those projections, compare the average to the sportsbook line, and bet when there’s a big enough difference.
There’s a problem with that.
A projection of 80 yards does not mean 80 is the most likely outcome
This is the biggest issue that people make while betting props. There's a projection that I know and trust that has a receiver at 80 yards, and a sportsbook line offering that same player at 78. Great, I'll bet the over. Easy enough, right? But this is the most common error when prop betting.
Fantasy projections are mean based, and are generally trying to estimate a player’s average outcome.
Player prop lines behave much more like the middle of the distribution (the median), and those aren't necessarily the same number.
Take a running back projected for 70 rushing yards. His possible outcomes might look something like this:
- 15 yards after getting hurt
- 35 yards in a game where his team falls behind
- 55 yards on an average-ish day
- 75 yards in a good game
- 140 yards when he breaks a couple of huge runs
Those 140-yard games pull his average upward. There isn't an equivalent outcome below zero that can pull it back down.
So a player can legitimately average 70 yards while finishing below 70 in more than half of his games.
That's the entire foundation of the model, and it's also where the Monte Carlo simulation comes in.
Four projection sources, four separate simulations
Every week we collect projections from:
- FantasyPros
- PFF
- FTN
- 4for4
Each source gets its own simulation.
If PFF projects a running back for 72 yards and FTN projects him for 64, both projections survive. We simulate a distribution around PFF's expectation and another around FTN's.
The same applies to the other two sources.
That gives us four separate opinions about the probability of going Over or Under a particular sportsbook line.
This matters because these sources disagree. If all four independently point toward the same side after going through the simulation, that's a lot more interesting to us than an average that can hide disagreement.
One important caveat: these aren't four completely independent opinions. Projection systems use a lot of the same underlying NFL information, and their projections are highly correlated. We treat agreement as evidence of robustness, not as four independent coin flips.
Where the distributions come from
This was the biggest part of the project.
We pulled NFL play-by-play and player data going back to 2018 and used it to study how these stats actually behave.
Running backs, quarterbacks and receivers don't accumulate yards the same way, so we don't use one generic distribution for everything.
Running backs
For running backs, we use each projection source's expected:
- carries
- rushing yards
From there, we simulate how many carries the player actually gets and what happens on those carries.
The historical data gives us the shape of real NFL rushing outcomes. That includes the ugly stuff. Zero-yard carries, negative plays, ordinary gains and explosive runs all remain part of the distribution.
The final simulation is anchored back to the projection source's expected rushing-yard average.
So if a source says 70 yards, our simulation still averages approximately 70 yards. We're figuring out all the different ways a player can get there.
Quarterback rushing
Quarterbacks required their own model because QB rushing is weird.
Scrambles, designed runs and kneel-downs create a very different distribution from running back carries.
We tested a more complicated model that explicitly separated those components. It didn't improve the final calibration enough to justify the added complexity.
The version we're using models total quarterback carries and historical QB rushing outcomes directly.
We also restrict the model based on projected workload. Trying to model a quarterback projected for basically no rushing usage turned out to be pretty noisy, so the model is allowed to say, essentially, "we don't have enough here." With betting, it's ok to abstain.
Passing yards
For quarterbacks, we tested three different ways of building the distribution.
One simulated individual pass attempts.
Another separately modeled attempts, completions and yards per completion.
The third modeled pass attempts and game-level yards per attempt.
The game-level yards-per-attempt version won.
It was the simplest of the serious candidates and produced the best calibration when we tested it against a held-out set of 2024 FantasyPros projections.
For the current model, each source gives us expected pass attempts and passing yards. We simulate the uncertainty in passing volume and game-level efficiency, then anchor the resulting distribution to that source's projected passing-yard mean.
We currently require at least 20 projected attempts before using the distribution for our betting research.
Receiving yards
Receiving yards turned out to be one of the more complicated markets.
Targets and catches create a lot of zeroes and low-volume outcomes, particularly for players near the bottom of the depth chart. Simply throwing a wide distribution around projected receiving yards didn't work particularly well.
The current version models whether a player receives targets, his target volume when involved, catches and the resulting yardage distribution.
We also learned that very low projected receiving volume is its own animal. The model tracks different applicability ranges rather than pretending we're equally confident modeling Ja'Marr Chase and some TE3 projected for one catch.
Then we simulate the game 50,000 times
Once we have the model for a player and projection source, we run 50,000 simulations.
That gives us an estimated distribution instead of one number.
We can see:
- average simulated yards
- median simulated yards
- probability of going Over
- probability of going Under
- probability of landing exactly on an integer line
For our Week 1 passing-yard simulations, the four projection sources had average gaps between their projected mean and simulated median of roughly:
- 4for4: 9.7 yards
- FantasyPros: 9.9 yards
- FTN: 10.2 yards
- PFF: 10.0 yards
Kind of a big deal.
A quarterback projected for 250 passing yards might have a simulated median around 240. Comparing 250 directly to a sportsbook line of 245 would tell you to look Over.
The distribution can tell you the opposite, and we've already seen plenty of cases where exactly that happens.
Sportsbook price matters too
Once we have our probabilities, we match them against actual sportsbook markets.
And we care about the price, not just the line.
Over 250.5 at -110 and Over 250.5 at +105 are obviously different bets.
We take the prices on both sides of a sportsbook's market, convert them into implied probabilities, and remove the vig. That gives us an estimate of the probability represented by the market itself.
Then we compare that with each projection source's simulated probability.
For example, suppose a sportsbook's no-vig price implies:
Under 250.5: 51%
And our four simulations produce:
FantasyPros: 59%
PFF: 61%
FTN: 58%
4for4: 60%
Now we have something worth looking at.
What actually qualifies as a bet?
For our first version, we're being fairly conservative.
A bet needs either:
4 of 4 models agreeing
or
3 of 4 models agreeing, with all four models having enough information to make a real decision
That second distinction matters. If three models say Under and the fourth doesn't have enough data, we don't count that as 3-of-4 agreement. We want the fourth model to actually disagree.
Then there's another hurdle.
Every model we're counting as agreeing needs to beat the sportsbook's no-vig probability by at least 5 percentage points.
We use the weakest agreeing model for this test.
So if four models differ from the market by:
+9.2, +8.1, +6.4 and +3.8 percentage points
there's no bet.
The average looks good. Three models look good. The fourth doesn't clear our requirement.
For a 3-of-4 play, all three models on the chosen side have to clear 5 percentage points.
We're starting there and leaving it alone. We don't want to see how Week 1 goes and decide afterward that 4% or 6% would have been the perfect cutoff. That's a great way to build a model that was awesome at predicting games that already happened.
We shop the line
If something qualifies, we check the sportsbooks available to us in Massachusetts and take the best line.
It sounds obvious, but one yard can make a meaningful difference in the simulated probability. Over hundreds of bets, consistently taking worse numbers is just lighting money on fire.
Every play is also a flat one unit.
No five-star plays. No "max whale nuclear lock of the century." No pretending we know our probability estimates precisely enough to start firing 4.73 goddamn units.
Why are there so many Unders?
You may notice this pretty quickly.
Our first Week 1 card has 26 qualifying plays. 23 are Unders and only three are Overs.
I'm not particularly bothered by that.
Remember where we started: fantasy projections are estimates of the mean, and yardage distributions tend to have long right tails.
Big games drag the average upward.
Once we turn those means into full distributions, the median frequently moves lower. That's especially obvious in passing yards, where our Week 1 simulations put the median roughly 10 yards below the projection across all four sources.
If this process were spitting out overwhelmingly more Overs, I'd be asking a lot more questions.
That doesn't mean "Unders are better bets." It means the direction makes sense given how the model is constructed.
Does this mean we've cracked player props?
Fuck no.
This is the part I'm probably most interested in following this season.
We've done historical validation on the distributions themselves. We've tested whether the simulations reproduce real-world variance, median behavior, tails and other characteristics. We've deliberately held data out while making modeling decisions.
What we haven't established is that these probability differences translate into profitable bets.
That's what 2026 is for.
Week 1 is our first real prospective test. The models were built, the projections and odds were frozen, and the betting rules were established before we knew the results.
From here, we get to find out.
Maybe the 5-point requirement is too loose. Maybe it's too strict. Maybe 4-of-4 agreement crushes while 3-of-4 doesn't add anything. Maybe passing yards works and receiving yards gets murdered. Maybe the whole thing loses money.
All of those are useful answers.
The important part is that we're keeping the receipts.
Every projection, sportsbook line, price, simulated probability and qualifying bet gets saved. At the end of the season, we'll have a dataset that didn't exist when we started this project.
Then we can make the model better based on what actually happened instead of what we wish had happened.
