Evaluating Models (General Thoughts)
When we build models, we apply standard evaluation metrics against a holdout dataset (data not used in any action related to the model building) to evaluate performance. Once we have what we think is “good” performance, we compare the results of the players in our hold out dataset to their actual fantasy performance and the fantasy performance of 2 other sites, ESPN and 4for4. The overall dataset we have for ESPN and 4for4 is not complete, so we filter down to just the players that exist in all three sets of data.
Opinion on Transparency
Throughout the season, I generate a weekly postmortem (when I have time) so I can address both the successes and shortcomings of our model for the previous week.
My goal of those posts isn’t to cover up the poor performance or overhype successes. I try to highlight both. We feel our process is stable and the results, throughout the whole season, will be better than other sites or just using one single website. There is going to be a week to week variance in the model performance, that is unavoidable. Football is a sport that has lots of randomness in it.
I am always on our slack channel to answer questions about the current projections or process (link is at the bottom of this post).
Analysis on the Meta Data
I’ve spent a lot of time (probably too much) analyzing variance in our projections. Below are general trends that are apparent in the data:
Our models don’t overreact to one massive game
- Zach Ertz Week 10 2018 is an example:
- 4for4 generally undervalued Ertz all year and never projected over 10 points, which I think is insane.
- ESPN increased its next week projection by 42%, which seems like an over-reaction.
- All our projections include both upside and downside projections, so we can try to measure the breakout potential of a player without arbitrarily inflating projections based on a best case outcome, which is extremely difficult to predict.
- Zach Ertz Week 10 2018 is an example:
Our models gradually increase projected performance when a player's situation materially changes.
- For example, Damien Williams was slow to see a boost to his projections last year.
Comparision is strictly on PPR
Everyone is pretty close; the net difference is <1% on our test data. I am going to break it down by Position and add details per player in the next section. Admittedly, we do not have the highest overall score, driven by the performance of QBs and RBs (Drew Brees and LeSean McCoy).
Last year we were ~ 22% better over the whole season when factoring in positional ranks. This post is only evaluating projected points.
The details will help you get a feeling for what value we end up adding over other sites.
- We are best in TE and WR, ESPN is best in QB and 4for4 is best in RB.
- Overall totals do not have us as the highest site, which isn’t a great story in the aggregate but in general, the difference is within a the margin of error, <1% and mainly due to low rankings of Drew Brees (see QB section below)
- We’ve only been making projections for a year, while ESPN and 4for4 have been around for years.
- In my (biased) opinion, <1% in our projections vs. theirs shows their lack of innovation or a lack of evaluation on their projections vs. their competition. When comparing week to week over the whole season last year (including every player), we beat both of them most weeks.
The dataset we are using to compare, as noted above, is only the players and weeks that exist in the data I gathered from ESPN and 4for4 and are in the test set for my model.
During the season I’ll be able to evaluate all the players each site ranked to give a holistic view of performance.
Details for the Tables Below
Site Pt Diff = Point difference between predicted values and actual
Position = Position the point difference applies to, Total is the the sum of all positions. The lower the score, the better.
Total Samples = Total weeks evaluated for the position
QB Comparision
- Drew Brees drove the most significant variance for us relative to ESPN.
- Our model is currently too low on him every week, probably related to his age.
- I could quickly improve this by automatically adjusting his points every week to be more but that goes against what we are trying to do
- This is something we are going to work to improve throughout the season
RB Comparision
- LeSean McCoy was generally punished the whole season and we consistently projected lower point totals than ESPN and 4for4.
- I don’t think that is unreasonable.
- Our model was slow to pick up on Damien Williams
- Basically, due to a limitation in the code, the model has a hard time identifying starting players when it’s not a clean break.
- I’ve noticed it takes a few games of similar performance before it buys a player’s projections should materially increase.
- In the app, we have adjustments for this scenario and our projections for Damien Williams would have been adjusted up for the change in depth chart position, I did not simulate that in these results.
WR Comparision
For the fantasy-relevant players in this grouping, there are no real takeaways.
<table> <thead> <tr> <th> full\_name </th> <th> Position </th> <th> Gridiron Pt Diff </th> <th> ESPN Pt Diff </th> <th> 4for4 Pt Diff </th> </tr> </thead> <tbody> <tr> <td> Adam Humphries </td> <td> WR </td> <td> 51 </td> <td> 54 </td> <td> 66 </td> </tr> <tr> <td> Adam Thielen </td> <td> WR </td> <td> 46 </td> <td> 51 </td> <td> 52 </td> </tr> <tr> <td> Amari Cooper </td> <td> WR </td> <td> 89 </td> <td> 82 </td> <td> 91 </td> </tr> <tr> <td> Antonio Brown </td> <td> WR </td> <td> 32 </td> <td> 35 </td> <td> 46 </td> </tr> <tr> <td> Bennie Fowler </td> <td> WR </td> <td> 11 </td> <td> 9 </td> <td> 9 </td> </tr> <tr> <td> Brice Butler </td> <td> WR </td> <td> 6 </td> <td> 7 </td> <td> 5 </td> </tr> <tr> <td> Calvin Ridley </td> <td> WR </td> <td> 46 </td> <td> 42 </td> <td> 40 </td> </tr> <tr> <td> Corey Davis </td> <td> WR </td> <td> 55 </td> <td> 52 </td> <td> 52 </td> </tr> <tr> <td> DaeSean Hamilton </td> <td> WR </td> <td> 6 </td> <td> 3 </td> <td> 6 </td> </tr> <tr> <td> Dante Pettis </td> <td> WR </td> <td> 26 </td> <td> 24 </td> <td> 32 </td> </tr> <tr> <td> Darius Jennings </td> <td> WR </td> <td> 2 </td> <td> 3 </td> <td> 2 </td> </tr> <tr> <td> Geronimo Allison </td> <td> WR </td> <td> 9 </td> <td> 8 </td> <td> 4 </td> </tr> <tr> <td> Golden Tate </td> <td> WR </td> <td> 35 </td> <td> 34 </td> <td> 20 </td> </tr> <tr> <td> Jermaine Kearse </td> <td> WR </td> <td> 35 </td> <td> 49 </td> <td> 27 </td> </tr> <tr> <td> Jordy Nelson </td> <td> WR </td> <td> 35 </td> <td> 40 </td> <td> 34 </td> </tr> <tr> <td> Marquise Goodwin </td> <td> WR </td> <td> 22 </td> <td> 23 </td> <td> 21 </td> </tr> <tr> <td> Michael Crabtree </td> <td> WR </td> <td> 35 </td> <td> 35 </td> <td> 22 </td> </tr> <tr> <td> Michael Floyd </td> <td> WR </td> <td> 2 </td> <td> 3 </td> <td> 0 </td> </tr> <tr> <td> Richie James </td> <td> WR </td> <td> 6 </td> <td> 8 </td> <td> 6 </td> </tr> <tr> <td> T.Y. Hilton </td> <td> WR </td> <td> 64 </td> <td> 59 </td> <td> 81 </td> </tr> <tr> <td> Tajae Sharpe </td> <td> WR </td> <td> 52 </td> <td> 51 </td> <td> 40 </td> </tr> <tr> <td> Trey Quinn </td> <td> WR </td> <td> 8 </td> <td> 7 </td> <td> 8 </td> </tr> <tr> <td> Tyler Boyd </td> <td> WR </td> <td> 46 </td> <td> 45 </td> <td> 47 </td> </tr> <tr> <td> Zach Pascal </td> <td> WR </td> <td> 20 </td> <td> 11 </td> <td> 7 </td> </tr> </tbody> </table>TE Comparision
- Nothing major to report here!
Questions/Feedback
If anyone has any questions or wants to talk to us, you can find us on in our Slack Channel. Click here to join: PLEASE CLICK ME, WE LOVE TO TALK.