Moving from Machine Learning to Statistics: the case of Expected Points in American football

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brill, Ryan S., Yee, Ryan, Deshpande, Sameer K., Wyner, Abraham J.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909308595208192
author Brill, Ryan S.
Yee, Ryan
Deshpande, Sameer K.
Wyner, Abraham J.
author_facet Brill, Ryan S.
Yee, Ryan
Deshpande, Sameer K.
Wyner, Abraham J.
contents Expected points is a value function fundamental to player evaluation and strategic in-game decision-making across sports analytics, particularly in American football. To estimate expected points, football analysts use machine learning tools, which are not equipped to handle certain challenges. They suffer from selection bias, display counter-intuitive artifacts of overfitting, do not quantify uncertainty in point estimates, and do not account for the strong dependence structure of observational football data. These issues are not unique to American football or even sports analytics; they are general problems analysts encounter across various statistical applications, particularly when using machine learning in lieu of traditional statistical models. We explore these issues in detail and devise expected points models that account for them. We also introduce a widely applicable novel methodological approach to mitigate overfitting, using a catalytic prior to smooth our machine learning models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04889
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Moving from Machine Learning to Statistics: the case of Expected Points in American football
Brill, Ryan S.
Yee, Ryan
Deshpande, Sameer K.
Wyner, Abraham J.
Applications
Expected points is a value function fundamental to player evaluation and strategic in-game decision-making across sports analytics, particularly in American football. To estimate expected points, football analysts use machine learning tools, which are not equipped to handle certain challenges. They suffer from selection bias, display counter-intuitive artifacts of overfitting, do not quantify uncertainty in point estimates, and do not account for the strong dependence structure of observational football data. These issues are not unique to American football or even sports analytics; they are general problems analysts encounter across various statistical applications, particularly when using machine learning in lieu of traditional statistical models. We explore these issues in detail and devise expected points models that account for them. We also introduce a widely applicable novel methodological approach to mitigate overfitting, using a catalytic prior to smooth our machine learning models.
title Moving from Machine Learning to Statistics: the case of Expected Points in American football
topic Applications
url https://arxiv.org/abs/2409.04889