Statistical Inference for Gradient Boosting Regression

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fang, Haimo, Tan, Kevin, Hooker, Giles
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918149682626560
author Fang, Haimo
Tan, Kevin
Hooker, Giles
author_facet Fang, Haimo
Tan, Kevin
Hooker, Giles
contents Gradient boosting is widely popular due to its flexibility and predictive accuracy. However, statistical inference and uncertainty quantification for gradient boosting remain challenging and under-explored. We propose a unified framework for statistical inference in gradient boosting regression. Our framework integrates dropout or parallel training with a recently proposed regularization procedure that allows for a central limit theorem (CLT) for boosting. With these enhancements, we surprisingly find that increasing the dropout rate and the number of trees grown in parallel at each iteration substantially enhances signal recovery and overall performance. Our resulting algorithms enjoy similar CLTs, which we use to construct built-in confidence intervals, prediction intervals, and rigorous hypothesis tests for assessing variable importance. Numerical experiments demonstrate that our algorithms perform well, interpolate between regularized boosting and random forests, and confirm the validity of their built-in statistical inference procedures.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Statistical Inference for Gradient Boosting Regression
Fang, Haimo
Tan, Kevin
Hooker, Giles
Machine Learning
Statistics Theory
Methodology
Gradient boosting is widely popular due to its flexibility and predictive accuracy. However, statistical inference and uncertainty quantification for gradient boosting remain challenging and under-explored. We propose a unified framework for statistical inference in gradient boosting regression. Our framework integrates dropout or parallel training with a recently proposed regularization procedure that allows for a central limit theorem (CLT) for boosting. With these enhancements, we surprisingly find that increasing the dropout rate and the number of trees grown in parallel at each iteration substantially enhances signal recovery and overall performance. Our resulting algorithms enjoy similar CLTs, which we use to construct built-in confidence intervals, prediction intervals, and rigorous hypothesis tests for assessing variable importance. Numerical experiments demonstrate that our algorithms perform well, interpolate between regularized boosting and random forests, and confirm the validity of their built-in statistical inference procedures.
title Statistical Inference for Gradient Boosting Regression
topic Machine Learning
Statistics Theory
Methodology
url https://arxiv.org/abs/2509.23127