Gradient Boosted Mixed Models: Flexible Joint Estimation of Mean and Variance Components for Clustered Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prevett, Mitchell L., Hui, Francis K. C., Tho, Zhi Yang, Welsh, A. H., Westveld, Anton H.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911244179472384
author Prevett, Mitchell L.
Hui, Francis K. C.
Tho, Zhi Yang
Welsh, A. H.
Westveld, Anton H.
author_facet Prevett, Mitchell L.
Hui, Francis K. C.
Tho, Zhi Yang
Welsh, A. H.
Westveld, Anton H.
contents Linear mixed models are widely used for clustered data, but their reliance on parametric forms limits flexibility in complex and high-dimensional settings. In contrast, gradient boosting methods achieve high predictive accuracy through nonparametric estimation, but do not accommodate clustered data structures or provide uncertainty quantification. We introduce Gradient Boosted Mixed Models (GBMixed), a framework and algorithm that extends boosting to jointly estimate mean and variance components via likelihood-based gradients. In addition to nonparametric mean estimation, the method models both random effects and residual variances as potentially covariate-dependent functions using flexible base learners such as regression trees or splines, enabling nonparametric estimation while maintaining interpretability. Simulations and real-world applications demonstrate accurate recovery of variance components, calibrated prediction intervals, and improved predictive accuracy relative to standard linear mixed models and nonparametric methods. GBMixed provides heteroscedastic uncertainty quantification and introduces boosting for heterogeneous random effects. This enables covariate-dependent shrinkage for cluster-specific predictions to adapt between population and cluster-level data. Under standard causal assumptions, the framework enables estimation of heterogeneous treatment effects with reliable uncertainty quantification.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gradient Boosted Mixed Models: Flexible Joint Estimation of Mean and Variance Components for Clustered Data
Prevett, Mitchell L.
Hui, Francis K. C.
Tho, Zhi Yang
Welsh, A. H.
Westveld, Anton H.
Machine Learning
Computation
Methodology
Linear mixed models are widely used for clustered data, but their reliance on parametric forms limits flexibility in complex and high-dimensional settings. In contrast, gradient boosting methods achieve high predictive accuracy through nonparametric estimation, but do not accommodate clustered data structures or provide uncertainty quantification. We introduce Gradient Boosted Mixed Models (GBMixed), a framework and algorithm that extends boosting to jointly estimate mean and variance components via likelihood-based gradients. In addition to nonparametric mean estimation, the method models both random effects and residual variances as potentially covariate-dependent functions using flexible base learners such as regression trees or splines, enabling nonparametric estimation while maintaining interpretability. Simulations and real-world applications demonstrate accurate recovery of variance components, calibrated prediction intervals, and improved predictive accuracy relative to standard linear mixed models and nonparametric methods. GBMixed provides heteroscedastic uncertainty quantification and introduces boosting for heterogeneous random effects. This enables covariate-dependent shrinkage for cluster-specific predictions to adapt between population and cluster-level data. Under standard causal assumptions, the framework enables estimation of heterogeneous treatment effects with reliable uncertainty quantification.
title Gradient Boosted Mixed Models: Flexible Joint Estimation of Mean and Variance Components for Clustered Data
topic Machine Learning
Computation
Methodology
url https://arxiv.org/abs/2511.00217