Super Ensemble Learning Using the Highly-Adaptive-Lasso

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Zeyi, Zhang, Wenxin, Caffo, Brian S, Lindquist, Martin, van der Laan, Mark
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911075424796672
author Wang, Zeyi
Zhang, Wenxin
Caffo, Brian S
Lindquist, Martin
van der Laan, Mark
author_facet Wang, Zeyi
Zhang, Wenxin
Caffo, Brian S
Lindquist, Martin
van der Laan, Mark
contents We introduce the Meta Highly-Adaptive-Lasso Minimum Loss Estimator (M-HAL-MLE), a novel ensemble approach for estimating functional parameters of realistically modeled data distribution from independent and identically distributed observations. Given $J$ initial estimators, candidate ensembles are generated by finite-sectional-variation cadlag functions. Using $V$-fold cross-validation, the M-HAL-MLE selects the optimal cadlag ensemble minimizing the cross-validated empirical risk, with the sectional variation bound as a tuning parameter. The final estimator, M-HAL super-learner, is obtained by averaging ensemble compositions across folds. In contrast, the oracle ensemble and oracle estimator are defined by minimizing the population excess risk relative to the true function. We establish following theoretical properties: 1) the M-HAL super-learner converges to the oracle estimator at rate $n^{-2/3}$ in excess risk, up to log-n factors; 2) by appropriate undersmoothing, target features of the M-HAL super-learner are asymptotically linear for corresponding target features of the oracle estimator; 3) the excess risk between the oracle estimator and true function, along with the difference between their target features, is generally second-order. Simulations validate the theoretical results, demonstrating effectiveness in high-dimensional settings. We further illustrate the method in a real-data application involving mediation analysis of functional MRI from human pain studies.
format Preprint
id arxiv_https___arxiv_org_abs_2312_16953
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Super Ensemble Learning Using the Highly-Adaptive-Lasso
Wang, Zeyi
Zhang, Wenxin
Caffo, Brian S
Lindquist, Martin
van der Laan, Mark
Methodology
We introduce the Meta Highly-Adaptive-Lasso Minimum Loss Estimator (M-HAL-MLE), a novel ensemble approach for estimating functional parameters of realistically modeled data distribution from independent and identically distributed observations. Given $J$ initial estimators, candidate ensembles are generated by finite-sectional-variation cadlag functions. Using $V$-fold cross-validation, the M-HAL-MLE selects the optimal cadlag ensemble minimizing the cross-validated empirical risk, with the sectional variation bound as a tuning parameter. The final estimator, M-HAL super-learner, is obtained by averaging ensemble compositions across folds. In contrast, the oracle ensemble and oracle estimator are defined by minimizing the population excess risk relative to the true function. We establish following theoretical properties: 1) the M-HAL super-learner converges to the oracle estimator at rate $n^{-2/3}$ in excess risk, up to log-n factors; 2) by appropriate undersmoothing, target features of the M-HAL super-learner are asymptotically linear for corresponding target features of the oracle estimator; 3) the excess risk between the oracle estimator and true function, along with the difference between their target features, is generally second-order. Simulations validate the theoretical results, demonstrating effectiveness in high-dimensional settings. We further illustrate the method in a real-data application involving mediation analysis of functional MRI from human pain studies.
title Super Ensemble Learning Using the Highly-Adaptive-Lasso
topic Methodology
url https://arxiv.org/abs/2312.16953