Surrogate-Powered Inference: Regularization and Adaptivity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jianmin, Wang, Huiyuan, Lumley, Thomas, Dai, Xiaowu, Chen, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917169769480192
author Chen, Jianmin
Wang, Huiyuan
Lumley, Thomas
Dai, Xiaowu
Chen, Yong
author_facet Chen, Jianmin
Wang, Huiyuan
Lumley, Thomas
Dai, Xiaowu
Chen, Yong
contents High-quality labeled data are essential for reliable statistical inference, but are often limited by validation costs. While surrogate labels provide cost-effective alternatives, their noise can introduce non-negligible bias. To address this challenge, we propose the surrogate-powered inference (SPI) toolbox, a unified framework that leverages both the validity of high-quality labels and the abundance of surrogates to enable reliable statistical inference. SPI comprises three progressively enhanced versions. Base-SPI integrates validated labels and surrogates through augmentation to improve estimation efficiency. SPI+ incorporates regularized regression to safely handle multiple surrogates, preventing performance degradation due to error accumulation. SPI++ further optimizes efficiency under limited validation budgets through an adaptive, multiwave labeling procedure that prioritizes informative subjects for labeling. Compared to traditional methods, SPI substantially reduces the estimation error and increases the power in risk factor identification. These results demonstrate the value of SPI in improving the reproducibility. Theoretical guarantees and extensive simulation studies further illustrate the properties of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21826
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Surrogate-Powered Inference: Regularization and Adaptivity
Chen, Jianmin
Wang, Huiyuan
Lumley, Thomas
Dai, Xiaowu
Chen, Yong
Methodology
High-quality labeled data are essential for reliable statistical inference, but are often limited by validation costs. While surrogate labels provide cost-effective alternatives, their noise can introduce non-negligible bias. To address this challenge, we propose the surrogate-powered inference (SPI) toolbox, a unified framework that leverages both the validity of high-quality labels and the abundance of surrogates to enable reliable statistical inference. SPI comprises three progressively enhanced versions. Base-SPI integrates validated labels and surrogates through augmentation to improve estimation efficiency. SPI+ incorporates regularized regression to safely handle multiple surrogates, preventing performance degradation due to error accumulation. SPI++ further optimizes efficiency under limited validation budgets through an adaptive, multiwave labeling procedure that prioritizes informative subjects for labeling. Compared to traditional methods, SPI substantially reduces the estimation error and increases the power in risk factor identification. These results demonstrate the value of SPI in improving the reproducibility. Theoretical guarantees and extensive simulation studies further illustrate the properties of our approach.
title Surrogate-Powered Inference: Regularization and Adaptivity
topic Methodology
url https://arxiv.org/abs/2512.21826