Towards the Efficient Inference by Incorporating Automated Computational Phenotypes under Covariate Shift

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ying, Chao, Jin, Jun, Guo, Yi, Li, Xiudi, Liang, Muxuan, Zhao, Jiwei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913864423047168
author Ying, Chao
Jin, Jun
Guo, Yi
Li, Xiudi
Liang, Muxuan
Zhao, Jiwei
author_facet Ying, Chao
Jin, Jun
Guo, Yi
Li, Xiudi
Liang, Muxuan
Zhao, Jiwei
contents Collecting gold-standard phenotype data via manual extraction is typically labor-intensive and slow, whereas automated computational phenotypes (ACPs) offer a systematic and much faster alternative. However, simply replacing the gold-standard with ACPs, without acknowledging their differences, could lead to biased results and misleading conclusions. Motivated by the complexity of incorporating ACPs while maintaining the validity of downstream analyses, in this paper, we consider a semi-supervised learning setting that consists of both labeled data (with gold-standard) and unlabeled data (without gold-standard), under the covariate shift framework. We develop doubly robust and semiparametrically efficient estimators that leverage ACPs for general target parameters in the unlabeled and combined populations. In addition, we carefully analyze the efficiency gains achieved by incorporating ACPs, comparing scenarios with and without their inclusion. Notably, we identify that ACPs for the unlabeled data, instead of for the labeled data, drive the enhanced efficiency gains. To validate our theoretical findings, we conduct comprehensive synthetic experiments and apply our method to multiple real-world datasets, confirming the practical advantages of our approach. \hfill{\texttt{Code}: \href{https://github.com/brucejunjin/ICML2025-ACPCS}{\faGithub}}
format Preprint
id arxiv_https___arxiv_org_abs_2505_22632
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards the Efficient Inference by Incorporating Automated Computational Phenotypes under Covariate Shift
Ying, Chao
Jin, Jun
Guo, Yi
Li, Xiudi
Liang, Muxuan
Zhao, Jiwei
Methodology
Collecting gold-standard phenotype data via manual extraction is typically labor-intensive and slow, whereas automated computational phenotypes (ACPs) offer a systematic and much faster alternative. However, simply replacing the gold-standard with ACPs, without acknowledging their differences, could lead to biased results and misleading conclusions. Motivated by the complexity of incorporating ACPs while maintaining the validity of downstream analyses, in this paper, we consider a semi-supervised learning setting that consists of both labeled data (with gold-standard) and unlabeled data (without gold-standard), under the covariate shift framework. We develop doubly robust and semiparametrically efficient estimators that leverage ACPs for general target parameters in the unlabeled and combined populations. In addition, we carefully analyze the efficiency gains achieved by incorporating ACPs, comparing scenarios with and without their inclusion. Notably, we identify that ACPs for the unlabeled data, instead of for the labeled data, drive the enhanced efficiency gains. To validate our theoretical findings, we conduct comprehensive synthetic experiments and apply our method to multiple real-world datasets, confirming the practical advantages of our approach. \hfill{\texttt{Code}: \href{https://github.com/brucejunjin/ICML2025-ACPCS}{\faGithub}}
title Towards the Efficient Inference by Incorporating Automated Computational Phenotypes under Covariate Shift
topic Methodology
url https://arxiv.org/abs/2505.22632