Scalable and Distributed Individualized Treatment Rules for Massive Datasets

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Qiao, Nan, Li, Wangcheng, Zhang, Jingxiao, Chen, Canyi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915606127706112
author Qiao, Nan
Li, Wangcheng
Zhang, Jingxiao
Chen, Canyi
author_facet Qiao, Nan
Li, Wangcheng
Zhang, Jingxiao
Chen, Canyi
contents Synthesizing information from multiple data sources is crucial for constructing accurate individualized treatment rules (ITRs). However, privacy concerns often present significant barriers to the integrative analysis of such multi-source data. Classical meta-learning, which averages local estimates to derive the final ITR, is frequently suboptimal due to biases in these local estimates. To address these challenges, we propose a convolution-smoothed weighted support vector machine for learning the optimal ITR. The accompanying loss function is both convex and smooth, which allows us to develop an efficient multi-round distributed learning procedure for ITRs. Such distributed learning ensures optimal statistical performance with a fixed number of communication rounds, thereby minimizing coordination costs across data centers while preserving data privacy. Our method avoids pooling subject-level raw data and instead requires only sharing summary statistics. Additionally, we develop an efficient coordinate gradient descent algorithm, which guarantees at least linear convergence for the resulting optimization problem. Extensive simulations and an application to sepsis treatment across multiple intensive care units validate the effectiveness of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scalable and Distributed Individualized Treatment Rules for Massive Datasets
Qiao, Nan
Li, Wangcheng
Zhang, Jingxiao
Chen, Canyi
Methodology
Computation
Synthesizing information from multiple data sources is crucial for constructing accurate individualized treatment rules (ITRs). However, privacy concerns often present significant barriers to the integrative analysis of such multi-source data. Classical meta-learning, which averages local estimates to derive the final ITR, is frequently suboptimal due to biases in these local estimates. To address these challenges, we propose a convolution-smoothed weighted support vector machine for learning the optimal ITR. The accompanying loss function is both convex and smooth, which allows us to develop an efficient multi-round distributed learning procedure for ITRs. Such distributed learning ensures optimal statistical performance with a fixed number of communication rounds, thereby minimizing coordination costs across data centers while preserving data privacy. Our method avoids pooling subject-level raw data and instead requires only sharing summary statistics. Additionally, we develop an efficient coordinate gradient descent algorithm, which guarantees at least linear convergence for the resulting optimization problem. Extensive simulations and an application to sepsis treatment across multiple intensive care units validate the effectiveness of the proposed method.
title Scalable and Distributed Individualized Treatment Rules for Massive Datasets
topic Methodology
Computation
url https://arxiv.org/abs/2511.05842