A Federated and Parameter-Efficient Framework for Large Language Model Training in Medicine

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Anran, Chen, Yuanyuan, Long, Wenjun, Yin, Yu, Hu, Yan, Kim, Hyunjae, Zhou, Weipeng, Zhou, Yujia, Peng, Hongyi, Ren, Yang, Ai, Xuguang, Qin, Zhenyue, Hu, Ming, Li, Xiaoxiao, Yu, Han, Tham, Yih-Chung, Ohno-Machado, Lucila, Xu, Hua, Chen, Qingyu
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910005036318720
author Li, Anran
Chen, Yuanyuan
Long, Wenjun
Yin, Yu
Hu, Yan
Kim, Hyunjae
Zhou, Weipeng
Zhou, Yujia
Peng, Hongyi
Ren, Yang
Ai, Xuguang
Qin, Zhenyue
Hu, Ming
Li, Xiaoxiao
Yu, Han
Tham, Yih-Chung
Ohno-Machado, Lucila
Xu, Hua
Chen, Qingyu
author_facet Li, Anran
Chen, Yuanyuan
Long, Wenjun
Yin, Yu
Hu, Yan
Kim, Hyunjae
Zhou, Weipeng
Zhou, Yujia
Peng, Hongyi
Ren, Yang
Ai, Xuguang
Qin, Zhenyue
Hu, Ming
Li, Xiaoxiao
Yu, Han
Tham, Yih-Chung
Ohno-Machado, Lucila
Xu, Hua
Chen, Qingyu
contents Large language models (LLMs) have demonstrated strong performance on medical benchmarks, including question answering and diagnosis. To enable their use in clinical settings, LLMs are typically further adapted through continued pretraining or post-training using clinical data. However, most medical LLMs are trained on data from a single institution, which faces limitations in generalizability and safety in heterogeneous systems. Federated learning (FL) is a promising solution for enabling collaborative model development across healthcare institutions. Yet applying FL to LLMs in medicine remains fundamentally limited. First, conventional FL requires transmitting the full model during each communication round, which becomes impractical for multi-billion-parameter LLMs given the limited computational resources. Second, many FL algorithms implicitly assume data homogeneity, whereas real-world clinical data are highly heterogeneous across patients, diseases, and institutional practices. We introduce the model-agnostic and parameter-efficient federated learning framework for adapting LLMs to medical applications. Fed-MedLoRA transmits only low-rank adapter parameters, reducing communication and computation overhead, while Fed-MedLoRA+ further incorporates adaptive, data-aware aggregation to improve convergence under cross-site heterogeneity. We apply the framework to clinical information extraction (IE), which transforms patient narratives into structured medical entities and relations. Accuracy was assessed across five patient cohorts through comparisons with BERT models, and LLaMA-3 and DeepSeek-R1, GPT-4o models. Evaluation settings included (1) in-domain training and testing, (2) external validation on independent cohorts, and (3) a low-resource new-site adaptation scenario using real-world clinical notes from the Yale New Haven Health System.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22124
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Federated and Parameter-Efficient Framework for Large Language Model Training in Medicine
Li, Anran
Chen, Yuanyuan
Long, Wenjun
Yin, Yu
Hu, Yan
Kim, Hyunjae
Zhou, Weipeng
Zhou, Yujia
Peng, Hongyi
Ren, Yang
Ai, Xuguang
Qin, Zhenyue
Hu, Ming
Li, Xiaoxiao
Yu, Han
Tham, Yih-Chung
Ohno-Machado, Lucila
Xu, Hua
Chen, Qingyu
Computation and Language
Distributed, Parallel, and Cluster Computing
Large language models (LLMs) have demonstrated strong performance on medical benchmarks, including question answering and diagnosis. To enable their use in clinical settings, LLMs are typically further adapted through continued pretraining or post-training using clinical data. However, most medical LLMs are trained on data from a single institution, which faces limitations in generalizability and safety in heterogeneous systems. Federated learning (FL) is a promising solution for enabling collaborative model development across healthcare institutions. Yet applying FL to LLMs in medicine remains fundamentally limited. First, conventional FL requires transmitting the full model during each communication round, which becomes impractical for multi-billion-parameter LLMs given the limited computational resources. Second, many FL algorithms implicitly assume data homogeneity, whereas real-world clinical data are highly heterogeneous across patients, diseases, and institutional practices. We introduce the model-agnostic and parameter-efficient federated learning framework for adapting LLMs to medical applications. Fed-MedLoRA transmits only low-rank adapter parameters, reducing communication and computation overhead, while Fed-MedLoRA+ further incorporates adaptive, data-aware aggregation to improve convergence under cross-site heterogeneity. We apply the framework to clinical information extraction (IE), which transforms patient narratives into structured medical entities and relations. Accuracy was assessed across five patient cohorts through comparisons with BERT models, and LLaMA-3 and DeepSeek-R1, GPT-4o models. Evaluation settings included (1) in-domain training and testing, (2) external validation on independent cohorts, and (3) a low-resource new-site adaptation scenario using real-world clinical notes from the Yale New Haven Health System.
title A Federated and Parameter-Efficient Framework for Large Language Model Training in Medicine
topic Computation and Language
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2601.22124