Generalization of Medical Large Language Models through Cross-Domain Weak Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Long, Robert, Gonzalez, Eric, Fuller, Harrison
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910810177011712
author Long, Robert
Gonzalez, Eric
Fuller, Harrison
author_facet Long, Robert
Gonzalez, Eric
Fuller, Harrison
contents The advancement of large language models (LLMs) has opened new frontiers in natural language processing, particularly in specialized domains like healthcare. In this paper, we propose the Incremental Curriculum-Based Fine-Tuning (ICFT) framework to enhance the generative capabilities of medical large language models (MLLMs). ICFT combines curriculum-based learning, dual-stage memory coordination, and parameter-efficient fine-tuning to enable a progressive transition from general linguistic knowledge to strong domain-specific expertise. Experimental results across diverse medical NLP tasks, including question answering, preference classification, and response generation, demonstrate that ICFT consistently outperforms state-of-the-art baselines, achieving improvements in both accuracy and efficiency. Further analysis reveals the framework's ability to generalize to unseen data, reduce errors, and deliver diverse, contextually relevant medical responses. These findings establish ICFT as a robust and scalable solution for adapting LLMs to the medical domain, offering practical benefits for real-world healthcare applications.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalization of Medical Large Language Models through Cross-Domain Weak Supervision
Long, Robert
Gonzalez, Eric
Fuller, Harrison
Computation and Language
The advancement of large language models (LLMs) has opened new frontiers in natural language processing, particularly in specialized domains like healthcare. In this paper, we propose the Incremental Curriculum-Based Fine-Tuning (ICFT) framework to enhance the generative capabilities of medical large language models (MLLMs). ICFT combines curriculum-based learning, dual-stage memory coordination, and parameter-efficient fine-tuning to enable a progressive transition from general linguistic knowledge to strong domain-specific expertise. Experimental results across diverse medical NLP tasks, including question answering, preference classification, and response generation, demonstrate that ICFT consistently outperforms state-of-the-art baselines, achieving improvements in both accuracy and efficiency. Further analysis reveals the framework's ability to generalize to unseen data, reduce errors, and deliver diverse, contextually relevant medical responses. These findings establish ICFT as a robust and scalable solution for adapting LLMs to the medical domain, offering practical benefits for real-world healthcare applications.
title Generalization of Medical Large Language Models through Cross-Domain Weak Supervision
topic Computation and Language
url https://arxiv.org/abs/2502.00832