WiNGPT-3.0 Technical Report

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhuang, Boqin, Song, Chenxiao, Lu, Huitong, Qiao, Jiacheng, Liu, Mingqian, Yu, Mingxing, Hong, Ping, Li, Rui, Song, Xiaoxia, Xu, Xiangjun, Chen, Xu, Ma, Yaoyao, Gao, Yujie
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908394061824000
author Zhuang, Boqin
Song, Chenxiao
Lu, Huitong
Qiao, Jiacheng
Liu, Mingqian
Yu, Mingxing
Hong, Ping
Li, Rui
Song, Xiaoxia
Xu, Xiangjun
Chen, Xu
Ma, Yaoyao
Gao, Yujie
author_facet Zhuang, Boqin
Song, Chenxiao
Lu, Huitong
Qiao, Jiacheng
Liu, Mingqian
Yu, Mingxing
Hong, Ping
Li, Rui
Song, Xiaoxia
Xu, Xiangjun
Chen, Xu
Ma, Yaoyao
Gao, Yujie
contents Current Large Language Models (LLMs) exhibit significant limitations, notably in structured, interpretable, and verifiable medical reasoning, alongside practical deployment challenges related to computational resources and data privacy. This report focused on the development of WiNGPT-3.0, the 32-billion parameter LLMs, engineered with the objective of enhancing its capacity for medical reasoning and exploring its potential for effective integration within healthcare IT infrastructures. The broader aim is to advance towards clinically applicable models. The approach involved a multi-stage training pipeline tailored for general, medical, and clinical reasoning. This pipeline incorporated supervised fine-tuning (SFT) and reinforcement learning (RL), leveraging curated Long Chain-of-Thought (CoT) datasets, auxiliary reward models, and an evidence-based diagnostic chain simulation. WiNGPT-3.0 demonstrated strong performance: specific model variants achieved scores of 66.6 on MedCalc and 87.1 on MedQA-USMLE. Furthermore, targeted training improved performance on a clinical reasoning task from a baseline score of 58.1 to 62.5. These findings suggest that reinforcement learning, even when applied with a limited dataset of only a few thousand examples, can enhance medical reasoning accuracy. Crucially, this demonstration of RL's efficacy with limited data and computation paves the way for more trustworthy and practically deployable LLMs within clinical workflows and health information infrastructures.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WiNGPT-3.0 Technical Report
Zhuang, Boqin
Song, Chenxiao
Lu, Huitong
Qiao, Jiacheng
Liu, Mingqian
Yu, Mingxing
Hong, Ping
Li, Rui
Song, Xiaoxia
Xu, Xiangjun
Chen, Xu
Ma, Yaoyao
Gao, Yujie
Computation and Language
Current Large Language Models (LLMs) exhibit significant limitations, notably in structured, interpretable, and verifiable medical reasoning, alongside practical deployment challenges related to computational resources and data privacy. This report focused on the development of WiNGPT-3.0, the 32-billion parameter LLMs, engineered with the objective of enhancing its capacity for medical reasoning and exploring its potential for effective integration within healthcare IT infrastructures. The broader aim is to advance towards clinically applicable models. The approach involved a multi-stage training pipeline tailored for general, medical, and clinical reasoning. This pipeline incorporated supervised fine-tuning (SFT) and reinforcement learning (RL), leveraging curated Long Chain-of-Thought (CoT) datasets, auxiliary reward models, and an evidence-based diagnostic chain simulation. WiNGPT-3.0 demonstrated strong performance: specific model variants achieved scores of 66.6 on MedCalc and 87.1 on MedQA-USMLE. Furthermore, targeted training improved performance on a clinical reasoning task from a baseline score of 58.1 to 62.5. These findings suggest that reinforcement learning, even when applied with a limited dataset of only a few thousand examples, can enhance medical reasoning accuracy. Crucially, this demonstration of RL's efficacy with limited data and computation paves the way for more trustworthy and practically deployable LLMs within clinical workflows and health information infrastructures.
title WiNGPT-3.0 Technical Report
topic Computation and Language
url https://arxiv.org/abs/2505.17387