ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junying, Cai, Zhenyang, Liu, Zhiheng, Yang, Yunjin, Wang, Rongsheng, Xiao, Qingying, Feng, Xiangyi, Su, Zhan, Guo, Jing, Wan, Xiang, Yu, Guangjun, Li, Haizhou, Wang, Benyou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911112234008576
author Chen, Junying
Cai, Zhenyang
Liu, Zhiheng
Yang, Yunjin
Wang, Rongsheng
Xiao, Qingying
Feng, Xiangyi
Su, Zhan
Guo, Jing
Wan, Xiang
Yu, Guangjun
Li, Haizhou
Wang, Benyou
author_facet Chen, Junying
Cai, Zhenyang
Liu, Zhiheng
Yang, Yunjin
Wang, Rongsheng
Xiao, Qingying
Feng, Xiangyi
Su, Zhan
Guo, Jing
Wan, Xiang
Yu, Guangjun
Li, Haizhou
Wang, Benyou
contents Despite the success of large language models (LLMs) in various domains, their potential in Traditional Chinese Medicine (TCM) remains largely underexplored due to two critical barriers: (1) the scarcity of high-quality TCM data and (2) the inherently multimodal nature of TCM diagnostics, which involve looking, listening, smelling, and pulse-taking. These sensory-rich modalities are beyond the scope of conventional LLMs. To address these challenges, we present ShizhenGPT, the first multimodal LLM tailored for TCM. To overcome data scarcity, we curate the largest TCM dataset to date, comprising 100GB+ of text and 200GB+ of multimodal data, including 1.2M images, 200 hours of audio, and physiological signals. ShizhenGPT is pretrained and instruction-tuned to achieve deep TCM knowledge and multimodal reasoning. For evaluation, we collect recent national TCM qualification exams and build a visual benchmark for Medicinal Recognition and Visual Diagnosis. Experiments demonstrate that ShizhenGPT outperforms comparable-scale LLMs and competes with larger proprietary models. Moreover, it leads in TCM visual understanding among existing multimodal LLMs and demonstrates unified perception across modalities like sound, pulse, smell, and vision, paving the way toward holistic multimodal perception and diagnosis in TCM. Datasets, models, and code are publicly available. We hope this work will inspire further exploration in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
Chen, Junying
Cai, Zhenyang
Liu, Zhiheng
Yang, Yunjin
Wang, Rongsheng
Xiao, Qingying
Feng, Xiangyi
Su, Zhan
Guo, Jing
Wan, Xiang
Yu, Guangjun
Li, Haizhou
Wang, Benyou
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Despite the success of large language models (LLMs) in various domains, their potential in Traditional Chinese Medicine (TCM) remains largely underexplored due to two critical barriers: (1) the scarcity of high-quality TCM data and (2) the inherently multimodal nature of TCM diagnostics, which involve looking, listening, smelling, and pulse-taking. These sensory-rich modalities are beyond the scope of conventional LLMs. To address these challenges, we present ShizhenGPT, the first multimodal LLM tailored for TCM. To overcome data scarcity, we curate the largest TCM dataset to date, comprising 100GB+ of text and 200GB+ of multimodal data, including 1.2M images, 200 hours of audio, and physiological signals. ShizhenGPT is pretrained and instruction-tuned to achieve deep TCM knowledge and multimodal reasoning. For evaluation, we collect recent national TCM qualification exams and build a visual benchmark for Medicinal Recognition and Visual Diagnosis. Experiments demonstrate that ShizhenGPT outperforms comparable-scale LLMs and competes with larger proprietary models. Moreover, it leads in TCM visual understanding among existing multimodal LLMs and demonstrates unified perception across modalities like sound, pulse, smell, and vision, paving the way toward holistic multimodal perception and diagnosis in TCM. Datasets, models, and code are publicly available. We hope this work will inspire further exploration in this field.
title ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
url https://arxiv.org/abs/2508.14706