_version_ 1866914300977741824
author Hua, Rui
Wei, Yu
Shu, Zixin
Chang, Kai
Yan, Dengying
Xia, Jianan
Liu, Zeyu
Zhu, Hui
Song, Shujie
Xiao, Mingzhong
Li, Xiaodong
Jia, Dongmei
Gao, Zhuye
Meng, Yanyan
Zhao, Naixuan
Fu, Yu
Yu, Haibin
Yu, Benman
Chen, Yuanyuan
Dong, Fei
Meng, Zhizhou
Yang, Pengcheng
Zhao, Songxue
Pei, Lijuan
Hu, Yunhui
Ding, Kan
Duan, Jiayuan
Yin, Wenmao
Gu, Yang
Zhang, Runshun
Zhu, Qiang
Yu, Jian
Li, Jiansheng
Liu, Baoyan
Wang, Wenjia
Zhou, Xuezhong
author_facet Hua, Rui
Wei, Yu
Shu, Zixin
Chang, Kai
Yan, Dengying
Xia, Jianan
Liu, Zeyu
Zhu, Hui
Song, Shujie
Xiao, Mingzhong
Li, Xiaodong
Jia, Dongmei
Gao, Zhuye
Meng, Yanyan
Zhao, Naixuan
Fu, Yu
Yu, Haibin
Yu, Benman
Chen, Yuanyuan
Dong, Fei
Meng, Zhizhou
Yang, Pengcheng
Zhao, Songxue
Pei, Lijuan
Hu, Yunhui
Ding, Kan
Duan, Jiayuan
Yin, Wenmao
Gu, Yang
Zhang, Runshun
Zhu, Qiang
Yu, Jian
Li, Jiansheng
Liu, Baoyan
Wang, Wenjia
Zhou, Xuezhong
contents Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are fragmented in coverage and scale and rely on non-unified or generation-heavy scoring that hinders fair comparison. We present the LingLanMiDian (LingLan) benchmark, a large-scale, expert-curated, multi-task suite that unifies evaluation across knowledge recall, multi-hop reasoning, information extraction, and real-world clinical decision-making. LingLan introduces a consistent metric design, a synonym-tolerant protocol for clinical labels, a per-dataset 400-item Hard subset, and a reframing of diagnosis and treatment recommendation into single-choice decision recognition. We conduct comprehensive, zero-shot evaluations on 14 leading open-source and proprietary LLMs, providing a unified perspective on their strengths and limitations in TCM commonsense knowledge understanding, reasoning, and clinical decision support; critically, the evaluation on Hard subset reveals a substantial gap between current models and human experts in TCM-specialized reasoning. By bridging fundamental knowledge and applied reasoning through standardized evaluation, LingLan establishes a unified, quantitative, and extensible foundation for advancing TCM LLMs and domain-specific medical AI research. All evaluation data and code are available at https://github.com/TCMAI-BJTU/LingLan and http://tcmnlp.com.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01779
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning
Hua, Rui
Wei, Yu
Shu, Zixin
Chang, Kai
Yan, Dengying
Xia, Jianan
Liu, Zeyu
Zhu, Hui
Song, Shujie
Xiao, Mingzhong
Li, Xiaodong
Jia, Dongmei
Gao, Zhuye
Meng, Yanyan
Zhao, Naixuan
Fu, Yu
Yu, Haibin
Yu, Benman
Chen, Yuanyuan
Dong, Fei
Meng, Zhizhou
Yang, Pengcheng
Zhao, Songxue
Pei, Lijuan
Hu, Yunhui
Ding, Kan
Duan, Jiayuan
Yin, Wenmao
Gu, Yang
Zhang, Runshun
Zhu, Qiang
Yu, Jian
Li, Jiansheng
Liu, Baoyan
Wang, Wenjia
Zhou, Xuezhong
Artificial Intelligence
Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are fragmented in coverage and scale and rely on non-unified or generation-heavy scoring that hinders fair comparison. We present the LingLanMiDian (LingLan) benchmark, a large-scale, expert-curated, multi-task suite that unifies evaluation across knowledge recall, multi-hop reasoning, information extraction, and real-world clinical decision-making. LingLan introduces a consistent metric design, a synonym-tolerant protocol for clinical labels, a per-dataset 400-item Hard subset, and a reframing of diagnosis and treatment recommendation into single-choice decision recognition. We conduct comprehensive, zero-shot evaluations on 14 leading open-source and proprietary LLMs, providing a unified perspective on their strengths and limitations in TCM commonsense knowledge understanding, reasoning, and clinical decision support; critically, the evaluation on Hard subset reveals a substantial gap between current models and human experts in TCM-specialized reasoning. By bridging fundamental knowledge and applied reasoning through standardized evaluation, LingLan establishes a unified, quantitative, and extensible foundation for advancing TCM LLMs and domain-specific medical AI research. All evaluation data and code are available at https://github.com/TCMAI-BJTU/LingLan and http://tcmnlp.com.
title LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2602.01779