Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.01779 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914300977741824 |
|---|---|
| author | Hua, Rui Wei, Yu Shu, Zixin Chang, Kai Yan, Dengying Xia, Jianan Liu, Zeyu Zhu, Hui Song, Shujie Xiao, Mingzhong Li, Xiaodong Jia, Dongmei Gao, Zhuye Meng, Yanyan Zhao, Naixuan Fu, Yu Yu, Haibin Yu, Benman Chen, Yuanyuan Dong, Fei Meng, Zhizhou Yang, Pengcheng Zhao, Songxue Pei, Lijuan Hu, Yunhui Ding, Kan Duan, Jiayuan Yin, Wenmao Gu, Yang Zhang, Runshun Zhu, Qiang Yu, Jian Li, Jiansheng Liu, Baoyan Wang, Wenjia Zhou, Xuezhong |
| author_facet | Hua, Rui Wei, Yu Shu, Zixin Chang, Kai Yan, Dengying Xia, Jianan Liu, Zeyu Zhu, Hui Song, Shujie Xiao, Mingzhong Li, Xiaodong Jia, Dongmei Gao, Zhuye Meng, Yanyan Zhao, Naixuan Fu, Yu Yu, Haibin Yu, Benman Chen, Yuanyuan Dong, Fei Meng, Zhizhou Yang, Pengcheng Zhao, Songxue Pei, Lijuan Hu, Yunhui Ding, Kan Duan, Jiayuan Yin, Wenmao Gu, Yang Zhang, Runshun Zhu, Qiang Yu, Jian Li, Jiansheng Liu, Baoyan Wang, Wenjia Zhou, Xuezhong |
| contents | Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are fragmented in coverage and scale and rely on non-unified or generation-heavy scoring that hinders fair comparison. We present the LingLanMiDian (LingLan) benchmark, a large-scale, expert-curated, multi-task suite that unifies evaluation across knowledge recall, multi-hop reasoning, information extraction, and real-world clinical decision-making. LingLan introduces a consistent metric design, a synonym-tolerant protocol for clinical labels, a per-dataset 400-item Hard subset, and a reframing of diagnosis and treatment recommendation into single-choice decision recognition. We conduct comprehensive, zero-shot evaluations on 14 leading open-source and proprietary LLMs, providing a unified perspective on their strengths and limitations in TCM commonsense knowledge understanding, reasoning, and clinical decision support; critically, the evaluation on Hard subset reveals a substantial gap between current models and human experts in TCM-specialized reasoning. By bridging fundamental knowledge and applied reasoning through standardized evaluation, LingLan establishes a unified, quantitative, and extensible foundation for advancing TCM LLMs and domain-specific medical AI research. All evaluation data and code are available at https://github.com/TCMAI-BJTU/LingLan and http://tcmnlp.com. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_01779 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning Hua, Rui Wei, Yu Shu, Zixin Chang, Kai Yan, Dengying Xia, Jianan Liu, Zeyu Zhu, Hui Song, Shujie Xiao, Mingzhong Li, Xiaodong Jia, Dongmei Gao, Zhuye Meng, Yanyan Zhao, Naixuan Fu, Yu Yu, Haibin Yu, Benman Chen, Yuanyuan Dong, Fei Meng, Zhizhou Yang, Pengcheng Zhao, Songxue Pei, Lijuan Hu, Yunhui Ding, Kan Duan, Jiayuan Yin, Wenmao Gu, Yang Zhang, Runshun Zhu, Qiang Yu, Jian Li, Jiansheng Liu, Baoyan Wang, Wenjia Zhou, Xuezhong Artificial Intelligence Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are fragmented in coverage and scale and rely on non-unified or generation-heavy scoring that hinders fair comparison. We present the LingLanMiDian (LingLan) benchmark, a large-scale, expert-curated, multi-task suite that unifies evaluation across knowledge recall, multi-hop reasoning, information extraction, and real-world clinical decision-making. LingLan introduces a consistent metric design, a synonym-tolerant protocol for clinical labels, a per-dataset 400-item Hard subset, and a reframing of diagnosis and treatment recommendation into single-choice decision recognition. We conduct comprehensive, zero-shot evaluations on 14 leading open-source and proprietary LLMs, providing a unified perspective on their strengths and limitations in TCM commonsense knowledge understanding, reasoning, and clinical decision support; critically, the evaluation on Hard subset reveals a substantial gap between current models and human experts in TCM-specialized reasoning. By bridging fundamental knowledge and applied reasoning through standardized evaluation, LingLan establishes a unified, quantitative, and extensible foundation for advancing TCM LLMs and domain-specific medical AI research. All evaluation data and code are available at https://github.com/TCMAI-BJTU/LingLan and http://tcmnlp.com. |
| title | LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2602.01779 |