Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913831732641792 |
|---|---|
| author | Ding, Chao Bian, Mouxiao Chen, Pengcheng Zhang, Hongliang Li, Tianbin Liu, Lihao Chen, Jiayuan Li, Zhuoran Zhong, Yabei Liu, Yongqi Huang, Haiqing Shan, Dongming He, Junjun Xu, Jie |
| author_facet | Ding, Chao Bian, Mouxiao Chen, Pengcheng Zhang, Hongliang Li, Tianbin Liu, Lihao Chen, Jiayuan Li, Zhuoran Zhong, Yabei Liu, Yongqi Huang, Haiqing Shan, Dongming He, Junjun Xu, Jie |
| contents | Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is compounded by the predominant reliance of current medical LLMs on corpora from scientific literature or synthetic data, which often lack the granular expert validation and high clinical relevance essential for advancing their specialized medical capabilities. To address these critical gaps, we introduce a highly clinically relevant dataset with 31,247 medical question-answer pairs, each accompanied by expert-validated chain-of-thought (CoT) explanations. This resource, spanning multiple clinical domains, was curated via a scalable human-LLM hybrid pipeline: LLM-generated rationales were iteratively reviewed, scored, and refined by medical experts against a structured rubric, with substandard outputs revised through human effort or guided LLM regeneration until expert consensus. This publicly available dataset provides a vital source for the development of medical LLMs that capable of transparent and verifiable reasoning, thereby advancing safer and more interpretable AI in medicine. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_06912 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI Ding, Chao Bian, Mouxiao Chen, Pengcheng Zhang, Hongliang Li, Tianbin Liu, Lihao Chen, Jiayuan Li, Zhuoran Zhong, Yabei Liu, Yongqi Huang, Haiqing Shan, Dongming He, Junjun Xu, Jie Computer Vision and Pattern Recognition Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is compounded by the predominant reliance of current medical LLMs on corpora from scientific literature or synthetic data, which often lack the granular expert validation and high clinical relevance essential for advancing their specialized medical capabilities. To address these critical gaps, we introduce a highly clinically relevant dataset with 31,247 medical question-answer pairs, each accompanied by expert-validated chain-of-thought (CoT) explanations. This resource, spanning multiple clinical domains, was curated via a scalable human-LLM hybrid pipeline: LLM-generated rationales were iteratively reviewed, scored, and refined by medical experts against a structured rubric, with substandard outputs revised through human effort or guided LLM regeneration until expert consensus. This publicly available dataset provides a vital source for the development of medical LLMs that capable of transparent and verifiable reasoning, thereby advancing safer and more interpretable AI in medicine. |
| title | Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.06912 |