Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ding, Chao, Bian, Mouxiao, Chen, Pengcheng, Zhang, Hongliang, Li, Tianbin, Liu, Lihao, Chen, Jiayuan, Li, Zhuoran, Zhong, Yabei, Liu, Yongqi, Huang, Haiqing, Shan, Dongming, He, Junjun, Xu, Jie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913831732641792
author Ding, Chao
Bian, Mouxiao
Chen, Pengcheng
Zhang, Hongliang
Li, Tianbin
Liu, Lihao
Chen, Jiayuan
Li, Zhuoran
Zhong, Yabei
Liu, Yongqi
Huang, Haiqing
Shan, Dongming
He, Junjun
Xu, Jie
author_facet Ding, Chao
Bian, Mouxiao
Chen, Pengcheng
Zhang, Hongliang
Li, Tianbin
Liu, Lihao
Chen, Jiayuan
Li, Zhuoran
Zhong, Yabei
Liu, Yongqi
Huang, Haiqing
Shan, Dongming
He, Junjun
Xu, Jie
contents Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is compounded by the predominant reliance of current medical LLMs on corpora from scientific literature or synthetic data, which often lack the granular expert validation and high clinical relevance essential for advancing their specialized medical capabilities. To address these critical gaps, we introduce a highly clinically relevant dataset with 31,247 medical question-answer pairs, each accompanied by expert-validated chain-of-thought (CoT) explanations. This resource, spanning multiple clinical domains, was curated via a scalable human-LLM hybrid pipeline: LLM-generated rationales were iteratively reviewed, scored, and refined by medical experts against a structured rubric, with substandard outputs revised through human effort or guided LLM regeneration until expert consensus. This publicly available dataset provides a vital source for the development of medical LLMs that capable of transparent and verifiable reasoning, thereby advancing safer and more interpretable AI in medicine.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
Ding, Chao
Bian, Mouxiao
Chen, Pengcheng
Zhang, Hongliang
Li, Tianbin
Liu, Lihao
Chen, Jiayuan
Li, Zhuoran
Zhong, Yabei
Liu, Yongqi
Huang, Haiqing
Shan, Dongming
He, Junjun
Xu, Jie
Computer Vision and Pattern Recognition
Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is compounded by the predominant reliance of current medical LLMs on corpora from scientific literature or synthetic data, which often lack the granular expert validation and high clinical relevance essential for advancing their specialized medical capabilities. To address these critical gaps, we introduce a highly clinically relevant dataset with 31,247 medical question-answer pairs, each accompanied by expert-validated chain-of-thought (CoT) explanations. This resource, spanning multiple clinical domains, was curated via a scalable human-LLM hybrid pipeline: LLM-generated rationales were iteratively reviewed, scored, and refined by medical experts against a structured rubric, with substandard outputs revised through human effort or guided LLM regeneration until expert consensus. This publicly available dataset provides a vital source for the development of medical LLMs that capable of transparent and verifiable reasoning, thereby advancing safer and more interpretable AI in medicine.
title Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.06912