DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yakun, Huang, Zhongzhen, Mu, Linjie, Huang, Yutong, Nie, Wei, Liu, Jiaji, Zhang, Shaoting, Liu, Pengfei, Zhang, Xiaofan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912401740267520
author Zhu, Yakun
Huang, Zhongzhen
Mu, Linjie
Huang, Yutong
Nie, Wei
Liu, Jiaji
Zhang, Shaoting
Liu, Pengfei
Zhang, Xiaofan
author_facet Zhu, Yakun
Huang, Zhongzhen
Mu, Linjie
Huang, Yutong
Nie, Wei
Liu, Jiaji
Zhang, Shaoting
Liu, Pengfei
Zhang, Xiaofan
contents The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios. To enable their safe and effective deployment in real-world healthcare settings, it is urgently necessary to benchmark the diagnostic capabilities of current models systematically. Given the limitations of existing medical benchmarks in evaluating advanced diagnostic reasoning, we present DiagnosisArena, a comprehensive and challenging benchmark designed to rigorously assess professional-level diagnostic competence. DiagnosisArena consists of 1,113 pairs of segmented patient cases and corresponding diagnoses, spanning 28 medical specialties, deriving from clinical case reports published in 10 top-tier medical journals. The benchmark is developed through a meticulous construction pipeline, involving multiple rounds of screening and review by both AI systems and human experts, with thorough checks conducted to prevent data leakage. Our study reveals that even the most advanced reasoning models, o3, o1, and DeepSeek-R1, achieve only 51.12%, 31.09%, and 17.79% accuracy, respectively. This finding highlights a significant generalization bottleneck in current large language models when faced with clinical diagnostic reasoning challenges. Through DiagnosisArena, we aim to drive further advancements in AI's diagnostic reasoning capabilities, enabling more effective solutions for real-world clinical diagnostic challenges. We provide the benchmark and evaluation tools for further research and development https://github.com/SPIRAL-MED/DiagnosisArena.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14107
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Zhu, Yakun
Huang, Zhongzhen
Mu, Linjie
Huang, Yutong
Nie, Wei
Liu, Jiaji
Zhang, Shaoting
Liu, Pengfei
Zhang, Xiaofan
Computation and Language
Artificial Intelligence
The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios. To enable their safe and effective deployment in real-world healthcare settings, it is urgently necessary to benchmark the diagnostic capabilities of current models systematically. Given the limitations of existing medical benchmarks in evaluating advanced diagnostic reasoning, we present DiagnosisArena, a comprehensive and challenging benchmark designed to rigorously assess professional-level diagnostic competence. DiagnosisArena consists of 1,113 pairs of segmented patient cases and corresponding diagnoses, spanning 28 medical specialties, deriving from clinical case reports published in 10 top-tier medical journals. The benchmark is developed through a meticulous construction pipeline, involving multiple rounds of screening and review by both AI systems and human experts, with thorough checks conducted to prevent data leakage. Our study reveals that even the most advanced reasoning models, o3, o1, and DeepSeek-R1, achieve only 51.12%, 31.09%, and 17.79% accuracy, respectively. This finding highlights a significant generalization bottleneck in current large language models when faced with clinical diagnostic reasoning challenges. Through DiagnosisArena, we aim to drive further advancements in AI's diagnostic reasoning capabilities, enabling more effective solutions for real-world clinical diagnostic challenges. We provide the benchmark and evaluation tools for further research and development https://github.com/SPIRAL-MED/DiagnosisArena.
title DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.14107