MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Jinru, Lu, Lu, Ding, Chao, Bian, Mouxiao, Chen, Jiayuan, Pang, Wenrao, Chen, Ruiyao, Peng, Xinwei, Lu, Renjie, Ren, Sijie, Zhu, Guanxu, Wu, Xiaoqin, Liu, Zhiqiang, Zhang, Rongzhao, Jiang, Luyi, Han, Bing, Wang, Yunqiu, Xu, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915626606395392
author Ding, Jinru
Lu, Lu
Ding, Chao
Bian, Mouxiao
Chen, Jiayuan
Pang, Wenrao
Chen, Ruiyao
Peng, Xinwei
Lu, Renjie
Ren, Sijie
Zhu, Guanxu
Wu, Xiaoqin
Liu, Zhiqiang
Zhang, Rongzhao
Jiang, Luyi
Han, Bing
Wang, Yunqiu
Xu, Jie
author_facet Ding, Jinru
Lu, Lu
Ding, Chao
Bian, Mouxiao
Chen, Jiayuan
Pang, Wenrao
Chen, Ruiyao
Peng, Xinwei
Lu, Renjie
Ren, Sijie
Zhu, Guanxu
Wu, Xiaoqin
Liu, Zhiqiang
Zhang, Rongzhao
Jiang, Luyi
Han, Bing
Wang, Yunqiu
Xu, Jie
contents Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We present MedBench v4, a nationwide, cloud-based benchmarking infrastructure comprising over 700,000 expert-curated tasks spanning 24 primary and 91 secondary specialties, with dedicated tracks for LLMs, multimodal models, and agents. Items undergo multi-stage refinement and multi-round review by clinicians from more than 500 institutions, and open-ended responses are scored by an LLM-as-a-judge calibrated to human ratings. We evaluate 15 frontier models. Base LLMs reach a mean overall score of 54.1/100 (best: Claude Sonnet 4.5, 62.5/100), but safety and ethics remain low (18.4/100). Multimodal models perform worse overall (mean 47.5/100; best: GPT-5, 54.9/100), with solid perception yet weaker cross-modal reasoning. Agents built on the same backbones substantially improve end-to-end performance (mean 79.8/100), with Claude Sonnet 4.5-based agents achieving up to 85.3/100 overall and 88.9/100 on safety tasks. MedBench v4 thus reveals persisting gaps in multimodal reasoning and safety for base models, while showing that governance-aware agentic orchestration can markedly enhance benchmarked clinical readiness without sacrificing capability. By aligning tasks with Chinese clinical guidelines and regulatory priorities, the platform offers a practical reference for hospitals, developers, and policymakers auditing medical AI.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
Ding, Jinru
Lu, Lu
Ding, Chao
Bian, Mouxiao
Chen, Jiayuan
Pang, Wenrao
Chen, Ruiyao
Peng, Xinwei
Lu, Renjie
Ren, Sijie
Zhu, Guanxu
Wu, Xiaoqin
Liu, Zhiqiang
Zhang, Rongzhao
Jiang, Luyi
Han, Bing
Wang, Yunqiu
Xu, Jie
Computation and Language
Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We present MedBench v4, a nationwide, cloud-based benchmarking infrastructure comprising over 700,000 expert-curated tasks spanning 24 primary and 91 secondary specialties, with dedicated tracks for LLMs, multimodal models, and agents. Items undergo multi-stage refinement and multi-round review by clinicians from more than 500 institutions, and open-ended responses are scored by an LLM-as-a-judge calibrated to human ratings. We evaluate 15 frontier models. Base LLMs reach a mean overall score of 54.1/100 (best: Claude Sonnet 4.5, 62.5/100), but safety and ethics remain low (18.4/100). Multimodal models perform worse overall (mean 47.5/100; best: GPT-5, 54.9/100), with solid perception yet weaker cross-modal reasoning. Agents built on the same backbones substantially improve end-to-end performance (mean 79.8/100), with Claude Sonnet 4.5-based agents achieving up to 85.3/100 overall and 88.9/100 on safety tasks. MedBench v4 thus reveals persisting gaps in multimodal reasoning and safety for base models, while showing that governance-aware agentic orchestration can markedly enhance benchmarked clinical readiness without sacrificing capability. By aligning tasks with Chinese clinical guidelines and regulatory priorities, the platform offers a practical reference for hospitals, developers, and policymakers auditing medical AI.
title MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
topic Computation and Language
url https://arxiv.org/abs/2511.14439