MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: He, Qing, Bi, Dongsheng, Lu, Jianrong, Yang, Minghui, Chen, Zixiao, Lu, Jiacheng, Chen, Jing, Du, Nannan, Cu, Xiao, Wu, Sijing, Xiang, Peng, Hu, Yinyin, Guo, Yi, Li, Chunpu, Li, Shaoyang, Dong, Zhuo, Jiang, Ming, Guo, Shuai, Feng, Liyun, Peng, Jin, Wang, Jian, Gu, Jinjie, Liu, Junwei
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918281139453952
author He, Qing
Bi, Dongsheng
Lu, Jianrong
Yang, Minghui
Chen, Zixiao
Lu, Jiacheng
Chen, Jing
Du, Nannan
Cu, Xiao
Wu, Sijing
Xiang, Peng
Hu, Yinyin
Guo, Yi
Li, Chunpu
Li, Shaoyang
Dong, Zhuo
Jiang, Ming
Guo, Shuai
Feng, Liyun
Peng, Jin
Wang, Jian
Gu, Jinjie
Liu, Junwei
author_facet He, Qing
Bi, Dongsheng
Lu, Jianrong
Yang, Minghui
Chen, Zixiao
Lu, Jiacheng
Chen, Jing
Du, Nannan
Cu, Xiao
Wu, Sijing
Xiang, Peng
Hu, Yinyin
Guo, Yi
Li, Chunpu
Li, Shaoyang
Dong, Zhuo
Jiang, Ming
Guo, Shuai
Feng, Liyun
Peng, Jin
Wang, Jian
Gu, Jinjie
Liu, Junwei
contents The proliferation of Large Language Models (LLMs) presents transformative potential for healthcare, yet practical deployment is hindered by the absence of frameworks that assess real-world clinical utility. Existing benchmarks test static knowledge, failing to capture the dynamic, application-oriented capabilities required in clinical practice. To bridge this gap, we introduce a Medical LLM Benchmark MLB, a comprehensive benchmark evaluating LLMs on both foundational knowledge and scenario-based reasoning. MLB is structured around five core dimensions: Medical Knowledge (MedKQA), Safety and Ethics (MedSE), Medical Record Understanding (MedRU), Smart Services (SmartServ), and Smart Healthcare (SmartCare). The benchmark integrates 22 datasets (17 newly curated) from diverse Chinese clinical sources, covering 64 clinical specialties. Its design features a rigorous curation pipeline involving 300 licensed physicians. Besides, we provide a scalable evaluation methodology, centered on a specialized judge model trained via Supervised Fine-Tuning (SFT) on expert annotations. Our comprehensive evaluation of 10 leading models reveals a critical translational gap: while the top-ranked model, Kimi-K2-Instruct (77.3% accuracy overall), excels in structured tasks like information extraction (87.8% accuracy in MedRU), performance plummets in patient-facing scenarios (61.3% in SmartServ). Moreover, the exceptional safety score (90.6% in MedSE) of the much smaller Baichuan-M2-32B highlights that targeted training is equally critical. Our specialized judge model, trained via SFT on a 19k expert-annotated medical dataset, achieves 92.1% accuracy, an F1-score of 94.37%, and a Cohen's Kappa of 81.3% for human-AI consistency, validating a reproducible and expert-aligned evaluation protocol. MLB thus provides a rigorous framework to guide the development of clinically viable LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06193
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications
He, Qing
Bi, Dongsheng
Lu, Jianrong
Yang, Minghui
Chen, Zixiao
Lu, Jiacheng
Chen, Jing
Du, Nannan
Cu, Xiao
Wu, Sijing
Xiang, Peng
Hu, Yinyin
Guo, Yi
Li, Chunpu
Li, Shaoyang
Dong, Zhuo
Jiang, Ming
Guo, Shuai
Feng, Liyun
Peng, Jin
Wang, Jian
Gu, Jinjie
Liu, Junwei
Machine Learning
Artificial Intelligence
Computation and Language
The proliferation of Large Language Models (LLMs) presents transformative potential for healthcare, yet practical deployment is hindered by the absence of frameworks that assess real-world clinical utility. Existing benchmarks test static knowledge, failing to capture the dynamic, application-oriented capabilities required in clinical practice. To bridge this gap, we introduce a Medical LLM Benchmark MLB, a comprehensive benchmark evaluating LLMs on both foundational knowledge and scenario-based reasoning. MLB is structured around five core dimensions: Medical Knowledge (MedKQA), Safety and Ethics (MedSE), Medical Record Understanding (MedRU), Smart Services (SmartServ), and Smart Healthcare (SmartCare). The benchmark integrates 22 datasets (17 newly curated) from diverse Chinese clinical sources, covering 64 clinical specialties. Its design features a rigorous curation pipeline involving 300 licensed physicians. Besides, we provide a scalable evaluation methodology, centered on a specialized judge model trained via Supervised Fine-Tuning (SFT) on expert annotations. Our comprehensive evaluation of 10 leading models reveals a critical translational gap: while the top-ranked model, Kimi-K2-Instruct (77.3% accuracy overall), excels in structured tasks like information extraction (87.8% accuracy in MedRU), performance plummets in patient-facing scenarios (61.3% in SmartServ). Moreover, the exceptional safety score (90.6% in MedSE) of the much smaller Baichuan-M2-32B highlights that targeted training is equally critical. Our specialized judge model, trained via SFT on a 19k expert-annotated medical dataset, achieves 92.1% accuracy, an F1-score of 94.37%, and a Cohen's Kappa of 81.3% for human-AI consistency, validating a reproducible and expert-aligned evaluation protocol. MLB thus provides a rigorous framework to guide the development of clinically viable LLMs.
title MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.06193