QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yao, Yin, Kangping, Dong, Liang, Ma, Zhenxin, Xu, Shuting, Wang, Xuehai, Jiang, Yuxuan, Yu, Tingting, Hong, Yunqing, Liu, Jiayi, Huang, Rianzhe, Zhao, Shuxin, Hu, Haiping, Shang, Wen, Xu, Jian, Jiang, Guanjun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912965999984640
author Wu, Yao
Yin, Kangping
Dong, Liang
Ma, Zhenxin
Xu, Shuting
Wang, Xuehai
Jiang, Yuxuan
Yu, Tingting
Hong, Yunqing
Liu, Jiayi
Huang, Rianzhe
Zhao, Shuxin
Hu, Haiping
Shang, Wen
Xu, Jian
Jiang, Guanjun
author_facet Wu, Yao
Yin, Kangping
Dong, Liang
Ma, Zhenxin
Xu, Shuting
Wang, Xuehai
Jiang, Yuxuan
Yu, Tingting
Hong, Yunqing
Liu, Jiayi
Huang, Rianzhe
Zhao, Shuxin
Hu, Haiping
Shang, Wen
Xu, Jian
Jiang, Guanjun
contents While Large Language Models (LLMs) excel on standardized medical exams, high scores often fail to translate to high-quality responses for real-world medical queries. Current evaluations rely heavily on multiple-choice questions, failing to capture the unstructured, ambiguous, and long-tail complexities inherent in genuine user inquiries. To bridge this gap, we introduce QuarkMedBench, an ecologically valid benchmark tailored for real-world medical LLM assessment. We compiled a massive dataset spanning Clinical Care, Wellness Health, and Professional Inquiry, comprising 20,821 single-turn queries and 3,853 multi-turn sessions. To objectively evaluate open-ended answers, we propose an automated scoring framework that integrates multi-model consensus with evidence-based retrieval to dynamically generate 220,617 fine-grained scoring rubrics (~9.8 per query). During evaluation, hierarchical weighting and safety constraints structurally quantify medical accuracy, key-point coverage, and risk interception, effectively mitigating the high costs and subjectivity of human grading. Experimental results demonstrate that the generated rubrics achieve a 91.8% concordance rate with clinical expert blind audits, establishing highly dependable medical reliability. Crucially, baseline evaluations on this benchmark reveal significant performance disparities among state-of-the-art models when navigating real-world clinical nuances, highlighting the limitations of conventional exam-based metrics. Ultimately, QuarkMedBench establishes a rigorous, reproducible yardstick for measuring LLM performance on complex health issues, while its framework inherently supports dynamic knowledge updates to prevent benchmark obsolescence.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13691
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models
Wu, Yao
Yin, Kangping
Dong, Liang
Ma, Zhenxin
Xu, Shuting
Wang, Xuehai
Jiang, Yuxuan
Yu, Tingting
Hong, Yunqing
Liu, Jiayi
Huang, Rianzhe
Zhao, Shuxin
Hu, Haiping
Shang, Wen
Xu, Jian
Jiang, Guanjun
Computation and Language
Artificial Intelligence
While Large Language Models (LLMs) excel on standardized medical exams, high scores often fail to translate to high-quality responses for real-world medical queries. Current evaluations rely heavily on multiple-choice questions, failing to capture the unstructured, ambiguous, and long-tail complexities inherent in genuine user inquiries. To bridge this gap, we introduce QuarkMedBench, an ecologically valid benchmark tailored for real-world medical LLM assessment. We compiled a massive dataset spanning Clinical Care, Wellness Health, and Professional Inquiry, comprising 20,821 single-turn queries and 3,853 multi-turn sessions. To objectively evaluate open-ended answers, we propose an automated scoring framework that integrates multi-model consensus with evidence-based retrieval to dynamically generate 220,617 fine-grained scoring rubrics (~9.8 per query). During evaluation, hierarchical weighting and safety constraints structurally quantify medical accuracy, key-point coverage, and risk interception, effectively mitigating the high costs and subjectivity of human grading. Experimental results demonstrate that the generated rubrics achieve a 91.8% concordance rate with clinical expert blind audits, establishing highly dependable medical reliability. Crucially, baseline evaluations on this benchmark reveal significant performance disparities among state-of-the-art models when navigating real-world clinical nuances, highlighting the limitations of conventional exam-based metrics. Ultimately, QuarkMedBench establishes a rigorous, reproducible yardstick for measuring LLM performance on complex health issues, while its framework inherently supports dynamic knowledge updates to prevent benchmark obsolescence.
title QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.13691