Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Luyi, Chen, Jiayuan, Lu, Lu, Peng, Xinwei, Liu, Lihao, He, Junjun, Xu, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TCM-3CEval: A Triaxial Benchmark for Assessing Responses from Large Language Models in Traditional Chinese Medicine
von: Huang, Tianai, et al.
Veröffentlicht: (2025)
von: Huang, Tianai, et al.
Veröffentlicht: (2025)
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
von: Ding, Jinru, et al.
Veröffentlicht: (2025)
von: Ding, Jinru, et al.
Veröffentlicht: (2025)
A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images
von: Liang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Liang, Xiaoyi, et al.
Veröffentlicht: (2025)
TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine
von: Huang, Tianai, et al.
Veröffentlicht: (2025)
von: Huang, Tianai, et al.
Veröffentlicht: (2025)
Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030
von: Bian, Mouxiao, et al.
Veröffentlicht: (2025)
von: Bian, Mouxiao, et al.
Veröffentlicht: (2025)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
von: Guo, Ruiling, et al.
Veröffentlicht: (2025)
von: Guo, Ruiling, et al.
Veröffentlicht: (2025)
MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models
von: Mao, Kangkun, et al.
Veröffentlicht: (2025)
von: Mao, Kangkun, et al.
Veröffentlicht: (2025)
Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering Benchmark
von: Xu, Zhikun, et al.
Veröffentlicht: (2024)
von: Xu, Zhikun, et al.
Veröffentlicht: (2024)
Chinese-Vicuna: A Chinese Instruction-following Llama-based Model
von: Fan, Chenghao, et al.
Veröffentlicht: (2025)
von: Fan, Chenghao, et al.
Veröffentlicht: (2025)
ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
LLMs Struggle with NLI for Perfect Aspect: A Cross-Linguistic Study in Chinese and Japanese
von: Lu, Jie, et al.
Veröffentlicht: (2025)
von: Lu, Jie, et al.
Veröffentlicht: (2025)
CMB: A Comprehensive Medical Benchmark in Chinese
von: Wang, Xidong, et al.
Veröffentlicht: (2023)
von: Wang, Xidong, et al.
Veröffentlicht: (2023)
MedAide: Information Fusion and Anatomy of Medical Intents via LLM-based Agent Collaboration
von: Yang, Dingkang, et al.
Veröffentlicht: (2024)
von: Yang, Dingkang, et al.
Veröffentlicht: (2024)
Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
VLLaVO: Mitigating Visual Gap through LLMs
von: Chen, Shuhao, et al.
Veröffentlicht: (2024)
von: Chen, Shuhao, et al.
Veröffentlicht: (2024)
Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation
von: Hong, Mengze, et al.
Veröffentlicht: (2025)
von: Hong, Mengze, et al.
Veröffentlicht: (2025)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
Benchmarking LLMs' Judgments with No Gold Standard
von: Xu, Shengwei, et al.
Veröffentlicht: (2024)
von: Xu, Shengwei, et al.
Veröffentlicht: (2024)
HLLM-Creator: Hierarchical LLM-based Personalized Creative Generation
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
von: Kumar, Somnath, et al.
Veröffentlicht: (2023)
von: Kumar, Somnath, et al.
Veröffentlicht: (2023)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
von: Zhu, Yiyun, et al.
Veröffentlicht: (2026)
von: Zhu, Yiyun, et al.
Veröffentlicht: (2026)
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
von: Ye, Xiao, et al.
Veröffentlicht: (2025)
von: Ye, Xiao, et al.
Veröffentlicht: (2025)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
von: He, Lehan, et al.
Veröffentlicht: (2024)
von: He, Lehan, et al.
Veröffentlicht: (2024)
Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2025)
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2025)
Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
von: He, Bowei, et al.
Veröffentlicht: (2025)
von: He, Bowei, et al.
Veröffentlicht: (2025)
Flames: Benchmarking Value Alignment of LLMs in Chinese
von: Huang, Kexin, et al.
Veröffentlicht: (2023)
von: Huang, Kexin, et al.
Veröffentlicht: (2023)
CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis
von: Xu, Xinzhe, et al.
Veröffentlicht: (2025)
von: Xu, Xinzhe, et al.
Veröffentlicht: (2025)
How Chinese are Chinese Language Models? The Puzzling Lack of Language Policy in China's LLMs
von: Wen-Yi, Andrea W, et al.
Veröffentlicht: (2024)
von: Wen-Yi, Andrea W, et al.
Veröffentlicht: (2024)
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
von: Liu, Zijie, et al.
Veröffentlicht: (2025)
von: Liu, Zijie, et al.
Veröffentlicht: (2025)
RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
von: Wu, Pengzuo, et al.
Veröffentlicht: (2025)
von: Wu, Pengzuo, et al.
Veröffentlicht: (2025)
MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
von: Hu, Gang, et al.
Veröffentlicht: (2026)
von: Hu, Gang, et al.
Veröffentlicht: (2026)
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making
von: Wang, Yihan, et al.
Veröffentlicht: (2025)
von: Wang, Yihan, et al.
Veröffentlicht: (2025)
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
von: Guo, Guangfu, et al.
Veröffentlicht: (2025)
von: Guo, Guangfu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TCM-3CEval: A Triaxial Benchmark for Assessing Responses from Large Language Models in Traditional Chinese Medicine
von: Huang, Tianai, et al.
Veröffentlicht: (2025) -
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
von: Ding, Jinru, et al.
Veröffentlicht: (2025) -
A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images
von: Liang, Xiaoyi, et al.
Veröffentlicht: (2025) -
TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine
von: Huang, Tianai, et al.
Veröffentlicht: (2025) -
Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030
von: Bian, Mouxiao, et al.
Veröffentlicht: (2025)