MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Kangkun, Ding, Jinru, Chen, Jiayuan, Bian, Mouxiao, Chen, Ruiyao, Peng, Xinwei, Ren, Sijie, Li, Linyang, Xu, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909879921278976
author Mao, Kangkun
Ding, Jinru
Chen, Jiayuan
Bian, Mouxiao
Chen, Ruiyao
Peng, Xinwei
Ren, Sijie
Li, Linyang
Xu, Jie
author_facet Mao, Kangkun
Ding, Jinru
Chen, Jiayuan
Bian, Mouxiao
Chen, Ruiyao
Peng, Xinwei
Ren, Sijie
Li, Linyang
Xu, Jie
contents As large language models (LLMs) enter the medical domain, most benchmarks evaluate them on question answering or descriptive reasoning, overlooking quantitative reasoning critical to clinical decision-making. Existing datasets like MedCalc-Bench cover few calculation tasks and fail to reflect real-world computational scenarios. We introduce MedCalc-Eval, the largest benchmark for assessing LLMs' medical calculation abilities, comprising 700+ tasks across two types: equation-based (e.g., Cockcroft-Gault, BMI, BSA) and rule-based scoring systems (e.g., Apgar, Glasgow Coma Scale). These tasks span diverse specialties including internal medicine, surgery, pediatrics, and cardiology, offering a broader and more challenging evaluation setting. To improve performance, we further develop MedCalc-Env, a reinforcement learning environment built on the InternBootcamp framework, enabling multi-step clinical reasoning and planning. Fine-tuning a Qwen2.5-32B model within this environment achieves state-of-the-art results on MedCalc-Eval, with notable gains in numerical sensitivity, formula selection, and reasoning robustness. Remaining challenges include unit conversion, multi-condition logic, and contextual understanding. Code and datasets are available at https://github.com/maokangkun/MedCalc-Eval.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27267
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models
Mao, Kangkun
Ding, Jinru
Chen, Jiayuan
Bian, Mouxiao
Chen, Ruiyao
Peng, Xinwei
Ren, Sijie
Li, Linyang
Xu, Jie
Computation and Language
Artificial Intelligence
As large language models (LLMs) enter the medical domain, most benchmarks evaluate them on question answering or descriptive reasoning, overlooking quantitative reasoning critical to clinical decision-making. Existing datasets like MedCalc-Bench cover few calculation tasks and fail to reflect real-world computational scenarios. We introduce MedCalc-Eval, the largest benchmark for assessing LLMs' medical calculation abilities, comprising 700+ tasks across two types: equation-based (e.g., Cockcroft-Gault, BMI, BSA) and rule-based scoring systems (e.g., Apgar, Glasgow Coma Scale). These tasks span diverse specialties including internal medicine, surgery, pediatrics, and cardiology, offering a broader and more challenging evaluation setting. To improve performance, we further develop MedCalc-Env, a reinforcement learning environment built on the InternBootcamp framework, enabling multi-step clinical reasoning and planning. Fine-tuning a Qwen2.5-32B model within this environment achieves state-of-the-art results on MedCalc-Eval, with notable gains in numerical sensitivity, formula selection, and reasoning robustness. Remaining challenges include unit conversion, multi-condition logic, and contextual understanding. Code and datasets are available at https://github.com/maokangkun/MedCalc-Eval.
title MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.27267