Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hang, Wang, Ruheng, Ji, Yuelyu, Kwak, Mingu, Wu, Xizhi, Li, Chenyu, Zhang, Li, Shi, Wenqi, Peng, Yifan, Wang, Yanshan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917227073110016
author Zhang, Hang
Wang, Ruheng
Ji, Yuelyu
Kwak, Mingu
Wu, Xizhi
Li, Chenyu
Zhang, Li
Shi, Wenqi
Peng, Yifan
Wang, Yanshan
author_facet Zhang, Hang
Wang, Ruheng
Ji, Yuelyu
Kwak, Mingu
Wu, Xizhi
Li, Chenyu
Zhang, Li
Shi, Wenqi
Peng, Yifan
Wang, Yanshan
contents Large language models have achieved strong performance on medical reasoning benchmarks, yet their deployment in clinical settings demands rigorous verification to ensure factual accuracy. While reward models offer a scalable approach for reasoning trace verification, existing methods face two limitations: they produce only scalar reward values without explicit justification, and they rely on single-pass retrieval that precludes adaptive knowledge access as verification unfolds. We introduce $\method$, an agentic framework that addresses these limitations by training medical reasoning verifiers to iteratively query external medical corpora during evaluation. Our approach combines tool-augmented verification with an iterative reinforcement learning paradigm that requires only trace-level supervision, alongside an adaptive curriculum mechanism that dynamically adjusts training data distribution. Across four medical reasoning benchmarks, $\method$ achieves substantial gains over existing methods, improving MedQA accuracy by 23.5% and MedXpertQA by 32.0% relative to the base generator in particular. Crucially, $\method$ demonstrates an $\mathbf{8\times}$ reduction in sampling budget requirement compared to prior reward model baselines. These findings establish that grounding verification in dynamically retrieved evidence offers a principled path toward more reliable medical reasoning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20221
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning
Zhang, Hang
Wang, Ruheng
Ji, Yuelyu
Kwak, Mingu
Wu, Xizhi
Li, Chenyu
Zhang, Li
Shi, Wenqi
Peng, Yifan
Wang, Yanshan
Artificial Intelligence
Computation and Language
Large language models have achieved strong performance on medical reasoning benchmarks, yet their deployment in clinical settings demands rigorous verification to ensure factual accuracy. While reward models offer a scalable approach for reasoning trace verification, existing methods face two limitations: they produce only scalar reward values without explicit justification, and they rely on single-pass retrieval that precludes adaptive knowledge access as verification unfolds. We introduce $\method$, an agentic framework that addresses these limitations by training medical reasoning verifiers to iteratively query external medical corpora during evaluation. Our approach combines tool-augmented verification with an iterative reinforcement learning paradigm that requires only trace-level supervision, alongside an adaptive curriculum mechanism that dynamically adjusts training data distribution. Across four medical reasoning benchmarks, $\method$ achieves substantial gains over existing methods, improving MedQA accuracy by 23.5% and MedXpertQA by 32.0% relative to the base generator in particular. Crucially, $\method$ demonstrates an $\mathbf{8\times}$ reduction in sampling budget requirement compared to prior reward model baselines. These findings establish that grounding verification in dynamically retrieved evidence offers a principled path toward more reliable medical reasoning systems.
title Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.20221