From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Wenhao, Tang, Zhentao, Li, Yafu, Kai, Shixiong, Yuan, Mingxuan, Chen, Chunlin, Wang, Zhi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914416130260992
author Wu, Wenhao
Tang, Zhentao
Li, Yafu
Kai, Shixiong
Yuan, Mingxuan
Chen, Chunlin
Wang, Zhi
author_facet Wu, Wenhao
Tang, Zhentao
Li, Yafu
Kai, Shixiong
Yuan, Mingxuan
Chen, Chunlin
Wang, Zhi
contents Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks in healthcare fields. While Retrieval-Augmented Generation (RAG) mitigates these issues, existing methods rely on noisy token-level signals and lack the multi-round refinement required for complex reasoning. In the paper, we propose MA-RAG (Multi-Round Agentic RAG), a framework that facilitates test-time scaling for complex medical reasoning by iteratively evolving both external evidence and internal reasoning history within an agentic refinement loop. At each round, the agent transforms semantic conflict among candidate responses into actionable queries to retrieve external evidence, while optimizing history reasoning traces to mitigate long-context degradation. MA-RAG extends the self-consistency principle by leveraging the lack of consistency as a proactive signal for multi-round agentic reasoning and retrieval, and mirrors a boosting mechanism that iteratively minimizes the residual error toward a stable, high-fidelity medical consensus. Extensive evaluations across 7 medical Q&A benchmarks show that MA-RAG consistently surpasses competitive inference-time scaling and RAG baselines, delivering substantial +6.8 points on average accuracy over the backbone model. Our code is available at https://github.com/NJU-RL/MA-RAG.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03292
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
Wu, Wenhao
Tang, Zhentao
Li, Yafu
Kai, Shixiong
Yuan, Mingxuan
Chen, Chunlin
Wang, Zhi
Computation and Language
Artificial Intelligence
Information Retrieval
Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks in healthcare fields. While Retrieval-Augmented Generation (RAG) mitigates these issues, existing methods rely on noisy token-level signals and lack the multi-round refinement required for complex reasoning. In the paper, we propose MA-RAG (Multi-Round Agentic RAG), a framework that facilitates test-time scaling for complex medical reasoning by iteratively evolving both external evidence and internal reasoning history within an agentic refinement loop. At each round, the agent transforms semantic conflict among candidate responses into actionable queries to retrieve external evidence, while optimizing history reasoning traces to mitigate long-context degradation. MA-RAG extends the self-consistency principle by leveraging the lack of consistency as a proactive signal for multi-round agentic reasoning and retrieval, and mirrors a boosting mechanism that iteratively minimizes the residual error toward a stable, high-fidelity medical consensus. Extensive evaluations across 7 medical Q&A benchmarks show that MA-RAG consistently surpasses competitive inference-time scaling and RAG baselines, delivering substantial +6.8 points on average accuracy over the backbone model. Our code is available at https://github.com/NJU-RL/MA-RAG.
title From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2603.03292