Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Hang, Wang, Ruheng, Ji, Yuelyu, Kwak, Mingu, Wu, Xizhi, Li, Chenyu, Zhang, Li, Shi, Wenqi, Peng, Yifan, Wang, Yanshan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
Bias Evaluation and Mitigation in Retrieval-Augmented Medical Question-Answering Systems
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
Generative Foundation Model for Structured and Unstructured Electronic Health Records
von: Sivarajkumar, Sonish, et al.
Veröffentlicht: (2025)
von: Sivarajkumar, Sonish, et al.
Veröffentlicht: (2025)
DeepRAG: Integrating Hierarchical Reasoning and Process Supervision for Biomedical Multi-Hop QA
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
Assertion Detection Large Language Model In-context Learning LoRA Fine-tuning
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
Orchestrator Multi-Agent Clinical Decision Support System for Secondary Headache Diagnosis in Primary Care
von: Wu, Xizhi, et al.
Veröffentlicht: (2025)
von: Wu, Xizhi, et al.
Veröffentlicht: (2025)
Mitigating the Risk of Health Inequity Exacerbated by Large Language Models
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
von: Li, Chenyu, et al.
Veröffentlicht: (2026)
von: Li, Chenyu, et al.
Veröffentlicht: (2026)
Automated Extraction of Fluoropyrimidine Treatment and Treatment-Related Toxicities from Clinical Notes Using Natural Language Processing
von: Wu, Xizhi, et al.
Veröffentlicht: (2025)
von: Wu, Xizhi, et al.
Veröffentlicht: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
von: Lu, Meng, et al.
Veröffentlicht: (2025)
von: Lu, Meng, et al.
Veröffentlicht: (2025)
AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning
von: Wei, Yifan, et al.
Veröffentlicht: (2025)
von: Wei, Yifan, et al.
Veröffentlicht: (2025)
ReasoningRank: Teaching Student Models to Rank through Reasoning-Based Knowledge Distillation
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement Learning
von: Lu, Meng, et al.
Veröffentlicht: (2026)
von: Lu, Meng, et al.
Veröffentlicht: (2026)
RAG-RLRC-LaySum at BioLaySumm: Integrating Retrieval-Augmented Generation and Readability Control for Layman Summarization of Biomedical Texts
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
von: Wang, Jianing, et al.
Veröffentlicht: (2026)
von: Wang, Jianing, et al.
Veröffentlicht: (2026)
Towards Safe AI Clinicians: A Comprehensive Study on Large Language Model Jailbreaking in Healthcare
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
Curriculum Guided Reinforcement Learning for Efficient Multi Hop Retrieval Augmented Generation
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
Learning from Committee: Reasoning Distillation from a Mixture of Teachers with Peer-Review
von: Li, Zhuochun, et al.
Veröffentlicht: (2024)
von: Li, Zhuochun, et al.
Veröffentlicht: (2024)
Retrieval--Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning
von: He, Yang, et al.
Veröffentlicht: (2026)
von: He, Yang, et al.
Veröffentlicht: (2026)
PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
von: Wang, Ruheng, et al.
Veröffentlicht: (2025)
von: Wang, Ruheng, et al.
Veröffentlicht: (2025)
When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning
von: Xu, Ruotao, et al.
Veröffentlicht: (2026)
von: Xu, Ruotao, et al.
Veröffentlicht: (2026)
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
von: Li, Lingxiao, et al.
Veröffentlicht: (2025)
von: Li, Lingxiao, et al.
Veröffentlicht: (2025)
Effect of Mind Mapping Combined With Behavior Rating Scale on Medical Compliance Behavior of Atrial Fibrillation Patients Underwent Radiofrequency Ablation
von: Huan Xu, et al.
Veröffentlicht: (2025)
von: Huan Xu, et al.
Veröffentlicht: (2025)
Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees
von: Li, Kun, et al.
Veröffentlicht: (2026)
von: Li, Kun, et al.
Veröffentlicht: (2026)
StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
CharTool: Tool-Integrated Visual Reasoning for Chart Understanding
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
Tool Verification for Test-Time Reinforcement Learning
von: Liao, Ruotong, et al.
Veröffentlicht: (2026)
von: Liao, Ruotong, et al.
Veröffentlicht: (2026)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning
von: Wang, Li, et al.
Veröffentlicht: (2026)
von: Wang, Li, et al.
Veröffentlicht: (2026)
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2025)
MedAdapter: Efficient Test-Time Adaptation of Large Language Models towards Medical Reasoning
von: Shi, Wenqi, et al.
Veröffentlicht: (2024)
von: Shi, Wenqi, et al.
Veröffentlicht: (2024)
Enhancing Large Language Models for Clinical Decision Support by Incorporating Clinical Practice Guidelines
von: Oniani, David, et al.
Veröffentlicht: (2024)
von: Oniani, David, et al.
Veröffentlicht: (2024)
Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
von: Wu, Wenxun, et al.
Veröffentlicht: (2025)
von: Wu, Wenxun, et al.
Veröffentlicht: (2025)
SDoH-GPT: Using Large Language Models to Extract Social Determinants of Health (SDoH)
von: Consoli, Bernardo, et al.
Veröffentlicht: (2024)
von: Consoli, Bernardo, et al.
Veröffentlicht: (2024)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2026)
von: Wu, Hang, et al.
Veröffentlicht: (2026)
Memory-Aware and Uncertainty-Guided Retrieval for Multi-Hop Question Answering
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
von: Chen, Jinpeng, et al.
Veröffentlicht: (2026)
von: Chen, Jinpeng, et al.
Veröffentlicht: (2026)
Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026) -
What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026) -
Bias Evaluation and Mitigation in Retrieval-Augmented Medical Question-Answering Systems
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025) -
Generative Foundation Model for Structured and Unstructured Electronic Health Records
von: Sivarajkumar, Sonish, et al.
Veröffentlicht: (2025) -
DeepRAG: Integrating Hierarchical Reasoning and Process Supervision for Biomedical Multi-Hop QA
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)