Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Gundersen, Benjamin, Deperrois, Nicolas, Ruiperez-Campillo, Samuel, Sutter, Thomas M., Vogt, Julia E., Moor, Michael, Nooralahzadeh, Farhad, Krauthammer, Michael
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915669421850624
author Gundersen, Benjamin
Deperrois, Nicolas
Ruiperez-Campillo, Samuel
Sutter, Thomas M.
Vogt, Julia E.
Moor, Michael
Nooralahzadeh, Farhad
Krauthammer, Michael
author_facet Gundersen, Benjamin
Deperrois, Nicolas
Ruiperez-Campillo, Samuel
Sutter, Thomas M.
Vogt, Julia E.
Moor, Michael
Nooralahzadeh, Farhad
Krauthammer, Michael
contents Recent advances in vision-language models (VLMs) have improved Chest X-ray (CXR) interpretation in multiple aspects. However, many medical VLMs rely solely on supervised fine-tuning (SFT), which optimizes next-token prediction without evaluating answer quality. In contrast, reinforcement learning (RL) can incorporate task-specific feedback, and its combination with explicit intermediate reasoning ("thinking") has demonstrated substantial gains on verifiable math and coding tasks. To investigate the effects of RL and thinking in a CXR VLM, we perform large-scale SFT on CXR data to build an updated RadVLM based on Qwen3-VL, followed by a cold-start SFT stage that equips the model with basic thinking ability. We then apply Group Relative Policy Optimization (GRPO) with clinically grounded, task-specific rewards for report generation and visual grounding, and run matched RL experiments on both domain-specific and general-domain Qwen3-VL variants, with and without thinking. Across these settings, we find that while strong SFT remains crucial for high base performance, RL provides additional gains on both tasks, whereas explicit thinking does not appear to further improve results. Under a unified evaluation pipeline, the RL-optimized RadVLM models outperform their baseline counterparts and reach state-of-the-art performance on both report generation and grounding, highlighting clinically aligned RL as a powerful complement to SFT for medical VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10691
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
Gundersen, Benjamin
Deperrois, Nicolas
Ruiperez-Campillo, Samuel
Sutter, Thomas M.
Vogt, Julia E.
Moor, Michael
Nooralahzadeh, Farhad
Krauthammer, Michael
Artificial Intelligence
Computer Vision and Pattern Recognition
Recent advances in vision-language models (VLMs) have improved Chest X-ray (CXR) interpretation in multiple aspects. However, many medical VLMs rely solely on supervised fine-tuning (SFT), which optimizes next-token prediction without evaluating answer quality. In contrast, reinforcement learning (RL) can incorporate task-specific feedback, and its combination with explicit intermediate reasoning ("thinking") has demonstrated substantial gains on verifiable math and coding tasks. To investigate the effects of RL and thinking in a CXR VLM, we perform large-scale SFT on CXR data to build an updated RadVLM based on Qwen3-VL, followed by a cold-start SFT stage that equips the model with basic thinking ability. We then apply Group Relative Policy Optimization (GRPO) with clinically grounded, task-specific rewards for report generation and visual grounding, and run matched RL experiments on both domain-specific and general-domain Qwen3-VL variants, with and without thinking. Across these settings, we find that while strong SFT remains crucial for high base performance, RL provides additional gains on both tasks, whereas explicit thinking does not appear to further improve results. Under a unified evaluation pipeline, the RL-optimized RadVLM models outperform their baseline counterparts and reach state-of-the-art performance on both report generation and grounding, highlighting clinically aligned RL as a powerful complement to SFT for medical VLMs.
title Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10691