CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Messina, Pablo, Villa, Andrés, Alcázar, Juan León, Sánchez, Karen, Hinojosa, Carlos, Parra, Denis, Soto, Álvaro, Ghanem, Bernard
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918299372093440
author Messina, Pablo
Villa, Andrés
Alcázar, Juan León
Sánchez, Karen
Hinojosa, Carlos
Parra, Denis
Soto, Álvaro
Ghanem, Bernard
author_facet Messina, Pablo
Villa, Andrés
Alcázar, Juan León
Sánchez, Karen
Hinojosa, Carlos
Parra, Denis
Soto, Álvaro
Ghanem, Bernard
contents Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present CURE, an error-aware curriculum learning framework that improves grounding and report quality without any additional data. CURE fine-tunes a multimodal instructional model on phrase grounding, grounded report generation, and anatomy-grounded report generation using public datasets. The method dynamically adjusts sampling based on model performance, emphasizing harder samples to improve spatial and textual alignment. CURE improves grounding accuracy by +0.37 IoU, boosts report quality by +0.188 CXRFEScore, and reduces hallucinations by 18.6%. CURE is a data-efficient framework that enhances both grounding accuracy and report reliability. Code is available at https://github.com/PabloMessina/CURE and model weights at https://huggingface.co/pamessina/medgemma-4b-it-cure
format Preprint
id arxiv_https___arxiv_org_abs_2601_15408
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation
Messina, Pablo
Villa, Andrés
Alcázar, Juan León
Sánchez, Karen
Hinojosa, Carlos
Parra, Denis
Soto, Álvaro
Ghanem, Bernard
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present CURE, an error-aware curriculum learning framework that improves grounding and report quality without any additional data. CURE fine-tunes a multimodal instructional model on phrase grounding, grounded report generation, and anatomy-grounded report generation using public datasets. The method dynamically adjusts sampling based on model performance, emphasizing harder samples to improve spatial and textual alignment. CURE improves grounding accuracy by +0.37 IoU, boosts report quality by +0.188 CXRFEScore, and reduces hallucinations by 18.6%. CURE is a data-efficient framework that enhances both grounding accuracy and report reliability. Code is available at https://github.com/PabloMessina/CURE and model weights at https://huggingface.co/pamessina/medgemma-4b-it-cure
title CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.15408