CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xi, Meng, Zaiqiao, Lever, Jake, Ho, Edmond S. L.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914098927632384
author Zhang, Xi
Meng, Zaiqiao
Lever, Jake
Ho, Edmond S. L.
author_facet Zhang, Xi
Meng, Zaiqiao
Lever, Jake
Ho, Edmond S. L.
contents Multimodal large language models (MLLMs) have recently achieved remarkable progress in radiology by integrating visual perception with natural language understanding. However, they often generate clinically unsupported descriptions, known as medical hallucinations, which pose serious risks in medical applications that demand accuracy and image-grounded outputs. Through empirical analysis, we find that prompt-induced hallucinations remain prevalent in radiology MLLMs, largely due to over-sensitivity to clinical sections. To address this, we introduce Clinical Contrastive Decoding (CCD), a training-free and retrieval-free inference framework that integrates structured clinical signals from task-specific radiology expert models. CCD introduces a dual-stage contrastive mechanism to refine token-level logits during generation, thereby enhancing clinical fidelity without modifying the base MLLM. Experiments on three datasets and multiple models demonstrate that CCD consistently improves overall performance on radiology report generation (RRG). On the MIMIC-CXR dataset, it yields up to a 17% improvement in RadGraph-F1 when applied to state-of-the-art RRG models. Our approach provides a lightweight and generalisable solution for mitigating medical hallucinations, effectively bridging expert models and MLLMs in radiology.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
Zhang, Xi
Meng, Zaiqiao
Lever, Jake
Ho, Edmond S. L.
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.10; J.3; I.5.4
Multimodal large language models (MLLMs) have recently achieved remarkable progress in radiology by integrating visual perception with natural language understanding. However, they often generate clinically unsupported descriptions, known as medical hallucinations, which pose serious risks in medical applications that demand accuracy and image-grounded outputs. Through empirical analysis, we find that prompt-induced hallucinations remain prevalent in radiology MLLMs, largely due to over-sensitivity to clinical sections. To address this, we introduce Clinical Contrastive Decoding (CCD), a training-free and retrieval-free inference framework that integrates structured clinical signals from task-specific radiology expert models. CCD introduces a dual-stage contrastive mechanism to refine token-level logits during generation, thereby enhancing clinical fidelity without modifying the base MLLM. Experiments on three datasets and multiple models demonstrate that CCD consistently improves overall performance on radiology report generation (RRG). On the MIMIC-CXR dataset, it yields up to a 17% improvement in RadGraph-F1 when applied to state-of-the-art RRG models. Our approach provides a lightweight and generalisable solution for mitigating medical hallucinations, effectively bridging expert models and MLLMs in radiology.
title CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.10; J.3; I.5.4
url https://arxiv.org/abs/2509.23379