When Background Matters: Breaking Medical Vision Language Models by Transferable Attack

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ghosh, Akash, Baidya, Subhadip, Saha, Sriparna, Chen, Xiuying
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908978123898880
author Ghosh, Akash
Baidya, Subhadip
Saha, Sriparna
Chen, Xiuying
author_facet Ghosh, Akash
Baidya, Subhadip
Saha, Sriparna
Chen, Xiuying
contents Vision-Language Models (VLMs) are increasingly used in clinical diagnostics, yet their robustness to adversarial attacks remains largely unexplored, posing serious risks. Existing medical attacks focus on secondary objectives such as model stealing or adversarial fine-tuning, while transferable attacks from natural images introduce visible distortions that clinicians can easily detect. To address this, we propose MedFocusLeak, a highly transferable black-box multimodal attack that induces incorrect yet clinically plausible diagnoses while keeping perturbations imperceptible. The method injects coordinated perturbations into non-diagnostic background regions and employs an attention distraction mechanism to shift the model's focus away from pathological areas. Extensive evaluations across six medical imaging modalities show that MedFocusLeak achieves state-of-the-art performance, generating misleading yet realistic diagnostic outputs across diverse VLMs. We further introduce a unified evaluation framework with novel metrics that jointly capture attack success and image fidelity, revealing a critical weakness in the reasoning capabilities of modern clinical VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17318
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Background Matters: Breaking Medical Vision Language Models by Transferable Attack
Ghosh, Akash
Baidya, Subhadip
Saha, Sriparna
Chen, Xiuying
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) are increasingly used in clinical diagnostics, yet their robustness to adversarial attacks remains largely unexplored, posing serious risks. Existing medical attacks focus on secondary objectives such as model stealing or adversarial fine-tuning, while transferable attacks from natural images introduce visible distortions that clinicians can easily detect. To address this, we propose MedFocusLeak, a highly transferable black-box multimodal attack that induces incorrect yet clinically plausible diagnoses while keeping perturbations imperceptible. The method injects coordinated perturbations into non-diagnostic background regions and employs an attention distraction mechanism to shift the model's focus away from pathological areas. Extensive evaluations across six medical imaging modalities show that MedFocusLeak achieves state-of-the-art performance, generating misleading yet realistic diagnostic outputs across diverse VLMs. We further introduce a unified evaluation framework with novel metrics that jointly capture attack success and image fidelity, revealing a critical weakness in the reasoning capabilities of modern clinical VLMs.
title When Background Matters: Breaking Medical Vision Language Models by Transferable Attack
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.17318