MadCLIP: Few-shot Medical Anomaly Detection with CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shiri, Mahshid, Beyan, Cigdem, Murino, Vittorio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909667189325824
author Shiri, Mahshid
Beyan, Cigdem
Murino, Vittorio
author_facet Shiri, Mahshid
Beyan, Cigdem
Murino, Vittorio
contents An innovative few-shot anomaly detection approach is presented, leveraging the pre-trained CLIP model for medical data, and adapting it for both image-level anomaly classification (AC) and pixel-level anomaly segmentation (AS). A dual-branch design is proposed to separately capture normal and abnormal features through learnable adapters in the CLIP vision encoder. To improve semantic alignment, learnable text prompts are employed to link visual features. Furthermore, SigLIP loss is applied to effectively handle the many-to-one relationship between images and unpaired text prompts, showcasing its adaptation in the medical field for the first time. Our approach is validated on multiple modalities, demonstrating superior performance over existing methods for AC and AS, in both same-dataset and cross-dataset evaluations. Unlike prior work, it does not rely on synthetic data or memory banks, and an ablation study confirms the contribution of each component. The code is available at https://github.com/mahshid1998/MadCLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23810
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MadCLIP: Few-shot Medical Anomaly Detection with CLIP
Shiri, Mahshid
Beyan, Cigdem
Murino, Vittorio
Computer Vision and Pattern Recognition
An innovative few-shot anomaly detection approach is presented, leveraging the pre-trained CLIP model for medical data, and adapting it for both image-level anomaly classification (AC) and pixel-level anomaly segmentation (AS). A dual-branch design is proposed to separately capture normal and abnormal features through learnable adapters in the CLIP vision encoder. To improve semantic alignment, learnable text prompts are employed to link visual features. Furthermore, SigLIP loss is applied to effectively handle the many-to-one relationship between images and unpaired text prompts, showcasing its adaptation in the medical field for the first time. Our approach is validated on multiple modalities, demonstrating superior performance over existing methods for AC and AS, in both same-dataset and cross-dataset evaluations. Unlike prior work, it does not rely on synthetic data or memory banks, and an ablation study confirms the contribution of each component. The code is available at https://github.com/mahshid1998/MadCLIP.
title MadCLIP: Few-shot Medical Anomaly Detection with CLIP
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23810