Domain Adaptation of Attention Heads for Zero-shot Anomaly Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jeong, Kiyoon, Heo, Jaehyuk, Son, Junyeong, Kang, Pilsung
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912864910966784
author Jeong, Kiyoon
Heo, Jaehyuk
Son, Junyeong
Kang, Pilsung
author_facet Jeong, Kiyoon
Heo, Jaehyuk
Son, Junyeong
Kang, Pilsung
contents Zero-shot anomaly detection (ZSAD) enables anomaly detection without normal samples from target categories, addressing scenarios where task-specific training data is unavailable. However, existing ZSAD methods either neglect adaptation of vision-language models to anomaly detection or implement only partial adaptation. This paper proposes Head-adaptive CLIP (HeadCLIP), which effectively adapts both text and image encoders. HeadCLIP employs learnable prompts in the text encoder to generalize normality and abnormality concepts, and introduces learnable head weights in the image encoder to dynamically adjust attention head features for task-specific adaptation. A joint anomaly score is further proposed to leverage adapted pixel-level information for enhanced image-level detection. Experiments on 17 datasets across industrial and medical domains demonstrate that HeadCLIP outperforms existing ZSAD methods at both pixel and image levels, achieving improvements of up to 4.9\%p in pixel-level mean anomaly detection score (mAD) and 3.7%p in image-level mAD in the industrial domain, with comparable gains (3.2%p, 3.2%p) in the medical domain. Code and pretrained weights are available at https://github.com/kiyoonjeong0305/HeadCLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22259
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Domain Adaptation of Attention Heads for Zero-shot Anomaly Detection
Jeong, Kiyoon
Heo, Jaehyuk
Son, Junyeong
Kang, Pilsung
Computer Vision and Pattern Recognition
Zero-shot anomaly detection (ZSAD) enables anomaly detection without normal samples from target categories, addressing scenarios where task-specific training data is unavailable. However, existing ZSAD methods either neglect adaptation of vision-language models to anomaly detection or implement only partial adaptation. This paper proposes Head-adaptive CLIP (HeadCLIP), which effectively adapts both text and image encoders. HeadCLIP employs learnable prompts in the text encoder to generalize normality and abnormality concepts, and introduces learnable head weights in the image encoder to dynamically adjust attention head features for task-specific adaptation. A joint anomaly score is further proposed to leverage adapted pixel-level information for enhanced image-level detection. Experiments on 17 datasets across industrial and medical domains demonstrate that HeadCLIP outperforms existing ZSAD methods at both pixel and image levels, achieving improvements of up to 4.9\%p in pixel-level mean anomaly detection score (mAD) and 3.7%p in image-level mAD in the industrial domain, with comparable gains (3.2%p, 3.2%p) in the medical domain. Code and pretrained weights are available at https://github.com/kiyoonjeong0305/HeadCLIP.
title Domain Adaptation of Attention Heads for Zero-shot Anomaly Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22259