ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Ke, Long, Jun, Fei, Hongxiao, Hua, Liujie, Dai, Zhen, Luo, Yueyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918411994398720
author Ma, Ke
Long, Jun
Fei, Hongxiao
Hua, Liujie
Dai, Zhen
Luo, Yueyi
author_facet Ma, Ke
Long, Jun
Fei, Hongxiao
Hua, Liujie
Dai, Zhen
Luo, Yueyi
contents Pre-trained Vision-Language Models (VLMs) struggle with Zero-Shot Anomaly Detection (ZSAD) due to a critical adaptation gap: they lack the local inductive biases required for dense prediction and employ inflexible feature fusion paradigms. We address these limitations through an Architectural Co-Design framework that jointly refines feature representation and cross-modal fusion. Our method proposes a parameter-efficient Convolutional Low-Rank Adaptation (Conv-LoRA) adapter to inject local inductive biases for fine-grained representation, and introduces a Dynamic Fusion Gateway (DFG) that leverages visual context to adaptively modulate text prompts, enabling a powerful bidirectional fusion. Extensive experiments on diverse industrial and medical benchmarks demonstrate superior accuracy and robustness, validating that this synergistic co-design is critical for robustly adapting foundation models to dense perception tasks. The source code is available at https://github.com/cockmake/ACD-CLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07819
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection
Ma, Ke
Long, Jun
Fei, Hongxiao
Hua, Liujie
Dai, Zhen
Luo, Yueyi
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Pre-trained Vision-Language Models (VLMs) struggle with Zero-Shot Anomaly Detection (ZSAD) due to a critical adaptation gap: they lack the local inductive biases required for dense prediction and employ inflexible feature fusion paradigms. We address these limitations through an Architectural Co-Design framework that jointly refines feature representation and cross-modal fusion. Our method proposes a parameter-efficient Convolutional Low-Rank Adaptation (Conv-LoRA) adapter to inject local inductive biases for fine-grained representation, and introduces a Dynamic Fusion Gateway (DFG) that leverages visual context to adaptively modulate text prompts, enabling a powerful bidirectional fusion. Extensive experiments on diverse industrial and medical benchmarks demonstrate superior accuracy and robustness, validating that this synergistic co-design is critical for robustly adapting foundation models to dense perception tasks. The source code is available at https://github.com/cockmake/ACD-CLIP.
title ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.07819