DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Qu, Zhen, Tao, Xian, Gong, Xinyi, Qu, ShiChen, Zhang, Xiaopei, Wang, Xingang, Shen, Fei, Zhang, Zhengtao, Prasad, Mukesh, Ding, Guiguang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911113355984896
author Qu, Zhen
Tao, Xian
Gong, Xinyi
Qu, ShiChen
Zhang, Xiaopei
Wang, Xingang
Shen, Fei
Zhang, Zhengtao
Prasad, Mukesh
Ding, Guiguang
author_facet Qu, Zhen
Tao, Xian
Gong, Xinyi
Qu, ShiChen
Zhang, Xiaopei
Wang, Xingang
Shen, Fei
Zhang, Zhengtao
Prasad, Mukesh
Ding, Guiguang
contents Recent vision-language models (e.g., CLIP) have demonstrated remarkable class-generalizable ability to unseen classes in few-shot anomaly segmentation (FSAS), leveraging supervised prompt learning or fine-tuning on seen classes. However, their cross-category generalization largely depends on prior knowledge of real seen anomaly samples. In this paper, we propose a novel framework, namely DictAS, which enables a unified model to detect visual anomalies in unseen object categories without any retraining on the target data, only employing a few normal reference images as visual prompts. The insight behind DictAS is to transfer dictionary lookup capabilities to the FSAS task for unseen classes via self-supervised learning, instead of merely memorizing the normal and abnormal feature patterns from the training set. Specifically, DictAS mainly consists of three components: (1) Dictionary Construction - to simulate the index and content of a real dictionary using features from normal reference images. (2) Dictionary Lookup - to retrieve queried region features from the dictionary via a sparse lookup strategy. When a query feature cannot be retrieved, it is classified as an anomaly. (3) Query Discrimination Regularization - to enhance anomaly discrimination by making abnormal features harder to retrieve from the dictionary. To achieve this, Contrastive Query Constraint and Text Alignment Constraint are further proposed. Extensive experiments on seven public industrial and medical datasets demonstrate that DictAS consistently outperforms state-of-the-art FSAS methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13560
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup
Qu, Zhen
Tao, Xian
Gong, Xinyi
Qu, ShiChen
Zhang, Xiaopei
Wang, Xingang
Shen, Fei
Zhang, Zhengtao
Prasad, Mukesh
Ding, Guiguang
Computer Vision and Pattern Recognition
Recent vision-language models (e.g., CLIP) have demonstrated remarkable class-generalizable ability to unseen classes in few-shot anomaly segmentation (FSAS), leveraging supervised prompt learning or fine-tuning on seen classes. However, their cross-category generalization largely depends on prior knowledge of real seen anomaly samples. In this paper, we propose a novel framework, namely DictAS, which enables a unified model to detect visual anomalies in unseen object categories without any retraining on the target data, only employing a few normal reference images as visual prompts. The insight behind DictAS is to transfer dictionary lookup capabilities to the FSAS task for unseen classes via self-supervised learning, instead of merely memorizing the normal and abnormal feature patterns from the training set. Specifically, DictAS mainly consists of three components: (1) Dictionary Construction - to simulate the index and content of a real dictionary using features from normal reference images. (2) Dictionary Lookup - to retrieve queried region features from the dictionary via a sparse lookup strategy. When a query feature cannot be retrieved, it is classified as an anomaly. (3) Query Discrimination Regularization - to enhance anomaly discrimination by making abnormal features harder to retrieve from the dictionary. To achieve this, Contrastive Query Constraint and Text Alignment Constraint are further proposed. Extensive experiments on seven public industrial and medical datasets demonstrate that DictAS consistently outperforms state-of-the-art FSAS methods.
title DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.13560