Referring Industrial Anomaly Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yue, Pengfei, Jiang, Xiaokang, Lu, Yilin, Lin, Jianghang, Zhang, Shengchuan, Cao, Liujuan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910010513031168
author Yue, Pengfei
Jiang, Xiaokang
Lu, Yilin
Lin, Jianghang
Zhang, Shengchuan
Cao, Liujuan
author_facet Yue, Pengfei
Jiang, Xiaokang
Lu, Yilin
Lin, Jianghang
Zhang, Shengchuan
Cao, Liujuan
contents Industrial Anomaly Detection (IAD) is vital for manufacturing, yet traditional methods face significant challenges: unsupervised approaches yield rough localizations requiring manual thresholds, while supervised methods overfit due to scarce, imbalanced data. Both suffer from the "One Anomaly Class, One Model" limitation. To address this, we propose Referring Industrial Anomaly Segmentation (RIAS), a paradigm leveraging language to guide detection. RIAS generates precise masks from text descriptions without manual thresholds and uses universal prompts to detect diverse anomalies with a single model. We introduce the MVTec-Ref dataset to support this, designed with diverse referring expressions and focusing on anomaly patterns, notably with 95% small anomalies. We also propose the Dual Query Token with Mask Group Transformer (DQFormer) benchmark, enhanced by Language-Gated Multi-Level Aggregation (LMA) to improve multi-scale segmentation. Unlike traditional methods using redundant queries, DQFormer employs only "Anomaly" and "Background" tokens for efficient visual-textual integration. Experiments demonstrate RIAS's effectiveness in advancing IAD toward open-set capabilities. Code: https://github.com/swagger-coder/RIAS-MVTec-Ref.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03673
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Referring Industrial Anomaly Segmentation
Yue, Pengfei
Jiang, Xiaokang
Lu, Yilin
Lin, Jianghang
Zhang, Shengchuan
Cao, Liujuan
Computer Vision and Pattern Recognition
Industrial Anomaly Detection (IAD) is vital for manufacturing, yet traditional methods face significant challenges: unsupervised approaches yield rough localizations requiring manual thresholds, while supervised methods overfit due to scarce, imbalanced data. Both suffer from the "One Anomaly Class, One Model" limitation. To address this, we propose Referring Industrial Anomaly Segmentation (RIAS), a paradigm leveraging language to guide detection. RIAS generates precise masks from text descriptions without manual thresholds and uses universal prompts to detect diverse anomalies with a single model. We introduce the MVTec-Ref dataset to support this, designed with diverse referring expressions and focusing on anomaly patterns, notably with 95% small anomalies. We also propose the Dual Query Token with Mask Group Transformer (DQFormer) benchmark, enhanced by Language-Gated Multi-Level Aggregation (LMA) to improve multi-scale segmentation. Unlike traditional methods using redundant queries, DQFormer employs only "Anomaly" and "Background" tokens for efficient visual-textual integration. Experiments demonstrate RIAS's effectiveness in advancing IAD toward open-set capabilities. Code: https://github.com/swagger-coder/RIAS-MVTec-Ref.
title Referring Industrial Anomaly Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.03673