Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Kun, Sun, Tianyu, Wang, Wenhong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913592357421056
author Qian, Kun
Sun, Tianyu
Wang, Wenhong
author_facet Qian, Kun
Sun, Tianyu
Wang, Wenhong
contents Industrial anomaly detection (IAD) plays a crucial role in the maintenance and quality control of manufacturing processes. In this paper, we propose a novel approach, Vision-Language Anomaly Detection via Contrastive Cross-Modal Training (CLAD), which leverages large vision-language models (LVLMs) to improve both anomaly detection and localization in industrial settings. CLAD aligns visual and textual features into a shared embedding space using contrastive learning, ensuring that normal instances are grouped together while anomalies are pushed apart. Through extensive experiments on two benchmark industrial datasets, MVTec-AD and VisA, we demonstrate that CLAD outperforms state-of-the-art methods in both image-level anomaly detection and pixel-level anomaly localization. Additionally, we provide ablation studies and human evaluation to validate the importance of key components in our method. Our approach not only achieves superior performance but also enhances interpretability by accurately localizing anomalies, making it a promising solution for real-world industrial applications.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00890
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection
Qian, Kun
Sun, Tianyu
Wang, Wenhong
Computer Vision and Pattern Recognition
Industrial anomaly detection (IAD) plays a crucial role in the maintenance and quality control of manufacturing processes. In this paper, we propose a novel approach, Vision-Language Anomaly Detection via Contrastive Cross-Modal Training (CLAD), which leverages large vision-language models (LVLMs) to improve both anomaly detection and localization in industrial settings. CLAD aligns visual and textual features into a shared embedding space using contrastive learning, ensuring that normal instances are grouped together while anomalies are pushed apart. Through extensive experiments on two benchmark industrial datasets, MVTec-AD and VisA, we demonstrate that CLAD outperforms state-of-the-art methods in both image-level anomaly detection and pixel-level anomaly localization. Additionally, we provide ablation studies and human evaluation to validate the importance of key components in our method. Our approach not only achieves superior performance but also enhances interpretability by accurately localizing anomalies, making it a promising solution for real-world industrial applications.
title Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00890