LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jin, Er, Feng, Qihui, Mou, Yongli, Decker, Stefan, Lakemeyer, Gerhard, Simons, Oliver, Stegmaier, Johannes
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915094374383616
author Jin, Er
Feng, Qihui
Mou, Yongli
Decker, Stefan
Lakemeyer, Gerhard
Simons, Oliver
Stegmaier, Johannes
author_facet Jin, Er
Feng, Qihui
Mou, Yongli
Decker, Stefan
Lakemeyer, Gerhard
Simons, Oliver
Stegmaier, Johannes
contents Logical image understanding involves interpreting and reasoning about the relationships and consistency within an image's visual content. This capability is essential in applications such as industrial inspection, where logical anomaly detection is critical for maintaining high-quality standards and minimizing costly recalls. Previous research in anomaly detection (AD) has relied on prior knowledge for designing algorithms, which often requires extensive manual annotations, significant computing power, and large amounts of data for training. Autoregressive, multimodal Vision Language Models (AVLMs) offer a promising alternative due to their exceptional performance in visual reasoning across various domains. Despite this, their application to logical AD remains unexplored. In this work, we investigate using AVLMs for logical AD and demonstrate that they are well-suited to the task. Combining AVLMs with format embedding and a logic reasoner, we achieve SOTA performance on public benchmarks, MVTec LOCO AD, with an AUROC of 86.0% and F1-max of 83.7%, along with explanations of anomalies. This significantly outperforms the existing SOTA method by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01767
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction
Jin, Er
Feng, Qihui
Mou, Yongli
Decker, Stefan
Lakemeyer, Gerhard
Simons, Oliver
Stegmaier, Johannes
Computer Vision and Pattern Recognition
Logical image understanding involves interpreting and reasoning about the relationships and consistency within an image's visual content. This capability is essential in applications such as industrial inspection, where logical anomaly detection is critical for maintaining high-quality standards and minimizing costly recalls. Previous research in anomaly detection (AD) has relied on prior knowledge for designing algorithms, which often requires extensive manual annotations, significant computing power, and large amounts of data for training. Autoregressive, multimodal Vision Language Models (AVLMs) offer a promising alternative due to their exceptional performance in visual reasoning across various domains. Despite this, their application to logical AD remains unexplored. In this work, we investigate using AVLMs for logical AD and demonstrate that they are well-suited to the task. Combining AVLMs with format embedding and a logic reasoner, we achieve SOTA performance on public benchmarks, MVTec LOCO AD, with an AUROC of 86.0% and F1-max of 83.7%, along with explanations of anomalies. This significantly outperforms the existing SOTA method by a large margin.
title LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.01767