OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Shifang, Lin, Yiheng, Han, Lu, Zhao, Yao, Wei, Yunchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918037190344704
author Zhao, Shifang
Lin, Yiheng
Han, Lu
Zhao, Yao
Wei, Yunchao
author_facet Zhao, Shifang
Lin, Yiheng
Han, Lu
Zhao, Yao
Wei, Yunchao
contents While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce OmniAD, a novel framework that unifies anomaly detection and understanding for fine-grained analysis. OmniAD is a multimodal reasoner that combines visual and textual reasoning processes. The visual reasoning provides detailed inspection by leveraging Text-as-Mask Encoding to perform anomaly detection through text generation without manually selected thresholds. Following this, Visual Guided Textual Reasoning conducts comprehensive analysis by integrating visual perception. To enhance few-shot generalization, we employ an integrated training strategy that combines supervised fine-tuning (SFT) with reinforcement learning (GRPO), incorporating three sophisticated reward functions. Experimental results demonstrate that OmniAD achieves a performance of 79.1 on the MMAD benchmark, surpassing models such as Qwen2.5-VL-7B and GPT-4o. It also shows strong results across multiple anomaly detection benchmarks. These results highlight the importance of enhancing visual perception for effective reasoning in anomaly understanding. All codes and models will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
Zhao, Shifang
Lin, Yiheng
Han, Lu
Zhao, Yao
Wei, Yunchao
Computer Vision and Pattern Recognition
While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce OmniAD, a novel framework that unifies anomaly detection and understanding for fine-grained analysis. OmniAD is a multimodal reasoner that combines visual and textual reasoning processes. The visual reasoning provides detailed inspection by leveraging Text-as-Mask Encoding to perform anomaly detection through text generation without manually selected thresholds. Following this, Visual Guided Textual Reasoning conducts comprehensive analysis by integrating visual perception. To enhance few-shot generalization, we employ an integrated training strategy that combines supervised fine-tuning (SFT) with reinforcement learning (GRPO), incorporating three sophisticated reward functions. Experimental results demonstrate that OmniAD achieves a performance of 79.1 on the MMAD benchmark, surpassing models such as Qwen2.5-VL-7B and GPT-4o. It also shows strong results across multiple anomaly detection benchmarks. These results highlight the importance of enhancing visual perception for effective reasoning in anomaly understanding. All codes and models will be publicly available.
title OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22039