Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yuxin, Luo, Wei, Zhang, Hui, Chen, Qiyu, Yao, Haiming, Shen, Weiming, Cao, Yunkang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911263415599104
author Jiang, Yuxin
Luo, Wei
Zhang, Hui
Chen, Qiyu
Yao, Haiming
Shen, Weiming
Cao, Yunkang
author_facet Jiang, Yuxin
Luo, Wei
Zhang, Hui
Chen, Qiyu
Yao, Haiming
Shen, Weiming
Cao, Yunkang
contents We propose Anomagic, a zero-shot anomaly generation method that produces semantically coherent anomalies without requiring any exemplar anomalies. By unifying both visual and textual cues through a crossmodal prompt encoding scheme, Anomagic leverages rich contextual information to steer an inpainting-based generation pipeline. A subsequent contrastive refinement strategy enforces precise alignment between synthesized anomalies and their masks, thereby bolstering downstream anomaly detection accuracy. To facilitate training, we introduce AnomVerse, a collection of 12,987 anomaly-mask-caption triplets assembled from 13 publicly available datasets, where captions are automatically generated by multimodal large language models using structured visual prompts and template-based textual hints. Extensive experiments demonstrate that Anomagic trained on AnomVerse can synthesize more realistic and varied anomalies than prior methods, yielding superior improvements in downstream anomaly detection. Furthermore, Anomagic can generate anomalies for any normal-category image using user-defined prompts, establishing a versatile foundation model for anomaly generation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10020
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation
Jiang, Yuxin
Luo, Wei
Zhang, Hui
Chen, Qiyu
Yao, Haiming
Shen, Weiming
Cao, Yunkang
Computer Vision and Pattern Recognition
Artificial Intelligence
We propose Anomagic, a zero-shot anomaly generation method that produces semantically coherent anomalies without requiring any exemplar anomalies. By unifying both visual and textual cues through a crossmodal prompt encoding scheme, Anomagic leverages rich contextual information to steer an inpainting-based generation pipeline. A subsequent contrastive refinement strategy enforces precise alignment between synthesized anomalies and their masks, thereby bolstering downstream anomaly detection accuracy. To facilitate training, we introduce AnomVerse, a collection of 12,987 anomaly-mask-caption triplets assembled from 13 publicly available datasets, where captions are automatically generated by multimodal large language models using structured visual prompts and template-based textual hints. Extensive experiments demonstrate that Anomagic trained on AnomVerse can synthesize more realistic and varied anomalies than prior methods, yielding superior improvements in downstream anomaly detection. Furthermore, Anomagic can generate anomalies for any normal-category image using user-defined prompts, establishing a versatile foundation model for anomaly generation.
title Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.10020