Saved in:
Bibliographic Details
Main Authors: Ahn, Sunghyun, Jo, Youngwan, Lee, Kijung, Kwon, Sein, Hong, Inpyo, Park, Sanghyun
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.04504
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914180792057856
author Ahn, Sunghyun
Jo, Youngwan
Lee, Kijung
Kwon, Sein
Hong, Inpyo
Park, Sanghyun
author_facet Ahn, Sunghyun
Jo, Youngwan
Lee, Kijung
Kwon, Sein
Hong, Inpyo
Park, Sanghyun
contents Video anomaly detection (VAD) is crucial for video analysis and surveillance in computer vision. However, existing VAD models rely on learned normal patterns, which makes them difficult to apply to diverse environments. Consequently, users should retrain models or develop separate AI models for new environments, which requires expertise in machine learning, high-performance hardware, and extensive data collection, limiting the practical usability of VAD. To address these challenges, this study proposes customizable video anomaly detection (C-VAD) technique and the AnyAnomaly model. C-VAD considers user-defined text as an abnormal event and detects frames containing a specified event in a video. We effectively implemented AnyAnomaly using a context-aware visual question answering without fine-tuning the large vision language model. To validate the effectiveness of the proposed model, we constructed C-VAD datasets and demonstrated the superiority of AnyAnomaly. Furthermore, our approach showed competitive results on VAD benchmarks, achieving state-of-the-art performance on UBnormal and UCF-Crime and surpassing other methods in generalization across all datasets. Our code is available online at github.com/SkiddieAhn/Paper-AnyAnomaly.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04504
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
Ahn, Sunghyun
Jo, Youngwan
Lee, Kijung
Kwon, Sein
Hong, Inpyo
Park, Sanghyun
Computer Vision and Pattern Recognition
Video anomaly detection (VAD) is crucial for video analysis and surveillance in computer vision. However, existing VAD models rely on learned normal patterns, which makes them difficult to apply to diverse environments. Consequently, users should retrain models or develop separate AI models for new environments, which requires expertise in machine learning, high-performance hardware, and extensive data collection, limiting the practical usability of VAD. To address these challenges, this study proposes customizable video anomaly detection (C-VAD) technique and the AnyAnomaly model. C-VAD considers user-defined text as an abnormal event and detects frames containing a specified event in a video. We effectively implemented AnyAnomaly using a context-aware visual question answering without fine-tuning the large vision language model. To validate the effectiveness of the proposed model, we constructed C-VAD datasets and demonstrated the superiority of AnyAnomaly. Furthermore, our approach showed competitive results on VAD benchmarks, achieving state-of-the-art performance on UBnormal and UCF-Crime and surpassing other methods in generalization across all datasets. Our code is available online at github.com/SkiddieAhn/Paper-AnyAnomaly.
title AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.04504