EventVAD: Training-Free Event-Aware Video Anomaly Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shao, Yihua, He, Haojin, Li, Sijie, Chen, Siyu, Long, Xinwei, Zeng, Fanhu, Fan, Yuxuan, Zhang, Muyang, Yan, Ziyang, Ma, Ao, Wang, Xiaochen, Tang, Hao, Wang, Yan, Li, Shuyan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911078994149376
author Shao, Yihua
He, Haojin
Li, Sijie
Chen, Siyu
Long, Xinwei
Zeng, Fanhu
Fan, Yuxuan
Zhang, Muyang
Yan, Ziyang
Ma, Ao
Wang, Xiaochen
Tang, Hao
Wang, Yan
Li, Shuyan
author_facet Shao, Yihua
He, Haojin
Li, Sijie
Chen, Siyu
Long, Xinwei
Zeng, Fanhu
Fan, Yuxuan
Zhang, Muyang
Yan, Ziyang
Ma, Ao
Wang, Xiaochen
Tang, Hao
Wang, Yan
Li, Shuyan
contents Video Anomaly Detection~(VAD) focuses on identifying anomalies within videos. Supervised methods require an amount of in-domain training data and often struggle to generalize to unseen anomalies. In contrast, training-free methods leverage the intrinsic world knowledge of large language models (LLMs) to detect anomalies but face challenges in localizing fine-grained visual transitions and diverse events. Therefore, we propose EventVAD, an event-aware video anomaly detection framework that combines tailored dynamic graph architectures and multimodal LLMs through temporal-event reasoning. Specifically, EventVAD first employs dynamic spatiotemporal graph modeling with time-decay constraints to capture event-aware video features. Then, it performs adaptive noise filtering and uses signal ratio thresholding to detect event boundaries via unsupervised statistical features. The statistical boundary detection module reduces the complexity of processing long videos for MLLMs and improves their temporal reasoning through event consistency. Finally, it utilizes a hierarchical prompting strategy to guide MLLMs in performing reasoning before determining final decisions. We conducted extensive experiments on the UCF-Crime and XD-Violence datasets. The results demonstrate that EventVAD with a 7B MLLM achieves state-of-the-art (SOTA) in training-free settings, outperforming strong baselines that use 7B or larger MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13092
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EventVAD: Training-Free Event-Aware Video Anomaly Detection
Shao, Yihua
He, Haojin
Li, Sijie
Chen, Siyu
Long, Xinwei
Zeng, Fanhu
Fan, Yuxuan
Zhang, Muyang
Yan, Ziyang
Ma, Ao
Wang, Xiaochen
Tang, Hao
Wang, Yan
Li, Shuyan
Computer Vision and Pattern Recognition
Video Anomaly Detection~(VAD) focuses on identifying anomalies within videos. Supervised methods require an amount of in-domain training data and often struggle to generalize to unseen anomalies. In contrast, training-free methods leverage the intrinsic world knowledge of large language models (LLMs) to detect anomalies but face challenges in localizing fine-grained visual transitions and diverse events. Therefore, we propose EventVAD, an event-aware video anomaly detection framework that combines tailored dynamic graph architectures and multimodal LLMs through temporal-event reasoning. Specifically, EventVAD first employs dynamic spatiotemporal graph modeling with time-decay constraints to capture event-aware video features. Then, it performs adaptive noise filtering and uses signal ratio thresholding to detect event boundaries via unsupervised statistical features. The statistical boundary detection module reduces the complexity of processing long videos for MLLMs and improves their temporal reasoning through event consistency. Finally, it utilizes a hierarchical prompting strategy to guide MLLMs in performing reasoning before determining final decisions. We conducted extensive experiments on the UCF-Crime and XD-Violence datasets. The results demonstrate that EventVAD with a 7B MLLM achieves state-of-the-art (SOTA) in training-free settings, outperforming strong baselines that use 7B or larger MLLMs.
title EventVAD: Training-Free Event-Aware Video Anomaly Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.13092