GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Guangyu, Chen, Dong, Tang, Siliang, Zhuang, Yueting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915672157585408
author Dai, Guangyu
Chen, Dong
Tang, Siliang
Zhuang, Yueting
author_facet Dai, Guangyu
Chen, Dong
Tang, Siliang
Zhuang, Yueting
contents Video anomaly detection (VAD) is a challenging task that detects anomalous frames in continuous surveillance videos. Most previous work utilizes the spatio-temporal correlation of visual features to distinguish whether there are abnormalities in video snippets. Recently, some works attempt to introduce multi-modal information, like text feature, to enhance the results of video anomaly detection. However, these works merely incorporate text features into video snippets in a coarse manner, overlooking the significant amount of redundant information that may exist within the video snippets. Therefore, we propose to leverage the diversity among multi-modal information to further refine the extracted features, reducing the redundancy in visual features, and we propose Grained Multi-modal Feature for Video Anomaly Detection (GMFVAD). Specifically, we generate more grained multi-modal feature based on the video snippet, which summarizes the main content, and text features based on the captions of original video will be introduced to further enhance the visual features of highlighted portions. Experiments show that the proposed GMFVAD achieves state-of-the-art performance on four mainly datasets. Ablation experiments also validate that the improvement of GMFVAD is due to the reduction of redundant information.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20268
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
Dai, Guangyu
Chen, Dong
Tang, Siliang
Zhuang, Yueting
Computer Vision and Pattern Recognition
Multimedia
Video anomaly detection (VAD) is a challenging task that detects anomalous frames in continuous surveillance videos. Most previous work utilizes the spatio-temporal correlation of visual features to distinguish whether there are abnormalities in video snippets. Recently, some works attempt to introduce multi-modal information, like text feature, to enhance the results of video anomaly detection. However, these works merely incorporate text features into video snippets in a coarse manner, overlooking the significant amount of redundant information that may exist within the video snippets. Therefore, we propose to leverage the diversity among multi-modal information to further refine the extracted features, reducing the redundancy in visual features, and we propose Grained Multi-modal Feature for Video Anomaly Detection (GMFVAD). Specifically, we generate more grained multi-modal feature based on the video snippet, which summarizes the main content, and text features based on the captions of original video will be introduced to further enhance the visual features of highlighted portions. Experiments show that the proposed GMFVAD achieves state-of-the-art performance on four mainly datasets. Ablation experiments also validate that the improvement of GMFVAD is due to the reduction of redundant information.
title GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2510.20268