Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghadiya, Ayush, Kar, Purbayan, Chudasama, Vishal, Wasnik, Pankaj
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917881160138752
author Ghadiya, Ayush
Kar, Purbayan
Chudasama, Vishal
Wasnik, Pankaj
author_facet Ghadiya, Ayush
Kar, Purbayan
Chudasama, Vishal
Wasnik, Pankaj
contents Recently, weakly supervised video anomaly detection (WS-VAD) has emerged as a contemporary research direction to identify anomaly events like violence and nudity in videos using only video-level labels. However, this task has substantial challenges, including addressing imbalanced modality information and consistently distinguishing between normal and abnormal features. In this paper, we address these challenges and propose a multi-modal WS-VAD framework to accurately detect anomalies such as violence and nudity. Within the proposed framework, we introduce a new fusion mechanism known as the Cross-modal Fusion Adapter (CFA), which dynamically selects and enhances highly relevant audio-visual features in relation to the visual modality. Additionally, we introduce a Hyperbolic Lorentzian Graph Attention (HLGAtt) to effectively capture the hierarchical relationships between normal and abnormal representations, thereby enhancing feature separation accuracy. Through extensive experiments, we demonstrate that the proposed model achieves state-of-the-art results on benchmark datasets of violence and nudity detection.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20455
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
Ghadiya, Ayush
Kar, Purbayan
Chudasama, Vishal
Wasnik, Pankaj
Computer Vision and Pattern Recognition
Recently, weakly supervised video anomaly detection (WS-VAD) has emerged as a contemporary research direction to identify anomaly events like violence and nudity in videos using only video-level labels. However, this task has substantial challenges, including addressing imbalanced modality information and consistently distinguishing between normal and abnormal features. In this paper, we address these challenges and propose a multi-modal WS-VAD framework to accurately detect anomalies such as violence and nudity. Within the proposed framework, we introduce a new fusion mechanism known as the Cross-modal Fusion Adapter (CFA), which dynamically selects and enhances highly relevant audio-visual features in relation to the visual modality. Additionally, we introduce a Hyperbolic Lorentzian Graph Attention (HLGAtt) to effectively capture the hierarchical relationships between normal and abnormal representations, thereby enhancing feature separation accuracy. Through extensive experiments, we demonstrate that the proposed model achieves state-of-the-art results on benchmark datasets of violence and nudity detection.
title Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.20455