Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Shengyang, Hua, Jiashen, Feng, Junyi, Gong, Xiaojin
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912896951255040
author Sun, Shengyang
Hua, Jiashen
Feng, Junyi
Gong, Xiaojin
author_facet Sun, Shengyang
Hua, Jiashen
Feng, Junyi
Gong, Xiaojin
contents Weakly supervised multimodal video anomaly detection has gained significant attention, yet the potential of the text modality remains under-explored. Text provides explicit semantic information that can enhance anomaly characterization and reduce false alarms. However, extracting effective text features is challenging due to the inability of general-purpose language models to capture anomaly-specific nuances and the scarcity of relevant descriptions. Furthermore, multimodal fusion often suffers from redundancy and imbalance. To address these issues, we propose a novel text-guided framework. First, we introduce an in-context learning-based multi-stage text augmentation mechanism to generate high-quality anomaly text samples for fine-tuning the text feature extractor. Second, we design a multi-scale bottleneck Transformer fusion module that uses compressed bottleneck tokens to progressively integrate information across modalities, mitigating redundancy and imbalance. Experiments on UCF-Crime and XD-Violence demonstrate state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10549
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance
Sun, Shengyang
Hua, Jiashen
Feng, Junyi
Gong, Xiaojin
Computer Vision and Pattern Recognition
Artificial Intelligence
Weakly supervised multimodal video anomaly detection has gained significant attention, yet the potential of the text modality remains under-explored. Text provides explicit semantic information that can enhance anomaly characterization and reduce false alarms. However, extracting effective text features is challenging due to the inability of general-purpose language models to capture anomaly-specific nuances and the scarcity of relevant descriptions. Furthermore, multimodal fusion often suffers from redundancy and imbalance. To address these issues, we propose a novel text-guided framework. First, we introduce an in-context learning-based multi-stage text augmentation mechanism to generate high-quality anomaly text samples for fine-tuning the text feature extractor. Second, we design a multi-scale bottleneck Transformer fusion module that uses compressed bottleneck tokens to progressively integrate information across modalities, mitigating redundancy and imbalance. Experiments on UCF-Crime and XD-Violence demonstrate state-of-the-art performance.
title Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.10549