Trust-Gated RGB-Depth Fusion and Self-Supervised Spatiotemporal Modelling for Crowd Anomaly Detection

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ms.Sindhu T, Dr.KrishnaPriya P
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901399980212224
author Ms.Sindhu T
Dr.KrishnaPriya P
author_facet Ms.Sindhu T
Dr.KrishnaPriya P
contents <p><strong><span>Abstract. </span></strong><span>Detecting unusual crowd behaviour in surveillance settings is difficult because abnormal events are unpredictable, and relying on only one type of visual input has clear limitations. To address this, we introduce a multimodal framework that combines RGB and Depth data using a dynamic trust-gated fusion mechanism. This allows the system to adjust the level of trust it places in each input stream based on signal quality and environmental conditions. Separate feature extractors and a hybrid temporal modelling approach preserve the unique strengths of each modality, while a spatiotemporal attention mechanism directs the model’s focus toward moving subjects and reduces interference from static background elements. Instead of requiring detailed definitions of every possible abnormal behaviour, the system is trained in a self-supervised way to learn and reconstruct normal crowd patterns. Any significant deviation from these learned patterns—measured through reconstruction errors across both modalities—is treated as a potential anomaly. Tests on crowd surveillance datasets show that this method is more robust in challenging conditions such as low lighting and complex scenes. It also outperforms traditional fusion techniques and standard autoencoder models, while providing interpretable results by highlighting abnormal regions through attention‑based spatiotemporal maps.</span></p> <p><strong><span>Key Words</span></strong><span>: </span><span>Multimodal Fusion, RGB‑Depth Surveillance, Abnormal Crowd Behaviour Detection, Self‑Supervised Learning, Spatiotemporal Attention</span></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18677147
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Trust-Gated RGB-Depth Fusion and Self-Supervised Spatiotemporal Modelling for Crowd Anomaly Detection
Ms.Sindhu T
Dr.KrishnaPriya P
Multimodal Fusion, RGB‑Depth Surveillance, Abnormal Crowd Behaviour Detection, Self‑Supervised Learning, Spatiotemporal Attention
<p><strong><span>Abstract. </span></strong><span>Detecting unusual crowd behaviour in surveillance settings is difficult because abnormal events are unpredictable, and relying on only one type of visual input has clear limitations. To address this, we introduce a multimodal framework that combines RGB and Depth data using a dynamic trust-gated fusion mechanism. This allows the system to adjust the level of trust it places in each input stream based on signal quality and environmental conditions. Separate feature extractors and a hybrid temporal modelling approach preserve the unique strengths of each modality, while a spatiotemporal attention mechanism directs the model’s focus toward moving subjects and reduces interference from static background elements. Instead of requiring detailed definitions of every possible abnormal behaviour, the system is trained in a self-supervised way to learn and reconstruct normal crowd patterns. Any significant deviation from these learned patterns—measured through reconstruction errors across both modalities—is treated as a potential anomaly. Tests on crowd surveillance datasets show that this method is more robust in challenging conditions such as low lighting and complex scenes. It also outperforms traditional fusion techniques and standard autoencoder models, while providing interpretable results by highlighting abnormal regions through attention‑based spatiotemporal maps.</span></p> <p><strong><span>Key Words</span></strong><span>: </span><span>Multimodal Fusion, RGB‑Depth Surveillance, Abnormal Crowd Behaviour Detection, Self‑Supervised Learning, Spatiotemporal Attention</span></p>
title Trust-Gated RGB-Depth Fusion and Self-Supervised Spatiotemporal Modelling for Crowd Anomaly Detection
topic Multimodal Fusion, RGB‑Depth Surveillance, Abnormal Crowd Behaviour Detection, Self‑Supervised Learning, Spatiotemporal Attention
url https://doi.org/10.5281/zenodo.18677147