Trust-Gated RGB-Depth Fusion and Self-Supervised Spatiotemporal Modelling for Crowd Anomaly Detection
Fuente:
Zenodo
Gespeichert in:
| Hauptverfasser: | , |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901399980212224 |
|---|---|
| author | Ms.Sindhu T Dr.KrishnaPriya P |
| author_facet | Ms.Sindhu T Dr.KrishnaPriya P |
| contents | <p><strong><span>Abstract. </span></strong><span>Detecting unusual crowd behaviour in surveillance settings is difficult because abnormal events are unpredictable, and relying on only one type of visual input has clear limitations. To address this, we introduce a multimodal framework that combines RGB and Depth data using a dynamic trust-gated fusion mechanism. This allows the system to adjust the level of trust it places in each input stream based on signal quality and environmental conditions. Separate feature extractors and a hybrid temporal modelling approach preserve the unique strengths of each modality, while a spatiotemporal attention mechanism directs the model’s focus toward moving subjects and reduces interference from static background elements. Instead of requiring detailed definitions of every possible abnormal behaviour, the system is trained in a self-supervised way to learn and reconstruct normal crowd patterns. Any significant deviation from these learned patterns—measured through reconstruction errors across both modalities—is treated as a potential anomaly. Tests on crowd surveillance datasets show that this method is more robust in challenging conditions such as low lighting and complex scenes. It also outperforms traditional fusion techniques and standard autoencoder models, while providing interpretable results by highlighting abnormal regions through attention‑based spatiotemporal maps.</span></p> <p><strong><span>Key Words</span></strong><span>: </span><span>Multimodal Fusion, RGB‑Depth Surveillance, Abnormal Crowd Behaviour Detection, Self‑Supervised Learning, Spatiotemporal Attention</span></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18677147 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Trust-Gated RGB-Depth Fusion and Self-Supervised Spatiotemporal Modelling for Crowd Anomaly Detection Ms.Sindhu T Dr.KrishnaPriya P Multimodal Fusion, RGB‑Depth Surveillance, Abnormal Crowd Behaviour Detection, Self‑Supervised Learning, Spatiotemporal Attention <p><strong><span>Abstract. </span></strong><span>Detecting unusual crowd behaviour in surveillance settings is difficult because abnormal events are unpredictable, and relying on only one type of visual input has clear limitations. To address this, we introduce a multimodal framework that combines RGB and Depth data using a dynamic trust-gated fusion mechanism. This allows the system to adjust the level of trust it places in each input stream based on signal quality and environmental conditions. Separate feature extractors and a hybrid temporal modelling approach preserve the unique strengths of each modality, while a spatiotemporal attention mechanism directs the model’s focus toward moving subjects and reduces interference from static background elements. Instead of requiring detailed definitions of every possible abnormal behaviour, the system is trained in a self-supervised way to learn and reconstruct normal crowd patterns. Any significant deviation from these learned patterns—measured through reconstruction errors across both modalities—is treated as a potential anomaly. Tests on crowd surveillance datasets show that this method is more robust in challenging conditions such as low lighting and complex scenes. It also outperforms traditional fusion techniques and standard autoencoder models, while providing interpretable results by highlighting abnormal regions through attention‑based spatiotemporal maps.</span></p> <p><strong><span>Key Words</span></strong><span>: </span><span>Multimodal Fusion, RGB‑Depth Surveillance, Abnormal Crowd Behaviour Detection, Self‑Supervised Learning, Spatiotemporal Attention</span></p> |
| title | Trust-Gated RGB-Depth Fusion and Self-Supervised Spatiotemporal Modelling for Crowd Anomaly Detection |
| topic | Multimodal Fusion, RGB‑Depth Surveillance, Abnormal Crowd Behaviour Detection, Self‑Supervised Learning, Spatiotemporal Attention |
| url | https://doi.org/10.5281/zenodo.18677147 |