Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chang, Ligang, Xu, Shengkai, Shen, Liangchang, Xu, Binhan, Wang, Junqiao, Shi, Tianyu, Du, Yanhui
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911157960310784
author Chang, Ligang
Xu, Shengkai
Shen, Liangchang
Xu, Binhan
Wang, Junqiao
Shi, Tianyu
Du, Yanhui
author_facet Chang, Ligang
Xu, Shengkai
Shen, Liangchang
Xu, Binhan
Wang, Junqiao
Shi, Tianyu
Du, Yanhui
contents Violence detection in public surveillance is critical for public safety. This study addresses challenges such as small-scale targets, complex environments, and real-time temporal analysis. We propose Vi-SAFE, a spatial-temporal framework that integrates an enhanced YOLOv8 with a Temporal Segment Network (TSN) for video surveillance. The YOLOv8 model is optimized with GhostNetV3 as a lightweight backbone, an exponential moving average (EMA) attention mechanism, and pruning to reduce computational cost while maintaining accuracy. YOLOv8 and TSN are trained separately on pedestrian and violence datasets, where YOLOv8 extracts human regions and TSN performs binary classification of violent behavior. Experiments on the RWF-2000 dataset show that Vi-SAFE achieves an accuracy of 0.88, surpassing TSN alone (0.77) and outperforming existing methods in both accuracy and efficiency, demonstrating its effectiveness for public safety surveillance. Code is available at https://anonymous.4open.science/r/Vi-SAFE-3B42/README.md.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13210
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance
Chang, Ligang
Xu, Shengkai
Shen, Liangchang
Xu, Binhan
Wang, Junqiao
Shi, Tianyu
Du, Yanhui
Computer Vision and Pattern Recognition
I.2.10; I.4.8
Violence detection in public surveillance is critical for public safety. This study addresses challenges such as small-scale targets, complex environments, and real-time temporal analysis. We propose Vi-SAFE, a spatial-temporal framework that integrates an enhanced YOLOv8 with a Temporal Segment Network (TSN) for video surveillance. The YOLOv8 model is optimized with GhostNetV3 as a lightweight backbone, an exponential moving average (EMA) attention mechanism, and pruning to reduce computational cost while maintaining accuracy. YOLOv8 and TSN are trained separately on pedestrian and violence datasets, where YOLOv8 extracts human regions and TSN performs binary classification of violent behavior. Experiments on the RWF-2000 dataset show that Vi-SAFE achieves an accuracy of 0.88, surpassing TSN alone (0.77) and outperforming existing methods in both accuracy and efficiency, demonstrating its effectiveness for public safety surveillance. Code is available at https://anonymous.4open.science/r/Vi-SAFE-3B42/README.md.
title Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance
topic Computer Vision and Pattern Recognition
I.2.10; I.4.8
url https://arxiv.org/abs/2509.13210