Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909901073154048 |
|---|---|
| author | Jung, Seoik Song, Taekyung Lee, Yangro Lee, Sungjun |
| author_facet | Jung, Seoik Song, Taekyung Lee, Yangro Lee, Sungjun |
| contents | This paper proposes a Short-Window Sliding Learning framework for real-time violence detection in CCTV footages. Unlike conventional long-video training approaches, the proposed method divides videos into 1-2 second clips and applies Large Language Model (LLM)-based auto-caption labeling to construct fine-grained datasets. Each short clip fully utilizes all frames to preserve temporal continuity, enabling precise recognition of rapid violent events. Experiments demonstrate that the proposed method achieves 95.25\% accuracy on RWF-2000 and significantly improves performance on long videos (UCF-Crime: 83.25\%), confirming its strong generalization and real-time applicability in intelligent surveillance systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_10866 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling Jung, Seoik Song, Taekyung Lee, Yangro Lee, Sungjun Computer Vision and Pattern Recognition Artificial Intelligence 68T45, 68T07 I.2.10; I.4.8; I.2.6 This paper proposes a Short-Window Sliding Learning framework for real-time violence detection in CCTV footages. Unlike conventional long-video training approaches, the proposed method divides videos into 1-2 second clips and applies Large Language Model (LLM)-based auto-caption labeling to construct fine-grained datasets. Each short clip fully utilizes all frames to preserve temporal continuity, enabling precise recognition of rapid violent events. Experiments demonstrate that the proposed method achieves 95.25\% accuracy on RWF-2000 and significantly improves performance on long videos (UCF-Crime: 83.25\%), confirming its strong generalization and real-time applicability in intelligent surveillance systems. |
| title | Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence 68T45, 68T07 I.2.10; I.4.8; I.2.6 |
| url | https://arxiv.org/abs/2511.10866 |