PiercingEye: Dual-Space Video Violence Detection with Hyperbolic Vision-Language Guidance

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Leng, Jiaxu, Wu, Zhanjie, Tan, Mingpi, Mo, Mengjingcheng, Zheng, Jiankang, Li, Qingqing, Gan, Ji, Gao, Xinbo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916708912988160
author Leng, Jiaxu
Wu, Zhanjie
Tan, Mingpi
Mo, Mengjingcheng
Zheng, Jiankang
Li, Qingqing
Gan, Ji
Gao, Xinbo
author_facet Leng, Jiaxu
Wu, Zhanjie
Tan, Mingpi
Mo, Mengjingcheng
Zheng, Jiankang
Li, Qingqing
Gan, Ji
Gao, Xinbo
contents Existing weakly supervised video violence detection (VVD) methods primarily rely on Euclidean representation learning, which often struggles to distinguish visually similar yet semantically distinct events due to limited hierarchical modeling and insufficient ambiguous training samples. To address this challenge, we propose PiercingEye, a novel dual-space learning framework that synergizes Euclidean and hyperbolic geometries to enhance discriminative feature representation. Specifically, PiercingEye introduces a layer-sensitive hyperbolic aggregation strategy with hyperbolic Dirichlet energy constraints to progressively model event hierarchies, and a cross-space attention mechanism to facilitate complementary feature interactions between Euclidean and hyperbolic spaces. Furthermore, to mitigate the scarcity of ambiguous samples, we leverage large language models to generate logic-guided ambiguous event descriptions, enabling explicit supervision through a hyperbolic vision-language contrastive loss that prioritizes high-confusion samples via dynamic similarity-aware weighting. Extensive experiments on XD-Violence and UCF-Crime benchmarks demonstrate that PiercingEye achieves state-of-the-art performance, with particularly strong results on a newly curated ambiguous event subset, validating its superior capability in fine-grained violence detection.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18866
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PiercingEye: Dual-Space Video Violence Detection with Hyperbolic Vision-Language Guidance
Leng, Jiaxu
Wu, Zhanjie
Tan, Mingpi
Mo, Mengjingcheng
Zheng, Jiankang
Li, Qingqing
Gan, Ji
Gao, Xinbo
Computer Vision and Pattern Recognition
Existing weakly supervised video violence detection (VVD) methods primarily rely on Euclidean representation learning, which often struggles to distinguish visually similar yet semantically distinct events due to limited hierarchical modeling and insufficient ambiguous training samples. To address this challenge, we propose PiercingEye, a novel dual-space learning framework that synergizes Euclidean and hyperbolic geometries to enhance discriminative feature representation. Specifically, PiercingEye introduces a layer-sensitive hyperbolic aggregation strategy with hyperbolic Dirichlet energy constraints to progressively model event hierarchies, and a cross-space attention mechanism to facilitate complementary feature interactions between Euclidean and hyperbolic spaces. Furthermore, to mitigate the scarcity of ambiguous samples, we leverage large language models to generate logic-guided ambiguous event descriptions, enabling explicit supervision through a hyperbolic vision-language contrastive loss that prioritizes high-confusion samples via dynamic similarity-aware weighting. Extensive experiments on XD-Violence and UCF-Crime benchmarks demonstrate that PiercingEye achieves state-of-the-art performance, with particularly strong results on a newly curated ambiguous event subset, validating its superior capability in fine-grained violence detection.
title PiercingEye: Dual-Space Video Violence Detection with Hyperbolic Vision-Language Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.18866