HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jun, Wang, Jinpeng, Tan, Chaolei, Lian, Niu, Chen, Long, Wang, Yaowei, Zhang, Min, Xia, Shu-Tao, Chen, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908469105262592
author Li, Jun
Wang, Jinpeng
Tan, Chaolei
Lian, Niu
Chen, Long
Wang, Yaowei
Zhang, Min
Xia, Shu-Tao
Chen, Bin
author_facet Li, Jun
Wang, Jinpeng
Tan, Chaolei
Lian, Niu
Chen, Long
Wang, Yaowei
Zhang, Min
Xia, Shu-Tao
Chen, Bin
contents Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos and overlooks certain hierarchical semantics, ultimately leading to suboptimal temporal modeling. To address this issue, we propose the first hyperbolic modeling framework for PRVR, namely HLFormer, which leverages hyperbolic space learning to compensate for the suboptimal hierarchical modeling capabilities of Euclidean space. Specifically, HLFormer integrates the Lorentz Attention Block and Euclidean Attention Block to encode video embeddings in hybrid spaces, using the Mean-Guided Adaptive Interaction Module to dynamically fuse features. Additionally, we introduce a Partial Order Preservation Loss to enforce "text < video" hierarchy through Lorentzian cone constraints. This approach further enhances cross-modal matching by reinforcing partial relevance between video content and text queries. Extensive experiments show that HLFormer outperforms state-of-the-art methods. Code is released at https://github.com/lijun2005/ICCV25-HLFormer.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
Li, Jun
Wang, Jinpeng
Tan, Chaolei
Lian, Niu
Chen, Long
Wang, Yaowei
Zhang, Min
Xia, Shu-Tao
Chen, Bin
Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos and overlooks certain hierarchical semantics, ultimately leading to suboptimal temporal modeling. To address this issue, we propose the first hyperbolic modeling framework for PRVR, namely HLFormer, which leverages hyperbolic space learning to compensate for the suboptimal hierarchical modeling capabilities of Euclidean space. Specifically, HLFormer integrates the Lorentz Attention Block and Euclidean Attention Block to encode video embeddings in hybrid spaces, using the Mean-Guided Adaptive Interaction Module to dynamically fuse features. Additionally, we introduce a Partial Order Preservation Loss to enforce "text < video" hierarchy through Lorentzian cone constraints. This approach further enhances cross-modal matching by reinforcing partial relevance between video content and text queries. Extensive experiments show that HLFormer outperforms state-of-the-art methods. Code is released at https://github.com/lijun2005/ICCV25-HLFormer.
title HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
topic Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
url https://arxiv.org/abs/2507.17402