Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Zequn, Zhang, Boyun, Lin, Yuxiao, Jin, Tao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917209703448576
author Xie, Zequn
Zhang, Boyun
Lin, Yuxiao
Jin, Tao
author_facet Xie, Zequn
Zhang, Boyun
Lin, Yuxiao
Jin, Tao
contents Video-text retrieval (VTR) aims to locate relevant videos using natural language queries. Current methods, often based on pre-trained models like CLIP, are hindered by video's inherent redundancy and their reliance on coarse, final-layer features, limiting matching accuracy. To address this, we introduce the HVP-Net (Hierarchical Visual Perception Network), a framework that mines richer video semantics by extracting and refining features from multiple intermediate layers of a vision encoder. Our approach progressively distills salient visual concepts from raw patch-tokens at different semantic levels, mitigating redundancy while preserving crucial details for alignment. This results in a more robust video representation, leading to new state-of-the-art performance on challenging benchmarks including MSRVTT, DiDeMo, and ActivityNet. Our work validates the effectiveness of exploiting hierarchical features for advancing video-text retrieval. Our codes are available at https://github.com/boyun-zhang/HVP-Net.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12768
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
Xie, Zequn
Zhang, Boyun
Lin, Yuxiao
Jin, Tao
Computer Vision and Pattern Recognition
Multimedia
Video-text retrieval (VTR) aims to locate relevant videos using natural language queries. Current methods, often based on pre-trained models like CLIP, are hindered by video's inherent redundancy and their reliance on coarse, final-layer features, limiting matching accuracy. To address this, we introduce the HVP-Net (Hierarchical Visual Perception Network), a framework that mines richer video semantics by extracting and refining features from multiple intermediate layers of a vision encoder. Our approach progressively distills salient visual concepts from raw patch-tokens at different semantic levels, mitigating redundancy while preserving crucial details for alignment. This results in a more robust video representation, leading to new state-of-the-art performance on challenging benchmarks including MSRVTT, DiDeMo, and ActivityNet. Our work validates the effectiveness of exploiting hierarchical features for advancing video-text retrieval. Our codes are available at https://github.com/boyun-zhang/HVP-Net.
title Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2601.12768