Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Haoyu, Gu, Jiaxi, Wang, Shicong, Zhang, Xing, Xu, Hang, Wu, Zuxuan, Jiang, Yu-Gang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908496439541760
author Zhao, Haoyu
Gu, Jiaxi
Wang, Shicong
Zhang, Xing
Xu, Hang
Wu, Zuxuan
Jiang, Yu-Gang
author_facet Zhao, Haoyu
Gu, Jiaxi
Wang, Shicong
Zhang, Xing
Xu, Hang
Wu, Zuxuan
Jiang, Yu-Gang
contents The explosive growth of video streaming presents challenges in achieving high accuracy and low training costs for video-language retrieval. However, existing methods rely on large-scale pre-training to improve video retrieval performance, resulting in significant computational demands. Additionally, the fine-grained information in videos and texts remains underexplored. To alleviate these problems, we propose a novel framework to learn fine-grained features for better alignment and introduce an inference pipeline to improve performance without additional training. Specifically, we employ coarse-to-fine objectives to understand the semantic information of video-text pairs, including contrastive and matching learning. The fine-grained data used for training is obtained through the Granularity-Aware Representation module, which is designed based on similarity analysis between video frames and words in captions. Furthermore, we observe that the repetition of keywords in the original captions, referred to as "Repetition", can enhance retrieval performance and improve alignment between video and text. Based on this insight, we propose a novel and effective inference pipeline that incorporates a voting mechanism and a new Matching Entropy metric to achieve better retrieval performance without requiring additional pre-training. Experimental results on four benchmarks demonstrate that the proposed method outperforms previous approaches. Additionally, our inference pipeline achieves significant performance improvements, with a 2.1% increase in Recall@1 on the MSR-VTT dataset and a 1.6% increase on the DiDeMo dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14812
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
Zhao, Haoyu
Gu, Jiaxi
Wang, Shicong
Zhang, Xing
Xu, Hang
Wu, Zuxuan
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
The explosive growth of video streaming presents challenges in achieving high accuracy and low training costs for video-language retrieval. However, existing methods rely on large-scale pre-training to improve video retrieval performance, resulting in significant computational demands. Additionally, the fine-grained information in videos and texts remains underexplored. To alleviate these problems, we propose a novel framework to learn fine-grained features for better alignment and introduce an inference pipeline to improve performance without additional training. Specifically, we employ coarse-to-fine objectives to understand the semantic information of video-text pairs, including contrastive and matching learning. The fine-grained data used for training is obtained through the Granularity-Aware Representation module, which is designed based on similarity analysis between video frames and words in captions. Furthermore, we observe that the repetition of keywords in the original captions, referred to as "Repetition", can enhance retrieval performance and improve alignment between video and text. Based on this insight, we propose a novel and effective inference pipeline that incorporates a voting mechanism and a new Matching Entropy metric to achieve better retrieval performance without requiring additional pre-training. Experimental results on four benchmarks demonstrate that the proposed method outperforms previous approaches. Additionally, our inference pipeline achieves significant performance improvements, with a 2.1% increase in Recall@1 on the MSR-VTT dataset and a 1.6% increase on the DiDeMo dataset.
title Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.14812