DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Leqi, Gong, Guoqiang, Hao, Tianxiang, He, Tao, Zhang, Yifeng, Liu, Pengzhang, Zhao, Sicheng, Han, Jungong, Ding, Guiguang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916788356251648
author Shen, Leqi
Gong, Guoqiang
Hao, Tianxiang
He, Tao
Zhang, Yifeng
Liu, Pengzhang
Zhao, Sicheng
Han, Jungong
Ding, Guiguang
author_facet Shen, Leqi
Gong, Guoqiang
Hao, Tianxiang
He, Tao
Zhang, Yifeng
Liu, Pengzhang
Zhao, Sicheng
Han, Jungong
Ding, Guiguang
contents The parameter-efficient adaptation of the image-text pretraining model CLIP for video-text retrieval is a prominent area of research. While CLIP is focused on image-level vision-language matching, video-text retrieval demands comprehensive understanding at the video level. Three key discrepancies emerge in the transfer from image-level to video-level: vision, language, and alignment. However, existing methods mainly focus on vision while neglecting language and alignment. In this paper, we propose Discrepancy Reduction in Vision, Language, and Alignment (DiscoVLA), which simultaneously mitigates all three discrepancies. Specifically, we introduce Image-Video Features Fusion to integrate image-level and video-level features, effectively tackling both vision and language discrepancies. Additionally, we generate pseudo image captions to learn fine-grained image-level alignment. To mitigate alignment discrepancies, we propose Image-to-Video Alignment Distillation, which leverages image-level alignment knowledge to enhance video-level alignment. Extensive experiments demonstrate the superiority of our DiscoVLA. In particular, on MSRVTT with CLIP (ViT-B/16), DiscoVLA outperforms previous methods by 1.5% in R@1, reaching a final score of 50.5% R@1. The code is available at https://github.com/LunarShen/DsicoVLA.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
Shen, Leqi
Gong, Guoqiang
Hao, Tianxiang
He, Tao
Zhang, Yifeng
Liu, Pengzhang
Zhao, Sicheng
Han, Jungong
Ding, Guiguang
Computer Vision and Pattern Recognition
The parameter-efficient adaptation of the image-text pretraining model CLIP for video-text retrieval is a prominent area of research. While CLIP is focused on image-level vision-language matching, video-text retrieval demands comprehensive understanding at the video level. Three key discrepancies emerge in the transfer from image-level to video-level: vision, language, and alignment. However, existing methods mainly focus on vision while neglecting language and alignment. In this paper, we propose Discrepancy Reduction in Vision, Language, and Alignment (DiscoVLA), which simultaneously mitigates all three discrepancies. Specifically, we introduce Image-Video Features Fusion to integrate image-level and video-level features, effectively tackling both vision and language discrepancies. Additionally, we generate pseudo image captions to learn fine-grained image-level alignment. To mitigate alignment discrepancies, we propose Image-to-Video Alignment Distillation, which leverages image-level alignment knowledge to enhance video-level alignment. Extensive experiments demonstrate the superiority of our DiscoVLA. In particular, on MSRVTT with CLIP (ViT-B/16), DiscoVLA outperforms previous methods by 1.5% in R@1, reaching a final score of 50.5% R@1. The code is available at https://github.com/LunarShen/DsicoVLA.
title DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08887