Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Ji Soo, Ko, Byungoh, Cho, Jaewon, Lee, Howoong, Byun, Jaewoon, Kim, Hyunwoo J.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915504636035072
author Lee, Ji Soo
Ko, Byungoh
Cho, Jaewon
Lee, Howoong
Byun, Jaewoon
Kim, Hyunwoo J.
author_facet Lee, Ji Soo
Ko, Byungoh
Cho, Jaewon
Lee, Howoong
Byun, Jaewoon
Kim, Hyunwoo J.
contents In text-video retrieval, auxiliary captions are often used to enhance video understanding, bridging the gap between the modalities. While recent advances in multi-modal large language models (MLLMs) have enabled strong zero-shot caption generation, we observe that such captions tend to be generic and indistinguishable across visually similar videos, limiting their utility for fine-grained retrieval. Moreover, conventional captioning approaches are typically evaluated using language generation metrics, such as BLEU, which are not typically tailored for retrieval tasks that require making discriminative distinctions between candidates. To address this, we propose $\textbf{CaRe-DPO}$, a retrieval framework that directly optimizes caption generation using retrieval relevance scores. At its core is Dual-Group Direct Preference Optimization (DG-DPO), a novel learning strategy that supervises captioning by modeling preferences across groups of distinct video and caption pairs. In addition, we present an MLLM-based retrieval model that incorporates role-embeddings to better distinguish between textual inputs with different functional roles, such as an auxiliary caption and a text query. Through extensive experiments, we demonstrate that CaRe-DPO significantly enhances retrieval performance by effectively leveraging auxiliary knowledge to generate fine-grained captions for retrieval. Code is available at https://github.com/mlvlab/CaReDPO.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16560
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
Lee, Ji Soo
Ko, Byungoh
Cho, Jaewon
Lee, Howoong
Byun, Jaewoon
Kim, Hyunwoo J.
Computer Vision and Pattern Recognition
In text-video retrieval, auxiliary captions are often used to enhance video understanding, bridging the gap between the modalities. While recent advances in multi-modal large language models (MLLMs) have enabled strong zero-shot caption generation, we observe that such captions tend to be generic and indistinguishable across visually similar videos, limiting their utility for fine-grained retrieval. Moreover, conventional captioning approaches are typically evaluated using language generation metrics, such as BLEU, which are not typically tailored for retrieval tasks that require making discriminative distinctions between candidates. To address this, we propose $\textbf{CaRe-DPO}$, a retrieval framework that directly optimizes caption generation using retrieval relevance scores. At its core is Dual-Group Direct Preference Optimization (DG-DPO), a novel learning strategy that supervises captioning by modeling preferences across groups of distinct video and caption pairs. In addition, we present an MLLM-based retrieval model that incorporates role-embeddings to better distinguish between textual inputs with different functional roles, such as an auxiliary caption and a text query. Through extensive experiments, we demonstrate that CaRe-DPO significantly enhances retrieval performance by effectively leveraging auxiliary knowledge to generate fine-grained captions for retrieval. Code is available at https://github.com/mlvlab/CaReDPO.
title Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.16560