Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fragomeni, Adriano, Damen, Dima, Wray, Michael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908384758857728
author Fragomeni, Adriano
Damen, Dima
Wray, Michael
author_facet Fragomeni, Adriano
Damen, Dima
Wray, Michael
contents Text-to-Video (T2V) retrieval aims to identify the most relevant item from a gallery of videos based on a user's text query. Traditional methods rely solely on aligning video and text modalities to compute the similarity and retrieve relevant items. However, recent advancements emphasise incorporating auxiliary information extracted from video and text modalities to improve retrieval performance and bridge the semantic gap between these modalities. Auxiliary information can include visual attributes, such as objects; temporal and spatial context; and textual descriptions, such as speech and rephrased captions. This survey comprehensively reviews 81 research papers on Text-to-Video retrieval that utilise such auxiliary information. It provides a detailed analysis of their methodologies; highlights state-of-the-art results on benchmark datasets; and discusses available datasets and their auxiliary information. Additionally, it proposes promising directions for future research, focusing on different ways to further enhance retrieval performance using this information.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
Fragomeni, Adriano
Damen, Dima
Wray, Michael
Computer Vision and Pattern Recognition
Text-to-Video (T2V) retrieval aims to identify the most relevant item from a gallery of videos based on a user's text query. Traditional methods rely solely on aligning video and text modalities to compute the similarity and retrieve relevant items. However, recent advancements emphasise incorporating auxiliary information extracted from video and text modalities to improve retrieval performance and bridge the semantic gap between these modalities. Auxiliary information can include visual attributes, such as objects; temporal and spatial context; and textual descriptions, such as speech and rephrased captions. This survey comprehensively reviews 81 research papers on Text-to-Video retrieval that utilise such auxiliary information. It provides a detailed analysis of their methodologies; highlights state-of-the-art results on benchmark datasets; and discusses available datasets and their auxiliary information. Additionally, it proposes promising directions for future research, focusing on different ways to further enhance retrieval performance using this information.
title Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23952