DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yimu, Yuan, Shuai, Xue, Bo, Jian, Xiangru, Pang, Wei, Wang, Mushi, Yu, Ning
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929696262848512
author Wang, Yimu
Yuan, Shuai
Xue, Bo
Jian, Xiangru
Pang, Wei
Wang, Mushi
Yu, Ning
author_facet Wang, Yimu
Yuan, Shuai
Xue, Bo
Jian, Xiangru
Pang, Wei
Wang, Mushi
Yu, Ning
contents Recent progress in video-text retrieval has been driven largely by advancements in model architectures and training strategies. However, the representation learning capabilities of videotext retrieval models remain constrained by lowquality and limited training data annotations. To address this issue, we present a novel ViDeoText Retrieval Paradigm with RElevance-based AugMentation, namely DREAM, which enhances video and text data using large foundation models to learn more generalized features. Specifically, we first adopt a simple augmentation method, which generates self-similar data by randomly duplicating or dropping subwords and frames. In addition, inspired by the recent advancement in visual and language generative models, we propose a more robust augmentation method through textual paraphrasing and video stylization using large language models (LLMs) and visual generative models (VGMs). To further enrich video and text information, we propose a relevance-based augmentation method, where LLMs and VGMs generate and integrate new relevant information into the original data. Leveraging this enriched data, extensive experiments on several video-text retrieval benchmarks demonstrate the superiority of DREAM over existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2404_05083
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models
Wang, Yimu
Yuan, Shuai
Xue, Bo
Jian, Xiangru
Pang, Wei
Wang, Mushi
Yu, Ning
Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
Machine Learning
Recent progress in video-text retrieval has been driven largely by advancements in model architectures and training strategies. However, the representation learning capabilities of videotext retrieval models remain constrained by lowquality and limited training data annotations. To address this issue, we present a novel ViDeoText Retrieval Paradigm with RElevance-based AugMentation, namely DREAM, which enhances video and text data using large foundation models to learn more generalized features. Specifically, we first adopt a simple augmentation method, which generates self-similar data by randomly duplicating or dropping subwords and frames. In addition, inspired by the recent advancement in visual and language generative models, we propose a more robust augmentation method through textual paraphrasing and video stylization using large language models (LLMs) and visual generative models (VGMs). To further enrich video and text information, we propose a relevance-based augmentation method, where LLMs and VGMs generate and integrate new relevant information into the original data. Leveraging this enriched data, extensive experiments on several video-text retrieval benchmarks demonstrate the superiority of DREAM over existing methods.
title DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models
topic Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2404.05083