Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Yuchen, Duan, Yingxuan
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929392148545536
author Yang, Yuchen
Duan, Yingxuan
author_facet Yang, Yuchen
Duan, Yingxuan
contents A more robust and holistic language-video representation is the key to pushing video understanding forward. Despite the improvement in training strategies, the quality of the language-video dataset is less attention to. The current plain and simple text descriptions and the visual-only focus for the language-video tasks result in a limited capacity in real-world natural language video retrieval tasks where queries are much more complex. This paper introduces a method to automatically enhance video-language datasets, making them more modality and context-aware for more sophisticated representation learning needs, hence helping all downstream tasks. Our multifaceted video captioning method captures entities, actions, speech transcripts, aesthetics, and emotional cues, providing detailed and correlating information from the text side to the video side for training. We also develop an agent-like strategy using language models to generate high-quality, factual textual descriptions, reducing human intervention and enabling scalability. The method's effectiveness in improving language-video representation is evaluated through text-video retrieval using the MSR-VTT dataset and several multi-modal retrieval models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13809
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset
Yang, Yuchen
Duan, Yingxuan
Multimedia
Computer Vision and Pattern Recognition
Information Retrieval
A more robust and holistic language-video representation is the key to pushing video understanding forward. Despite the improvement in training strategies, the quality of the language-video dataset is less attention to. The current plain and simple text descriptions and the visual-only focus for the language-video tasks result in a limited capacity in real-world natural language video retrieval tasks where queries are much more complex. This paper introduces a method to automatically enhance video-language datasets, making them more modality and context-aware for more sophisticated representation learning needs, hence helping all downstream tasks. Our multifaceted video captioning method captures entities, actions, speech transcripts, aesthetics, and emotional cues, providing detailed and correlating information from the text side to the video side for training. We also develop an agent-like strategy using language models to generate high-quality, factual textual descriptions, reducing human intervention and enabling scalability. The method's effectiveness in improving language-video representation is evaluated through text-video retrieval using the MSR-VTT dataset and several multi-modal retrieval models.
title Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset
topic Multimedia
Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2406.13809