CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yifan, Li, Xinhao, Yang, Yichun, Meng, Desen, Huang, Rui, Wang, Limin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912281101598720
author Xu, Yifan
Li, Xinhao
Yang, Yichun
Meng, Desen
Huang, Rui
Wang, Limin
author_facet Xu, Yifan
Li, Xinhao
Yang, Yichun
Meng, Desen
Huang, Rui
Wang, Limin
contents Video understanding, including video captioning and retrieval, is still a great challenge for video-language models (VLMs). The existing video retrieval and caption benchmarks only include short descriptions, limits their ability of detailed video understanding evaluation. To address this problem, we present CaReBench, a testing benchmark for fine-grained video captioning and retrieval with 1,000 high-quality pairs of videos and human-annotated detailed captions. Uniquely, it provides manually separated spatial annotations and temporal annotations for each video. Based on this design, we introduce two evaluation metrics, ReBias and CapST, specifically tailored for video retrieval and video captioning tasks, respectively. These metrics enable a comprehensive investigation into the spatial and temporal biases inherent in VLMs. In addition, to handle both video retrieval and video captioning tasks in a unified framework, we develop a simple baseline based on a Multimodal Language Model (MLLM). By implementing a two-stage Supervised Fine-Tuning (SFT), we fully unlock the potential of MLLM, enabling it not only to generate detailed video descriptions but also to extract video features. Surprisingly, experimental results demonstrate that, compared to the CLIP-based models designed for retrieval and the popular MLLMs skilled in video captioning, our baseline shows competitive performance in both fine-grained video retrieval and video detailed captioning.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00513
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
Xu, Yifan
Li, Xinhao
Yang, Yichun
Meng, Desen
Huang, Rui
Wang, Limin
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
Video understanding, including video captioning and retrieval, is still a great challenge for video-language models (VLMs). The existing video retrieval and caption benchmarks only include short descriptions, limits their ability of detailed video understanding evaluation. To address this problem, we present CaReBench, a testing benchmark for fine-grained video captioning and retrieval with 1,000 high-quality pairs of videos and human-annotated detailed captions. Uniquely, it provides manually separated spatial annotations and temporal annotations for each video. Based on this design, we introduce two evaluation metrics, ReBias and CapST, specifically tailored for video retrieval and video captioning tasks, respectively. These metrics enable a comprehensive investigation into the spatial and temporal biases inherent in VLMs. In addition, to handle both video retrieval and video captioning tasks in a unified framework, we develop a simple baseline based on a Multimodal Language Model (MLLM). By implementing a two-stage Supervised Fine-Tuning (SFT), we fully unlock the potential of MLLM, enabling it not only to generate detailed video descriptions but also to extract video features. Surprisingly, experimental results demonstrate that, compared to the CLIP-based models designed for retrieval and the popular MLLMs skilled in video captioning, our baseline shows competitive performance in both fine-grained video retrieval and video detailed captioning.
title CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
topic Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2501.00513