LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Qifeng, Liang, Hao, Han, Zhaoyang, Dong, Hejun, Qiang, Meiyi, An, Ruichuan, Xu, Quanqing, Cui, Bin, Zhang, Wentao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917246269390848
author Cai, Qifeng
Liang, Hao
Han, Zhaoyang
Dong, Hejun
Qiang, Meiyi
An, Ruichuan
Xu, Quanqing
Cui, Bin
Zhang, Wentao
author_facet Cai, Qifeng
Liang, Hao
Han, Zhaoyang
Dong, Hejun
Qiang, Meiyi
An, Ruichuan
Xu, Quanqing
Cui, Bin
Zhang, Wentao
contents Long videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning. However, existing benchmarks suffer from limited video duration, low-quality captions, and coarse annotation granularity, which hinder the evaluation of advanced video-text retrieval methods. To address these limitations, we introduce LoVR, a benchmark specifically designed for long video-text retrieval. LoVR contains 467 long videos and over 40,804 fine-grained clips with high-quality captions. To overcome the issue of poor machine-generated annotations, we propose an efficient caption generation framework that integrates VLM automatic generation, caption quality scoring, and dynamic refinement. This pipeline improves annotation accuracy while maintaining scalability. Furthermore, we introduce a semantic fusion method to generate coherent full-video captions without losing important contextual information. Our benchmark introduces longer videos, more detailed captions, and a larger-scale dataset, presenting new challenges for video understanding and retrieval. Extensive experiments on various advanced embedding models demonstrate that LoVR is a challenging benchmark, revealing the limitations of current approaches and providing valuable insights for future research. We release the code and dataset link at https://lovrbench.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2505_13928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
Cai, Qifeng
Liang, Hao
Han, Zhaoyang
Dong, Hejun
Qiang, Meiyi
An, Ruichuan
Xu, Quanqing
Cui, Bin
Zhang, Wentao
Computer Vision and Pattern Recognition
Information Retrieval
Long videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning. However, existing benchmarks suffer from limited video duration, low-quality captions, and coarse annotation granularity, which hinder the evaluation of advanced video-text retrieval methods. To address these limitations, we introduce LoVR, a benchmark specifically designed for long video-text retrieval. LoVR contains 467 long videos and over 40,804 fine-grained clips with high-quality captions. To overcome the issue of poor machine-generated annotations, we propose an efficient caption generation framework that integrates VLM automatic generation, caption quality scoring, and dynamic refinement. This pipeline improves annotation accuracy while maintaining scalability. Furthermore, we introduce a semantic fusion method to generate coherent full-video captions without losing important contextual information. Our benchmark introduces longer videos, more detailed captions, and a larger-scale dataset, presenting new challenges for video understanding and retrieval. Extensive experiments on various advanced embedding models demonstrate that LoVR is a challenging benchmark, revealing the limitations of current approaches and providing valuable insights for future research. We release the code and dataset link at https://lovrbench.github.io/
title LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
topic Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2505.13928