MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Huaying, Ni, Jian, Liu, Zheng, Wang, Yueze, Zhou, Junjie, Liang, Zhengyang, Zhao, Bo, Cao, Zhao, Dou, Zhicheng, Wen, Ji-Rong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911364794023936
author Yuan, Huaying
Ni, Jian
Liu, Zheng
Wang, Yueze
Zhou, Junjie
Liang, Zhengyang
Zhao, Bo
Cao, Zhao
Dou, Zhicheng
Wen, Ji-Rong
author_facet Yuan, Huaying
Ni, Jian
Liu, Zheng
Wang, Yueze
Zhou, Junjie
Liang, Zhengyang
Zhao, Bo
Cao, Zhao
Dou, Zhicheng
Wen, Ji-Rong
contents Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for evaluating whether key moments can be accurately accessed. To address this challenge, we propose MomentSeeker, a novel benchmark for long-video moment retrieval (LMVR), distinguished by the following features. First, it is created based on long and diverse videos, averaging over 1200 seconds in duration and collected from various domains, e.g., movie, anomaly, egocentric, and sports. Second, it covers a variety of real-world scenarios in three levels: global-level, event-level, object-level, covering common tasks like action recognition, object localization, and causal reasoning, etc. Third, it incorporates rich forms of queries, including text-only queries, image-conditioned queries, and video-conditioned queries. On top of MomentSeeker, we conduct comprehensive experiments for both generation-based approaches (directly using MLLMs) and retrieval-based approaches (leveraging video retrievers). Our results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning. We have publicly released MomentSeeker(https://yhy-2000.github.io/MomentSeeker/) to facilitate future research in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12558
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
Yuan, Huaying
Ni, Jian
Liu, Zheng
Wang, Yueze
Zhou, Junjie
Liang, Zhengyang
Zhao, Bo
Cao, Zhao
Dou, Zhicheng
Wen, Ji-Rong
Computer Vision and Pattern Recognition
Artificial Intelligence
Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for evaluating whether key moments can be accurately accessed. To address this challenge, we propose MomentSeeker, a novel benchmark for long-video moment retrieval (LMVR), distinguished by the following features. First, it is created based on long and diverse videos, averaging over 1200 seconds in duration and collected from various domains, e.g., movie, anomaly, egocentric, and sports. Second, it covers a variety of real-world scenarios in three levels: global-level, event-level, object-level, covering common tasks like action recognition, object localization, and causal reasoning, etc. Third, it incorporates rich forms of queries, including text-only queries, image-conditioned queries, and video-conditioned queries. On top of MomentSeeker, we conduct comprehensive experiments for both generation-based approaches (directly using MLLMs) and retrieval-based approaches (leveraging video retrievers). Our results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning. We have publicly released MomentSeeker(https://yhy-2000.github.io/MomentSeeker/) to facilitate future research in this area.
title MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2502.12558