Object-Centric Framework for Video Moment Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zongyao, Wong, Yongkang, Yamazaki, Satoshi, Liu, Jianquan, Kankanhalli, Mohan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918257364041728
author Li, Zongyao
Wong, Yongkang
Yamazaki, Satoshi
Liu, Jianquan
Kankanhalli, Mohan
author_facet Li, Zongyao
Wong, Yongkang
Yamazaki, Satoshi
Liu, Jianquan
Kankanhalli, Mohan
contents Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing moments described by object-oriented queries involving specific entities and their interactions. In particular, temporal dynamics at the object level have been largely overlooked, limiting the effectiveness of existing approaches in scenarios requiring detailed object-level reasoning. To address this limitation, we propose a novel object-centric framework for moment retrieval. Our method first extracts query-relevant objects using a scene graph parser and then generates scene graphs from video frames to represent these objects and their relationships. Based on the scene graphs, we construct object-level feature sequences that encode rich visual and semantic information. These sequences are processed by a relational tracklet transformer, which models spatio-temporal correlations among objects over time. By explicitly capturing object-level state changes, our framework enables more accurate localization of moments aligned with object-oriented queries. We evaluated our method on three benchmarks: Charades-STA, QVHighlights, and TACoS. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across all benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18448
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Object-Centric Framework for Video Moment Retrieval
Li, Zongyao
Wong, Yongkang
Yamazaki, Satoshi
Liu, Jianquan
Kankanhalli, Mohan
Computer Vision and Pattern Recognition
Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing moments described by object-oriented queries involving specific entities and their interactions. In particular, temporal dynamics at the object level have been largely overlooked, limiting the effectiveness of existing approaches in scenarios requiring detailed object-level reasoning. To address this limitation, we propose a novel object-centric framework for moment retrieval. Our method first extracts query-relevant objects using a scene graph parser and then generates scene graphs from video frames to represent these objects and their relationships. Based on the scene graphs, we construct object-level feature sequences that encode rich visual and semantic information. These sequences are processed by a relational tracklet transformer, which models spatio-temporal correlations among objects over time. By explicitly capturing object-level state changes, our framework enables more accurate localization of moments aligned with object-oriented queries. We evaluated our method on three benchmarks: Charades-STA, QVHighlights, and TACoS. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across all benchmarks.
title Object-Centric Framework for Video Moment Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.18448