OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Jiali, Deng, Xinran, Gu, Xin, Dai, Mengrui, Fan, Bing, Zhang, Zhipeng, Huang, Yan, Fan, Heng, Zhang, Libo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917956171071488
author Yao, Jiali
Deng, Xinran
Gu, Xin
Dai, Mengrui
Fan, Bing
Zhang, Zhipeng
Huang, Yan
Fan, Heng
Zhang, Libo
author_facet Yao, Jiali
Deng, Xinran
Gu, Xin
Dai, Mengrui
Fan, Bing
Zhang, Zhipeng
Huang, Yan
Fan, Heng
Zhang, Libo
contents In this paper, we propose spatio-temporal omni-object video grounding, dubbed OmniSTVG, a new STVG task that aims at localizing spatially and temporally all targets mentioned in the textual query from videos. Compared to classic STVG locating only a single target, OmniSTVG enables localization of not only an arbitrary number of text-referred targets but also their interacting counterparts in the query from the video, making it more flexible and practical in real scenarios for comprehensive understanding. In order to facilitate exploration of OmniSTVG, we introduce BOSTVG, a large-scale benchmark dedicated to OmniSTVG. Specifically, our BOSTVG consists of 10,018 videos with 10.2M frames and covers a wide selection of 287 classes from diverse scenarios. Each sequence in BOSTVG, paired with a free-form textual query, encompasses a varying number of targets ranging from 1 to 10. To ensure high quality, each video is manually annotated with meticulous inspection and refinement. To our best knowledge, BOSTVG is to date the first and the largest benchmark for OmniSTVG. To encourage future research, we introduce a simple yet effective approach, named OmniTube, which, drawing inspiration from Transformer-based STVG methods, is specially designed for OmniSTVG and demonstrates promising results. By releasing BOSTVG, we hope to go beyond classic STVG by locating every object appearing in the query for more comprehensive understanding, opening up a new direction for STVG. Our benchmark, model, and results will be released at https://github.com/JellyYao3000/OmniSTVG.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10500
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
Yao, Jiali
Deng, Xinran
Gu, Xin
Dai, Mengrui
Fan, Bing
Zhang, Zhipeng
Huang, Yan
Fan, Heng
Zhang, Libo
Computer Vision and Pattern Recognition
In this paper, we propose spatio-temporal omni-object video grounding, dubbed OmniSTVG, a new STVG task that aims at localizing spatially and temporally all targets mentioned in the textual query from videos. Compared to classic STVG locating only a single target, OmniSTVG enables localization of not only an arbitrary number of text-referred targets but also their interacting counterparts in the query from the video, making it more flexible and practical in real scenarios for comprehensive understanding. In order to facilitate exploration of OmniSTVG, we introduce BOSTVG, a large-scale benchmark dedicated to OmniSTVG. Specifically, our BOSTVG consists of 10,018 videos with 10.2M frames and covers a wide selection of 287 classes from diverse scenarios. Each sequence in BOSTVG, paired with a free-form textual query, encompasses a varying number of targets ranging from 1 to 10. To ensure high quality, each video is manually annotated with meticulous inspection and refinement. To our best knowledge, BOSTVG is to date the first and the largest benchmark for OmniSTVG. To encourage future research, we introduce a simple yet effective approach, named OmniTube, which, drawing inspiration from Transformer-based STVG methods, is specially designed for OmniSTVG and demonstrates promising results. By releasing BOSTVG, we hope to go beyond classic STVG by locating every object appearing in the query for more comprehensive understanding, opening up a new direction for STVG. Our benchmark, model, and results will be released at https://github.com/JellyYao3000/OmniSTVG.
title OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10500