Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gu, Xin, Zhang, Haoji, Fan, Qihang, Niu, Jingxuan, Zhang, Zhipeng, Zhang, Libo, Chen, Guang, Chen, Fan, Wen, Longyin, Zhu, Sijie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909926607028224
author Gu, Xin
Zhang, Haoji
Fan, Qihang
Niu, Jingxuan
Zhang, Zhipeng
Zhang, Libo
Chen, Guang
Chen, Fan
Wen, Longyin
Zhu, Sijie
author_facet Gu, Xin
Zhang, Haoji
Fan, Qihang
Niu, Jingxuan
Zhang, Zhipeng
Zhang, Libo
Chen, Guang
Chen, Fan
Wen, Longyin
Zhu, Sijie
contents Spatio-temporal video grounding (STVG) requires localizing a target object in untrimmed videos both temporally and spatially from natural language descriptions. Despite their strong language understanding, multimodal large language models (MLLMs) underperform on STVG due to misaligned training objectives and weak fine-grained region-word alignment in standard visual encoders. To address this, we propose STVG-o1, the first framework that enables off-the-shelf MLLMs to achieve state-of-the-art STVG performance without any architectural modifications. Our method introduces a bounding-box chain-of-thought mechanism that explicitly reasons about spatio-temporal locations in an intermediate step before producing the final prediction. We further design a multi-dimensional reinforcement reward function consisting of format, consistency, temporal, spatial, and think rewards, which provides geometry-aware supervision through reinforcement fine-tuning. Evaluated on HCSTVG-v1/v2 and VidSTG, STVG-o1 sets new state-of-the-art results on HCSTVG, outperforming the best task-specific method by 7.3\% m\_tIoU on HCSTVG-v1, matching specialized models on VidSTG, and surpassing all existing MLLM-based approaches by large margins. It also demonstrates strong open-vocabulary generalization across datasets, establishing MLLMs as viable and powerful backbones for precise spatio-temporal grounding. Our code and models will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21375
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
Gu, Xin
Zhang, Haoji
Fan, Qihang
Niu, Jingxuan
Zhang, Zhipeng
Zhang, Libo
Chen, Guang
Chen, Fan
Wen, Longyin
Zhu, Sijie
Computer Vision and Pattern Recognition
Spatio-temporal video grounding (STVG) requires localizing a target object in untrimmed videos both temporally and spatially from natural language descriptions. Despite their strong language understanding, multimodal large language models (MLLMs) underperform on STVG due to misaligned training objectives and weak fine-grained region-word alignment in standard visual encoders. To address this, we propose STVG-o1, the first framework that enables off-the-shelf MLLMs to achieve state-of-the-art STVG performance without any architectural modifications. Our method introduces a bounding-box chain-of-thought mechanism that explicitly reasons about spatio-temporal locations in an intermediate step before producing the final prediction. We further design a multi-dimensional reinforcement reward function consisting of format, consistency, temporal, spatial, and think rewards, which provides geometry-aware supervision through reinforcement fine-tuning. Evaluated on HCSTVG-v1/v2 and VidSTG, STVG-o1 sets new state-of-the-art results on HCSTVG, outperforming the best task-specific method by 7.3\% m\_tIoU on HCSTVG-v1, matching specialized models on VidSTG, and surpassing all existing MLLM-based approaches by large margins. It also demonstrates strong open-vocabulary generalization across datasets, establishing MLLMs as viable and powerful backbones for precise spatio-temporal grounding. Our code and models will be released.
title Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21375