VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Zishan, Guo, Yifu, Lu, Yuquan, Yang, Fengyu, Li, Junxin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912720460185600
author Xu, Zishan
Guo, Yifu
Lu, Yuquan
Yang, Fengyu
Li, Junxin
author_facet Xu, Zishan
Guo, Yifu
Lu, Yuquan
Yang, Fengyu
Li, Junxin
contents Traditional video reasoning segmentation methods rely on supervised fine-tuning, which limits generalization to out-of-distribution scenarios and lacks explicit reasoning. To address this, we propose \textbf{VideoSeg-R1}, the first framework to introduce reinforcement learning into video reasoning segmentation. It adopts a decoupled architecture that formulates the task as joint referring image segmentation and video mask propagation. It comprises three stages: (1) A hierarchical text-guided frame sampler to emulate human attention; (2) A reasoning model that produces spatial cues along with explicit reasoning chains; and (3) A segmentation-propagation stage using SAM2 and XMem. A task difficulty-aware mechanism adaptively controls reasoning length for better efficiency and accuracy. Extensive evaluations on multiple benchmarks demonstrate that VideoSeg-R1 achieves state-of-the-art performance in complex video reasoning and segmentation tasks. The code will be publicly available at https://github.com/euyis1019/VideoSeg-R1.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
Xu, Zishan
Guo, Yifu
Lu, Yuquan
Yang, Fengyu
Li, Junxin
Computer Vision and Pattern Recognition
Traditional video reasoning segmentation methods rely on supervised fine-tuning, which limits generalization to out-of-distribution scenarios and lacks explicit reasoning. To address this, we propose \textbf{VideoSeg-R1}, the first framework to introduce reinforcement learning into video reasoning segmentation. It adopts a decoupled architecture that formulates the task as joint referring image segmentation and video mask propagation. It comprises three stages: (1) A hierarchical text-guided frame sampler to emulate human attention; (2) A reasoning model that produces spatial cues along with explicit reasoning chains; and (3) A segmentation-propagation stage using SAM2 and XMem. A task difficulty-aware mechanism adaptively controls reasoning length for better efficiency and accuracy. Extensive evaluations on multiple benchmarks demonstrate that VideoSeg-R1 achieves state-of-the-art performance in complex video reasoning and segmentation tasks. The code will be publicly available at https://github.com/euyis1019/VideoSeg-R1.
title VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.16077