Guardado en:
Detalles Bibliográficos
Autores principales: Li, Yifan, Yin, Yingda, Zhu, Lingting, Chen, Weikai, Qian, Shengju, Wang, Xin, Fu, Yanwei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2512.02835
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912743552974848
author Li, Yifan
Yin, Yingda
Zhu, Lingting
Chen, Weikai
Qian, Shengju
Wang, Xin
Fu, Yanwei
author_facet Li, Yifan
Yin, Yingda
Zhu, Lingting
Chen, Weikai
Qian, Shengju
Wang, Xin
Fu, Yanwei
contents Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances. Yet existing solutions generally collapse these factors into simplified reasoning with latent embeddings, rendering the reasoning chain opaque and essentially intractable. We therefore adopt an explicit decomposition perspective and introduce ReVSeg, which executes reasoning as sequential decisions in the native interface of pretrained vision language models (VLMs). Rather than folding all reasoning into a single-step prediction, ReVSeg executes three explicit operations -- semantics interpretation, temporal evidence selection, and spatial grounding -- aligning pretrained capabilities. We further employ reinforcement learning to optimize the multi-step reasoning chain, enabling the model to self-refine its decision quality from outcome-driven signals. Experimental results demonstrate that ReVSeg attains state-of-the-art performances on standard video object segmentation benchmarks and yields interpretable reasoning trajectories. Project page is available at https://clementine24.github.io/ReVSeg/ .
format Preprint
id arxiv_https___arxiv_org_abs_2512_02835
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
Li, Yifan
Yin, Yingda
Zhu, Lingting
Chen, Weikai
Qian, Shengju
Wang, Xin
Fu, Yanwei
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances. Yet existing solutions generally collapse these factors into simplified reasoning with latent embeddings, rendering the reasoning chain opaque and essentially intractable. We therefore adopt an explicit decomposition perspective and introduce ReVSeg, which executes reasoning as sequential decisions in the native interface of pretrained vision language models (VLMs). Rather than folding all reasoning into a single-step prediction, ReVSeg executes three explicit operations -- semantics interpretation, temporal evidence selection, and spatial grounding -- aligning pretrained capabilities. We further employ reinforcement learning to optimize the multi-step reasoning chain, enabling the model to self-refine its decision quality from outcome-driven signals. Experimental results demonstrate that ReVSeg attains state-of-the-art performances on standard video object segmentation benchmarks and yields interpretable reasoning trajectories. Project page is available at https://clementine24.github.io/ReVSeg/ .
title ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.02835