ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yifan, Yin, Yingda, Zhu, Lingting, Chen, Weikai, Qian, Shengju, Wang, Xin, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
by: Jin, Li, et al.
Published: (2026)
by: Jin, Li, et al.
Published: (2026)
V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
by: Luo, Yang, et al.
Published: (2025)
by: Luo, Yang, et al.
Published: (2025)
MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Agentic Post-Processing
by: Zhu, Lingting, et al.
Published: (2025)
by: Zhu, Lingting, et al.
Published: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation
by: Liu, Chenyu, et al.
Published: (2025)
by: Liu, Chenyu, et al.
Published: (2025)
MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation
by: Chen, Jinnan, et al.
Published: (2025)
by: Chen, Jinnan, et al.
Published: (2025)
Prompt Highlighter: Interactive Control for Multi-Modal LLMs
by: Zhang, Yuechen, et al.
Published: (2023)
by: Zhang, Yuechen, et al.
Published: (2023)
SAP: Segment Any 4K Panorama
by: Jiang, Lutao, et al.
Published: (2026)
by: Jiang, Lutao, et al.
Published: (2026)
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
by: Yue, Yang, et al.
Published: (2025)
by: Yue, Yang, et al.
Published: (2025)
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation
by: Wen, Junwei, et al.
Published: (2026)
by: Wen, Junwei, et al.
Published: (2026)
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments
by: Lin, Mingxian, et al.
Published: (2025)
by: Lin, Mingxian, et al.
Published: (2025)
Rethinking Chain-of-Thought Reasoning for Videos
by: Zhong, Yiwu, et al.
Published: (2025)
by: Zhong, Yiwu, et al.
Published: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
by: Tao, Xingjian, et al.
Published: (2026)
by: Tao, Xingjian, et al.
Published: (2026)
AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer
by: Zhu, Lingting, et al.
Published: (2026)
by: Zhu, Lingting, et al.
Published: (2026)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
by: Fu, Xingyu, et al.
Published: (2025)
by: Fu, Xingyu, et al.
Published: (2025)
Generating Storytelling Images with Rich Chains-of-Reasoning
by: Song, Xiujie, et al.
Published: (2025)
by: Song, Xiujie, et al.
Published: (2025)
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning
by: Schroeder, Philip, et al.
Published: (2026)
by: Schroeder, Philip, et al.
Published: (2026)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
LumiTex: Towards High-Fidelity PBR Texture Generation with Illumination Context
by: Bao, Jingzhi, et al.
Published: (2025)
by: Bao, Jingzhi, et al.
Published: (2025)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
by: Luo, Fuwen, et al.
Published: (2025)
by: Luo, Fuwen, et al.
Published: (2025)
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging
by: Fu, Zihang, et al.
Published: (2026)
by: Fu, Zihang, et al.
Published: (2026)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
by: Cai, Zhenyang, et al.
Published: (2025)
by: Cai, Zhenyang, et al.
Published: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
Large Material Gaussian Model for Relightable 3D Generation
by: Ye, Jingrui, et al.
Published: (2025)
by: Ye, Jingrui, et al.
Published: (2025)
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval
by: Xu, Mingjun, et al.
Published: (2025)
by: Xu, Mingjun, et al.
Published: (2025)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
by: Shi, Weikang, et al.
Published: (2025)
by: Shi, Weikang, et al.
Published: (2025)
LaRe: Latent Refocusing for Multimodal Reasoning
by: Ma, Jizheng, et al.
Published: (2025)
by: Ma, Jizheng, et al.
Published: (2025)
[De|Re]constructing VLMs' Reasoning in Counting
by: Alghisi, Simone, et al.
Published: (2025)
by: Alghisi, Simone, et al.
Published: (2025)
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
by: Huang, Yixu, et al.
Published: (2026)
by: Huang, Yixu, et al.
Published: (2026)
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
by: Xu, Qinwu, et al.
Published: (2026)
by: Xu, Qinwu, et al.
Published: (2026)
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
by: Qi, Yukun, et al.
Published: (2025)
by: Qi, Yukun, et al.
Published: (2025)
ReMI: A Dataset for Reasoning with Multiple Images
by: Kazemi, Mehran, et al.
Published: (2024)
by: Kazemi, Mehran, et al.
Published: (2024)
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
by: Wang, Yan, et al.
Published: (2024)
by: Wang, Yan, et al.
Published: (2024)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Similar Items
-
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
by: Jin, Li, et al.
Published: (2026) -
V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
by: Luo, Yang, et al.
Published: (2025) -
MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Agentic Post-Processing
by: Zhu, Lingting, et al.
Published: (2025) -
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025) -
CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation
by: Liu, Chenyu, et al.
Published: (2025)