RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yi, Cao, Jiawang, Wu, Yongliang, Li, Bozheng, Tang, Licheng, Ji, Yangguang, Wu, Chong, Wu, Jay, Zhu, Wenbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913876861255680
author Lu, Yi
Cao, Jiawang
Wu, Yongliang
Li, Bozheng
Tang, Licheng
Ji, Yangguang
Wu, Chong
Wu, Jay
Zhu, Wenbo
author_facet Lu, Yi
Cao, Jiawang
Wu, Yongliang
Li, Bozheng
Tang, Licheng
Ji, Yangguang
Wu, Chong
Wu, Jay
Zhu, Wenbo
contents Multi-modal Large Language Models (MLLMs) have demonstrated remarkable reasoning capability while lack explicit mechanisms for visual grounding and segmentation, creating a gap between cognitive reasoning and visual perception. To bridge this gap, we introduce Reasoning Segmentation via Visual Prompting (RSVP), a novel framework that unifies multi-step multimodal reasoning with grounded visual understanding. RSVP is a two-stage structuralized framework that integrates reasoning-driven localization with segmentation refinement. In the reasoning stage, RSVP employs multimodal chain-of-thought visual prompts to help MLLMs understand queries and infer targets, generating interpretable region proposals that enhance visual grounding. In segmentation stage, RSVP refines these proposals with a Vision-Language Segmentation Module (VLSM), seamlessly integrates textual and visual cues to produce precise segmentation masks. By explicitly modelling the interaction between multimodal reasoning and segmentation, RSVP introduces a new paradigm for interpretable reasoning segmentation. It exploits MLLMs' inherent localization capabilities, enabling the models to not only reason about objects but also generate structured visual representations. Our extensive experiments demonstrate that RSVP achieves state-of-the-art performance, surpasses state-of-the-art methods by up to +6.5 gIoU and +9.2 cIoU on ReasonSeg, and achieves 49.7 mAP on SegInW under zero-shot settings. These results validate RSVP as an effective and scalable framework for integrating cognitive reasoning with structured visual understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
Lu, Yi
Cao, Jiawang
Wu, Yongliang
Li, Bozheng
Tang, Licheng
Ji, Yangguang
Wu, Chong
Wu, Jay
Zhu, Wenbo
Computer Vision and Pattern Recognition
Artificial Intelligence
Multi-modal Large Language Models (MLLMs) have demonstrated remarkable reasoning capability while lack explicit mechanisms for visual grounding and segmentation, creating a gap between cognitive reasoning and visual perception. To bridge this gap, we introduce Reasoning Segmentation via Visual Prompting (RSVP), a novel framework that unifies multi-step multimodal reasoning with grounded visual understanding. RSVP is a two-stage structuralized framework that integrates reasoning-driven localization with segmentation refinement. In the reasoning stage, RSVP employs multimodal chain-of-thought visual prompts to help MLLMs understand queries and infer targets, generating interpretable region proposals that enhance visual grounding. In segmentation stage, RSVP refines these proposals with a Vision-Language Segmentation Module (VLSM), seamlessly integrates textual and visual cues to produce precise segmentation masks. By explicitly modelling the interaction between multimodal reasoning and segmentation, RSVP introduces a new paradigm for interpretable reasoning segmentation. It exploits MLLMs' inherent localization capabilities, enabling the models to not only reason about objects but also generate structured visual representations. Our extensive experiments demonstrate that RSVP achieves state-of-the-art performance, surpasses state-of-the-art methods by up to +6.5 gIoU and +9.2 cIoU on ReasonSeg, and achieves 49.7 mAP on SegInW under zero-shot settings. These results validate RSVP as an effective and scalable framework for integrating cognitive reasoning with structured visual understanding.
title RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.04277