Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Junfei, Guan, Jian, Feng, Kaituo, Liu, Qiang, Wu, Shu, Wang, Liang, Wu, Wei, Tan, Tieniu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915350490120192
author Wu, Junfei
Guan, Jian
Feng, Kaituo
Liu, Qiang
Wu, Shu
Wang, Liang
Wu, Wei
Tan, Tieniu
author_facet Wu, Junfei
Guan, Jian
Feng, Kaituo
Liu, Qiang
Wu, Shu
Wang, Liang
Wu, Wei
Tan, Tieniu
contents As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods primarily approach multimodal reasoning in a straightforward, text-centric manner, where both reasoning and answer derivation are conducted purely through text, with the only difference being the presence of multimodal input. As a result, these methods often encounter fundamental limitations in spatial reasoning tasks that demand precise geometric understanding and continuous spatial tracking-capabilities that humans achieve through mental visualization and manipulation. To address the limitations, we propose drawing to reason in space, a novel paradigm that enables LVLMs to reason through elementary drawing operations in the visual space. By equipping models with basic drawing operations, including annotating bounding boxes and drawing auxiliary lines, we empower them to express and analyze spatial relationships through direct visual manipulation, meanwhile avoiding the performance ceiling imposed by specialized perception tools in previous tool-integrated reasoning approaches. To cultivate this capability, we develop a three-stage training framework: cold-start training with synthetic data to establish basic drawing abilities, reflective rejection sampling to enhance self-reflection behaviors, and reinforcement learning to directly optimize for target rewards. Extensive experiments demonstrate that our model, named VILASR, consistently outperforms existing methods across diverse spatial reasoning benchmarks, involving maze navigation, static spatial reasoning, video-based reasoning, and multi-view-based reasoning tasks, with an average improvement of 18.4%.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09965
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
Wu, Junfei
Guan, Jian
Feng, Kaituo
Liu, Qiang
Wu, Shu
Wang, Liang
Wu, Wei
Tan, Tieniu
Computer Vision and Pattern Recognition
Artificial Intelligence
As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods primarily approach multimodal reasoning in a straightforward, text-centric manner, where both reasoning and answer derivation are conducted purely through text, with the only difference being the presence of multimodal input. As a result, these methods often encounter fundamental limitations in spatial reasoning tasks that demand precise geometric understanding and continuous spatial tracking-capabilities that humans achieve through mental visualization and manipulation. To address the limitations, we propose drawing to reason in space, a novel paradigm that enables LVLMs to reason through elementary drawing operations in the visual space. By equipping models with basic drawing operations, including annotating bounding boxes and drawing auxiliary lines, we empower them to express and analyze spatial relationships through direct visual manipulation, meanwhile avoiding the performance ceiling imposed by specialized perception tools in previous tool-integrated reasoning approaches. To cultivate this capability, we develop a three-stage training framework: cold-start training with synthetic data to establish basic drawing abilities, reflective rejection sampling to enhance self-reflection behaviors, and reinforcement learning to directly optimize for target rewards. Extensive experiments demonstrate that our model, named VILASR, consistently outperforms existing methods across diverse spatial reasoning benchmarks, involving maze navigation, static spatial reasoning, video-based reasoning, and multi-view-based reasoning tasks, with an average improvement of 18.4%.
title Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.09965