Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Yifan, Liu, Yuanzhe, Zhu, Jingyuan, Cao, Xu, Zhang, Xiaofeng, He, Yixiao, Ye, Wenming, Rehg, James Matthew, Lourentzou, Ismini
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914232011849728
author Shen, Yifan
Liu, Yuanzhe
Zhu, Jingyuan
Cao, Xu
Zhang, Xiaofeng
He, Yixiao
Ye, Wenming
Rehg, James Matthew
Lourentzou, Ismini
author_facet Shen, Yifan
Liu, Yuanzhe
Zhu, Jingyuan
Cao, Xu
Zhang, Xiaofeng
He, Yixiao
Ye, Wenming
Rehg, James Matthew
Lourentzou, Ismini
contents Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment are required. In this work, we introduce SpatialReasoner-R1, a vision-language reasoning model designed to address these limitations. To construct high-quality supervision for spatial reasoning, we design a Multi-Model Monte Carlo Tree Search (M3CTS) method that generates diverse, logically consistent Long Chain-of-Thought (LongCOT) reasoning trajectories. In addition, we propose a fine-grained Direct Preference Optimization (fDPO) method that introduces segment-specific preference granularity for descriptive grounding and logical reasoning, guided by a spatial reward mechanism that evaluates candidate responses based on visual consistency, spatial grounding, and logical coherence. Experimental results demonstrate that fDPO achieves relative performance gains of 4.1% and 9.0% over standard DPO on spatial qualitative and quantitative tasks, respectively. SpatialReasoner-R1, trained with fDPO, sets a new SoTA on SpatialRGPT-Bench, outperforming the strongest baseline by 9.4% in average accuracy, while maintaining competitive performance on general vision-language tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21656
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
Shen, Yifan
Liu, Yuanzhe
Zhu, Jingyuan
Cao, Xu
Zhang, Xiaofeng
He, Yixiao
Ye, Wenming
Rehg, James Matthew
Lourentzou, Ismini
Computer Vision and Pattern Recognition
Computation and Language
Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment are required. In this work, we introduce SpatialReasoner-R1, a vision-language reasoning model designed to address these limitations. To construct high-quality supervision for spatial reasoning, we design a Multi-Model Monte Carlo Tree Search (M3CTS) method that generates diverse, logically consistent Long Chain-of-Thought (LongCOT) reasoning trajectories. In addition, we propose a fine-grained Direct Preference Optimization (fDPO) method that introduces segment-specific preference granularity for descriptive grounding and logical reasoning, guided by a spatial reward mechanism that evaluates candidate responses based on visual consistency, spatial grounding, and logical coherence. Experimental results demonstrate that fDPO achieves relative performance gains of 4.1% and 9.0% over standard DPO on spatial qualitative and quantitative tasks, respectively. SpatialReasoner-R1, trained with fDPO, sets a new SoTA on SpatialRGPT-Bench, outperforming the strongest baseline by 9.4% in average accuracy, while maintaining competitive performance on general vision-language tasks.
title Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.21656