SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Xianda, Zhang, Ruijun, Duan, Yiqun, He, Yuhang, Nie, Dujun, Huang, Wenke, Zhang, Chenming, Liu, Shuai, Zhao, Hao, Chen, Long
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908380942041088
author Guo, Xianda
Zhang, Ruijun
Duan, Yiqun
He, Yuhang
Nie, Dujun
Huang, Wenke
Zhang, Chenming
Liu, Shuai
Zhao, Hao
Chen, Long
author_facet Guo, Xianda
Zhang, Ruijun
Duan, Yiqun
He, Yuhang
Nie, Dujun
Huang, Wenke
Zhang, Chenming
Liu, Shuai
Zhao, Hao
Chen, Long
contents Accurate spatial reasoning in outdoor environments - covering geometry, object pose, and inter-object relationships - is fundamental to downstream tasks such as mapping, motion forecasting, and high-level planning in autonomous driving. We introduce SURDS, a large-scale benchmark designed to systematically evaluate the spatial reasoning capabilities of vision language models (VLMs). Built on the nuScenes dataset, SURDS comprises 41,080 vision-question-answer training instances and 9,250 evaluation samples, spanning six spatial categories: orientation, depth estimation, pixel-level localization, pairwise distance, lateral ordering, and front-behind relations. We benchmark leading general-purpose VLMs, including GPT, Gemini, and Qwen, revealing persistent limitations in fine-grained spatial understanding. To address these deficiencies, we go beyond static evaluation and explore whether alignment techniques can improve spatial reasoning performance. Specifically, we propose a reinforcement learning-based alignment scheme leveraging spatially grounded reward signals - capturing both perception-level accuracy (location) and reasoning consistency (logic). We further incorporate final-answer correctness and output-format rewards to guide fine-grained policy adaptation. Our GRPO-aligned variant achieves an overall score of 40.80 in the SURDS benchmark. Notably, it outperforms proprietary systems such as GPT-4o (13.30) and Gemini-2.0-flash (35.71). To our best knowledge, this is the first study to demonstrate that reinforcement learning-based alignment can significantly and consistently enhance the spatial reasoning capabilities of VLMs in real-world driving contexts. We release the SURDS benchmark, evaluation toolkit, and GRPO alignment code through: https://github.com/XiandaGuo/Drive-MLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13112
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
Guo, Xianda
Zhang, Ruijun
Duan, Yiqun
He, Yuhang
Nie, Dujun
Huang, Wenke
Zhang, Chenming
Liu, Shuai
Zhao, Hao
Chen, Long
Computer Vision and Pattern Recognition
Accurate spatial reasoning in outdoor environments - covering geometry, object pose, and inter-object relationships - is fundamental to downstream tasks such as mapping, motion forecasting, and high-level planning in autonomous driving. We introduce SURDS, a large-scale benchmark designed to systematically evaluate the spatial reasoning capabilities of vision language models (VLMs). Built on the nuScenes dataset, SURDS comprises 41,080 vision-question-answer training instances and 9,250 evaluation samples, spanning six spatial categories: orientation, depth estimation, pixel-level localization, pairwise distance, lateral ordering, and front-behind relations. We benchmark leading general-purpose VLMs, including GPT, Gemini, and Qwen, revealing persistent limitations in fine-grained spatial understanding. To address these deficiencies, we go beyond static evaluation and explore whether alignment techniques can improve spatial reasoning performance. Specifically, we propose a reinforcement learning-based alignment scheme leveraging spatially grounded reward signals - capturing both perception-level accuracy (location) and reasoning consistency (logic). We further incorporate final-answer correctness and output-format rewards to guide fine-grained policy adaptation. Our GRPO-aligned variant achieves an overall score of 40.80 in the SURDS benchmark. Notably, it outperforms proprietary systems such as GPT-4o (13.30) and Gemini-2.0-flash (35.71). To our best knowledge, this is the first study to demonstrate that reinforcement learning-based alignment can significantly and consistently enhance the spatial reasoning capabilities of VLMs in real-world driving contexts. We release the SURDS benchmark, evaluation toolkit, and GRPO alignment code through: https://github.com/XiandaGuo/Drive-MLLM.
title SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.13112