SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xiaolong, Liu, Yifei, Gong, Ziyang, Li, Jiarui, Zhao, Qiyue, Niu, Muyao, Gao, Yuanyuan, Ma, Le, Yang, Xue, Zhang, Hongjie, Zhong, Zhihang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917520190996480
author Zhou, Xiaolong
Liu, Yifei
Gong, Ziyang
Li, Jiarui
Zhao, Qiyue
Niu, Muyao
Gao, Yuanyuan
Ma, Le
Yang, Xue
Zhang, Hongjie
Zhong, Zhihang
author_facet Zhou, Xiaolong
Liu, Yifei
Gong, Ziyang
Li, Jiarui
Zhao, Qiyue
Niu, Muyao
Gao, Yuanyuan
Ma, Le
Yang, Xue
Zhang, Hongjie
Zhong, Zhihang
contents Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pristine visual inputs and overlook the degradations that commonly occur in real-world deployment, such as motion blur, low light, adverse weather, lens distortion, and compression artifacts. This raises a fundamental question: how robust is the spatial intelligence of current MLLMs when visual observations are imperfect? To answer this question, we introduce SpaceDG, the first large-scale dataset for degradation-aware spatial understanding. It is constructed with a physically grounded degradation synthesis engine that embeds degradation formation process into 3D Gaussian Splatting (3DGS) rendering, enabling realistic simulation of nine degradation types. The resulting dataset contains approximately 1M QA pairs from nearly 1,000 indoor scenes. We further introduce SpaceDG-Bench, an human-verified benchmark with 1,102 questions spanning 11 reasoning categories and 9 visual degradation types, yielding over 10K VQA instances. Evaluating 25 open- and closed-source MLLMs reveals that visual degradations consistently and substantially impair spatial reasoning, exposing a critical robustness gap. Finally, we show that finetuning on SpaceDG markedly improves degradation robustness and can even surpass human performance under degraded conditions without any performance drop on clean images, highlighting the promise of degradation-aware training for robust spatial intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22536
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
Zhou, Xiaolong
Liu, Yifei
Gong, Ziyang
Li, Jiarui
Zhao, Qiyue
Niu, Muyao
Gao, Yuanyuan
Ma, Le
Yang, Xue
Zhang, Hongjie
Zhong, Zhihang
Computer Vision and Pattern Recognition
Computation and Language
Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pristine visual inputs and overlook the degradations that commonly occur in real-world deployment, such as motion blur, low light, adverse weather, lens distortion, and compression artifacts. This raises a fundamental question: how robust is the spatial intelligence of current MLLMs when visual observations are imperfect? To answer this question, we introduce SpaceDG, the first large-scale dataset for degradation-aware spatial understanding. It is constructed with a physically grounded degradation synthesis engine that embeds degradation formation process into 3D Gaussian Splatting (3DGS) rendering, enabling realistic simulation of nine degradation types. The resulting dataset contains approximately 1M QA pairs from nearly 1,000 indoor scenes. We further introduce SpaceDG-Bench, an human-verified benchmark with 1,102 questions spanning 11 reasoning categories and 9 visual degradation types, yielding over 10K VQA instances. Evaluating 25 open- and closed-source MLLMs reveals that visual degradations consistently and substantially impair spatial reasoning, exposing a critical robustness gap. Finally, we show that finetuning on SpaceDG markedly improves degradation robustness and can even surpass human performance under degraded conditions without any performance drop on clean images, highlighting the promise of degradation-aware training for robust spatial intelligence.
title SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2605.22536