SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xiaolong, Liu, Yifei, Gong, Ziyang, Li, Jiarui, Zhao, Qiyue, Niu, Muyao, Gao, Yuanyuan, Ma, Le, Yang, Xue, Zhang, Hongjie, Zhong, Zhihang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
by: Gao, Yuanyuan, et al.
Published: (2026)
by: Gao, Yuanyuan, et al.
Published: (2026)
PhotoFlow: Agentic 3D Virtual Photography Missions
by: Guo, Jiarui, et al.
Published: (2026)
by: Guo, Jiarui, et al.
Published: (2026)
R3-Avatar: Record and Retrieve Temporal Codebook for Reconstructing Photorealistic Human Avatars
by: Zhan, Yifan, et al.
Published: (2025)
by: Zhan, Yifan, et al.
Published: (2025)
KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter
by: Zhan, Yifan, et al.
Published: (2024)
by: Zhan, Yifan, et al.
Published: (2024)
Representation Learning of Limit Order Book: A Comprehensive Study and Benchmarking
by: Zhong, Muyao, et al.
Published: (2025)
by: Zhong, Muyao, et al.
Published: (2025)
Motion-Aware Animatable Gaussian Avatars Deblurring
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
Proxy-GS: Unified Occlusion Priors for Training and Inference in Structured 3D Gaussian Splatting
by: Gao, Yuanyuan, et al.
Published: (2025)
by: Gao, Yuanyuan, et al.
Published: (2025)
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
by: Gong, Ziyang, et al.
Published: (2025)
by: Gong, Ziyang, et al.
Published: (2025)
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
ToMiE: Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars
by: Zhan, Yifan, et al.
Published: (2024)
by: Zhan, Yifan, et al.
Published: (2024)
ExGS: Extreme 3D Gaussian Compression with Diffusion Priors
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
Towards Vision-Language Geo-Foundation Model: A Survey
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
by: Niu, Muyao, et al.
Published: (2025)
by: Niu, Muyao, et al.
Published: (2025)
MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks
by: Liu, Yifei, et al.
Published: (2024)
by: Liu, Yifei, et al.
Published: (2024)
A Meta-analysis of College Students' Intention to Use Generative Artificial Intelligence
by: Diao, Yifei, et al.
Published: (2024)
by: Diao, Yifei, et al.
Published: (2024)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
SGA-INTERACT: A 3D Skeleton-based Benchmark for Group Activity Understanding in Modern Basketball Tactic
by: Yang, Yuchen, et al.
Published: (2025)
by: Yang, Yuchen, et al.
Published: (2025)
WEC-DG: Multi-Exposure Wavelet Correction Method Guided by Degradation Description
by: Zhao, Ming, et al.
Published: (2025)
by: Zhao, Ming, et al.
Published: (2025)
Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform
by: Gong, Yuning, et al.
Published: (2025)
by: Gong, Yuning, et al.
Published: (2025)
Adaptive 3D Gaussian Splatting Video Streaming: Visual Saliency-Aware Tiling and Meta-Learning-Based Bitrate Adaptation
by: Gong, Han, et al.
Published: (2025)
by: Gong, Han, et al.
Published: (2025)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
by: Wang, Siting, et al.
Published: (2025)
by: Wang, Siting, et al.
Published: (2025)
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
by: Yang, Jiahui, et al.
Published: (2024)
by: Yang, Jiahui, et al.
Published: (2024)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
by: Chen, Zhanpeng, et al.
Published: (2025)
by: Chen, Zhanpeng, et al.
Published: (2025)
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
by: Kargin, Turhan Can, et al.
Published: (2026)
by: Kargin, Turhan Can, et al.
Published: (2026)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
by: Liu, Zhihang, et al.
Published: (2025)
by: Liu, Zhihang, et al.
Published: (2025)
Hyper-3DG: Text-to-3D Gaussian Generation via Hypergraph
by: Di, Donglin, et al.
Published: (2024)
by: Di, Donglin, et al.
Published: (2024)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
Rethinking Exposure Correction for Spatially Non-uniform Degradation
by: Li, Ao, et al.
Published: (2026)
by: Li, Ao, et al.
Published: (2026)
ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments
by: Gong, Ziyang, et al.
Published: (2026)
by: Gong, Ziyang, et al.
Published: (2026)
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs
by: Li, Mingxiao, et al.
Published: (2025)
by: Li, Mingxiao, et al.
Published: (2025)
SITE: towards Spatial Intelligence Thorough Evaluation
by: Wang, Wenqi, et al.
Published: (2025)
by: Wang, Wenqi, et al.
Published: (2025)
CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction
by: Gao, Yuanyuan, et al.
Published: (2025)
by: Gao, Yuanyuan, et al.
Published: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025)
by: Mayer, Julius, et al.
Published: (2025)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
by: Zhang, Ziang, et al.
Published: (2025)
by: Zhang, Ziang, et al.
Published: (2025)
From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
by: Zhang, Le, et al.
Published: (2026)
by: Zhang, Le, et al.
Published: (2026)
NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Similar Items
-
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026) -
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
by: Gao, Yuanyuan, et al.
Published: (2026) -
PhotoFlow: Agentic 3D Virtual Photography Missions
by: Guo, Jiarui, et al.
Published: (2026) -
R3-Avatar: Record and Retrieve Temporal Codebook for Reconstructing Photorealistic Human Avatars
by: Zhan, Yifan, et al.
Published: (2025) -
KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter
by: Zhan, Yifan, et al.
Published: (2024)