Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Zhongbin, Liu, Jiahe, Li, Yushan, Gao, Wenyu, Yang, Zhen, Li, Chenzhi, Zhang, Xinyue, Jian, Ping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
by: Guo, Zhongbin, et al.
Published: (2026)
by: Guo, Zhongbin, et al.
Published: (2026)
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
by: Deng, Yonghong, et al.
Published: (2026)
by: Deng, Yonghong, et al.
Published: (2026)
Decoupled Competitive Framework for Semi-supervised Medical Image Segmentation
by: Chen, Jiahe, et al.
Published: (2025)
by: Chen, Jiahe, et al.
Published: (2025)
Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning
by: Ran, Xingjian, et al.
Published: (2025)
by: Ran, Xingjian, et al.
Published: (2025)
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
by: Gao, Yuanyuan, et al.
Published: (2026)
by: Gao, Yuanyuan, et al.
Published: (2026)
Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning
by: Jiang, Haoyi, et al.
Published: (2026)
by: Jiang, Haoyi, et al.
Published: (2026)
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
by: Guo, Yanjun, et al.
Published: (2026)
by: Guo, Yanjun, et al.
Published: (2026)
COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
by: Zhang, Zefeng, et al.
Published: (2025)
by: Zhang, Zefeng, et al.
Published: (2025)
Decoupled and Interactive Regression Modeling for High-performance One-stage 3D Object Detection
by: Xiao, Weiping, et al.
Published: (2024)
by: Xiao, Weiping, et al.
Published: (2024)
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
by: Jeon, Yerim, et al.
Published: (2025)
by: Jeon, Yerim, et al.
Published: (2025)
Deep Height Decoupling for Precise Vision-based 3D Occupancy Prediction
by: Wu, Yuan, et al.
Published: (2024)
by: Wu, Yuan, et al.
Published: (2024)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
by: Chen, Jiabin, et al.
Published: (2025)
by: Chen, Jiabin, et al.
Published: (2025)
Video Spatial Reasoning with Object-Centric 3D Rollout
by: Tang, Haoran, et al.
Published: (2025)
by: Tang, Haoran, et al.
Published: (2025)
DIMM: Decoupled Multi-hierarchy Kalman Filter for 3D Object Tracking
by: Zha, Jirong, et al.
Published: (2025)
by: Zha, Jirong, et al.
Published: (2025)
Decoupled Complementary Spectral-Spatial Learning for Background Representation Enhancement in Hyperspectral Anomaly Detection
by: Jin, Wenping, et al.
Published: (2025)
by: Jin, Wenping, et al.
Published: (2025)
Unsupervised Spatial-Temporal Feature Enrichment and Fidelity Preservation Network for Skeleton based Action Recognition
by: Li, Chuankun, et al.
Published: (2024)
by: Li, Chuankun, et al.
Published: (2024)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
by: He, Jixuan, et al.
Published: (2026)
by: He, Jixuan, et al.
Published: (2026)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
by: Zheng, Hongpei, et al.
Published: (2025)
by: Zheng, Hongpei, et al.
Published: (2025)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
by: Yang, Zhen, et al.
Published: (2026)
by: Yang, Zhen, et al.
Published: (2026)
D$^2$-World: An Efficient World Model through Decoupled Dynamic Flow
by: Zhang, Haiming, et al.
Published: (2024)
by: Zhang, Haiming, et al.
Published: (2024)
3D-UIR: 3D Gaussian for Underwater 3D Scene Reconstruction via Physics Based Appearance-Medium Decoupling
by: Yuan, Jieyu, et al.
Published: (2025)
by: Yuan, Jieyu, et al.
Published: (2025)
SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
by: Jian, Siyong, et al.
Published: (2025)
by: Jian, Siyong, et al.
Published: (2025)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
by: Wang, Pan, et al.
Published: (2025)
by: Wang, Pan, et al.
Published: (2025)
Think3D: Thinking with Space for Spatial Reasoning
by: Zhang, Zaibin, et al.
Published: (2026)
by: Zhang, Zaibin, et al.
Published: (2026)
How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
by: Zha, Jirong, et al.
Published: (2025)
by: Zha, Jirong, et al.
Published: (2025)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
by: Li, Siqi, et al.
Published: (2025)
by: Li, Siqi, et al.
Published: (2025)
Memoryless Multimodal Anomaly Detection via Student-Teacher Network and Signed Distance Learning
by: Sun, Zhongbin, et al.
Published: (2024)
by: Sun, Zhongbin, et al.
Published: (2024)
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
by: Yeh, Chun-Hsiao, et al.
Published: (2026)
by: Yeh, Chun-Hsiao, et al.
Published: (2026)
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
by: Liu, Zishan, et al.
Published: (2026)
by: Liu, Zishan, et al.
Published: (2026)
Similar Items
-
LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
by: Guo, Zhongbin, et al.
Published: (2025) -
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
by: Guo, Zhongbin, et al.
Published: (2026) -
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
by: Zhang, Jiahui, et al.
Published: (2025) -
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
by: Guo, Zhongbin, et al.
Published: (2025) -
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)