Saved in:
| Main Authors: | Zhang, Jian, Zhou, Shijie, Liu, Bangya, Kadambi, Achuta, Fan, Zhiwen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.27437 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
by: Jeon, Changwoo, et al.
Published: (2026)
by: Jeon, Changwoo, et al.
Published: (2026)
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
by: Zhou, Shijie, et al.
Published: (2023)
by: Zhou, Shijie, et al.
Published: (2023)
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
4K4DGen: Panoramic 4D Generation at 4K Resolution
by: Li, Renjie, et al.
Published: (2024)
by: Li, Renjie, et al.
Published: (2024)
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
by: He, Xuehai, et al.
Published: (2025)
by: He, Xuehai, et al.
Published: (2025)
InstantRestore: Single-Step Personalized Face Restoration with Shared-Image Attention
by: Zhang, Howard, et al.
Published: (2024)
by: Zhang, Howard, et al.
Published: (2024)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
Solutions to Deepfakes: Can Camera Hardware, Cryptography, and Deep Learning Verify Real Images?
by: Vilesov, Alexander, et al.
Published: (2024)
by: Vilesov, Alexander, et al.
Published: (2024)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
by: Guo, Jiajie, et al.
Published: (2025)
by: Guo, Jiajie, et al.
Published: (2025)
Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis
by: Li, Dayou, et al.
Published: (2026)
by: Li, Dayou, et al.
Published: (2026)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
by: Fan, Zhiwen, et al.
Published: (2025)
by: Fan, Zhiwen, et al.
Published: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
by: Zheng, Hongpei, et al.
Published: (2025)
by: Zheng, Hongpei, et al.
Published: (2025)
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
by: Liu, Zishan, et al.
Published: (2026)
by: Liu, Zishan, et al.
Published: (2026)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
Privacy-Aware Sharing of Raw Spatial Sensor Data for Cooperative Perception
by: Liu, Bangya, et al.
Published: (2025)
by: Liu, Bangya, et al.
Published: (2025)
SparseGS: Real-Time 360° Sparse View Synthesis using Gaussian Splatting
by: Xiong, Haolin, et al.
Published: (2023)
by: Xiong, Haolin, et al.
Published: (2023)
Make Geometry Matter for Spatial Reasoning
by: Zhang, Shihua, et al.
Published: (2026)
by: Zhang, Shihua, et al.
Published: (2026)
Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
by: Cai, Zhongyi, et al.
Published: (2025)
by: Cai, Zhongyi, et al.
Published: (2025)
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
by: Liu, Bangya, et al.
Published: (2024)
by: Liu, Bangya, et al.
Published: (2024)
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
by: Shi, Jian, et al.
Published: (2026)
by: Shi, Jian, et al.
Published: (2026)
TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
by: Ly, Vinh-Thuan, et al.
Published: (2025)
by: Ly, Vinh-Thuan, et al.
Published: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
by: Pan, Zhenyu, et al.
Published: (2025)
by: Pan, Zhenyu, et al.
Published: (2025)
DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
by: Wu, Changti, et al.
Published: (2025)
by: Wu, Changti, et al.
Published: (2025)
Think3D: Thinking with Space for Spatial Reasoning
by: Zhang, Zaibin, et al.
Published: (2026)
by: Zhang, Zaibin, et al.
Published: (2026)
pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
by: Luo, Zhanpeng, et al.
Published: (2026)
by: Luo, Zhanpeng, et al.
Published: (2026)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Vision-Language Memory for Spatial Reasoning
by: Liu, Zuntao, et al.
Published: (2025)
by: Liu, Zuntao, et al.
Published: (2025)
Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
by: Chen, Pingyi, et al.
Published: (2025)
by: Chen, Pingyi, et al.
Published: (2025)
Similar Items
-
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
by: Fan, Zhiwen, et al.
Published: (2024) -
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
by: Jeon, Changwoo, et al.
Published: (2026) -
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
by: Zhou, Shijie, et al.
Published: (2025) -
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
by: Zhou, Shijie, et al.
Published: (2023) -
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024)