SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Jian, Zhou, Shijie, Liu, Bangya, Kadambi, Achuta, Fan, Zhiwen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
di: Fan, Zhiwen, et al.
Pubblicazione: (2024)
di: Fan, Zhiwen, et al.
Pubblicazione: (2024)
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
di: Jeon, Changwoo, et al.
Pubblicazione: (2026)
di: Jeon, Changwoo, et al.
Pubblicazione: (2026)
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
di: Zhou, Shijie, et al.
Pubblicazione: (2023)
di: Zhou, Shijie, et al.
Pubblicazione: (2023)
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
di: Zhou, Shijie, et al.
Pubblicazione: (2024)
di: Zhou, Shijie, et al.
Pubblicazione: (2024)
4K4DGen: Panoramic 4D Generation at 4K Resolution
di: Li, Renjie, et al.
Pubblicazione: (2024)
di: Li, Renjie, et al.
Pubblicazione: (2024)
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
di: He, Xuehai, et al.
Pubblicazione: (2025)
di: He, Xuehai, et al.
Pubblicazione: (2025)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
di: Liu, Tianhui, et al.
Pubblicazione: (2026)
di: Liu, Tianhui, et al.
Pubblicazione: (2026)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
InstantRestore: Single-Step Personalized Face Restoration with Shared-Image Attention
di: Zhang, Howard, et al.
Pubblicazione: (2024)
di: Zhang, Howard, et al.
Pubblicazione: (2024)
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
di: Guo, Jiajie, et al.
Pubblicazione: (2025)
di: Guo, Jiajie, et al.
Pubblicazione: (2025)
Solutions to Deepfakes: Can Camera Hardware, Cryptography, and Deep Learning Verify Real Images?
di: Vilesov, Alexander, et al.
Pubblicazione: (2024)
di: Vilesov, Alexander, et al.
Pubblicazione: (2024)
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
di: Zheng, Hongpei, et al.
Pubblicazione: (2025)
di: Zheng, Hongpei, et al.
Pubblicazione: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
di: Zhang, Yiming, et al.
Pubblicazione: (2026)
di: Zhang, Yiming, et al.
Pubblicazione: (2026)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
di: Fan, Zhiwen, et al.
Pubblicazione: (2025)
di: Fan, Zhiwen, et al.
Pubblicazione: (2025)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
di: Ma, Wufei, et al.
Pubblicazione: (2025)
di: Ma, Wufei, et al.
Pubblicazione: (2025)
Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis
di: Li, Dayou, et al.
Pubblicazione: (2026)
di: Li, Dayou, et al.
Pubblicazione: (2026)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
di: Zhang, Wanyue, et al.
Pubblicazione: (2026)
di: Zhang, Wanyue, et al.
Pubblicazione: (2026)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
di: Li, Haoyuan, et al.
Pubblicazione: (2026)
di: Li, Haoyuan, et al.
Pubblicazione: (2026)
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
di: Liu, Zishan, et al.
Pubblicazione: (2026)
di: Liu, Zishan, et al.
Pubblicazione: (2026)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
di: Ma, Xueqi, et al.
Pubblicazione: (2026)
di: Ma, Xueqi, et al.
Pubblicazione: (2026)
Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
di: Cai, Zhongyi, et al.
Pubblicazione: (2025)
di: Cai, Zhongyi, et al.
Pubblicazione: (2025)
TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
di: Ly, Vinh-Thuan, et al.
Pubblicazione: (2025)
di: Ly, Vinh-Thuan, et al.
Pubblicazione: (2025)
Make Geometry Matter for Spatial Reasoning
di: Zhang, Shihua, et al.
Pubblicazione: (2026)
di: Zhang, Shihua, et al.
Pubblicazione: (2026)
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
di: Shi, Jian, et al.
Pubblicazione: (2026)
di: Shi, Jian, et al.
Pubblicazione: (2026)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
Think3D: Thinking with Space for Spatial Reasoning
di: Zhang, Zaibin, et al.
Pubblicazione: (2026)
di: Zhang, Zaibin, et al.
Pubblicazione: (2026)
pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
di: Luo, Zhanpeng, et al.
Pubblicazione: (2026)
di: Luo, Zhanpeng, et al.
Pubblicazione: (2026)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
Vision-Language Memory for Spatial Reasoning
di: Liu, Zuntao, et al.
Pubblicazione: (2025)
di: Liu, Zuntao, et al.
Pubblicazione: (2025)
Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression
di: Guo, Zhongbin, et al.
Pubblicazione: (2025)
di: Guo, Zhongbin, et al.
Pubblicazione: (2025)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
di: Chen, Wei, et al.
Pubblicazione: (2026)
di: Chen, Wei, et al.
Pubblicazione: (2026)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
di: Zhang, Weichen, et al.
Pubblicazione: (2025)
di: Zhang, Weichen, et al.
Pubblicazione: (2025)
Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
di: Bai, Weimin, et al.
Pubblicazione: (2025)
di: Bai, Weimin, et al.
Pubblicazione: (2025)
Privacy-Aware Sharing of Raw Spatial Sensor Data for Cooperative Perception
di: Liu, Bangya, et al.
Pubblicazione: (2025)
di: Liu, Bangya, et al.
Pubblicazione: (2025)
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
di: Liu, Bangya, et al.
Pubblicazione: (2024)
di: Liu, Bangya, et al.
Pubblicazione: (2024)
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
di: Ma, Chenyang, et al.
Pubblicazione: (2024)
di: Ma, Chenyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
di: Fan, Zhiwen, et al.
Pubblicazione: (2024) -
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
di: Jeon, Changwoo, et al.
Pubblicazione: (2026) -
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
di: Zhou, Shijie, et al.
Pubblicazione: (2025) -
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
di: Zhou, Shijie, et al.
Pubblicazione: (2023) -
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
di: Zhou, Shijie, et al.
Pubblicazione: (2024)