SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Tianhui, Feng, Jie, Zheng, Zhiheng, Wang, Shengyuan, Guo, Yiming, Xi, Yanxin, Fan, Hangyu, Li, Yong, Hui, Pan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
von: Feng, Jie, et al.
Veröffentlicht: (2025)
von: Feng, Jie, et al.
Veröffentlicht: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale
von: Wang, Shengyuan, et al.
Veröffentlicht: (2025)
von: Wang, Shengyuan, et al.
Veröffentlicht: (2025)
CityRiSE: Reasoning Urban Socio-Economic Status in Vision-Language Models via Reinforcement Learning
von: Liu, Tianhui, et al.
Veröffentlicht: (2025)
von: Liu, Tianhui, et al.
Veröffentlicht: (2025)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
von: Wang, Shengyuan, et al.
Veröffentlicht: (2025)
von: Wang, Shengyuan, et al.
Veröffentlicht: (2025)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
von: Bae, Kyungho, et al.
Veröffentlicht: (2025)
von: Bae, Kyungho, et al.
Veröffentlicht: (2025)
CityGPT: Empowering Urban Spatial Cognition of Large Language Models
von: Feng, Jie, et al.
Veröffentlicht: (2024)
von: Feng, Jie, et al.
Veröffentlicht: (2024)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science
von: Feng, Jie, et al.
Veröffentlicht: (2025)
von: Feng, Jie, et al.
Veröffentlicht: (2025)
SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
DCAU-Net: Differential Cross Attention and Channel-Spatial Feature Fusion for Medical Image Segmentation
von: Li, Yanxin, et al.
Veröffentlicht: (2026)
von: Li, Yanxin, et al.
Veröffentlicht: (2026)
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
von: Tan, Kaizhen, et al.
Veröffentlicht: (2026)
von: Tan, Kaizhen, et al.
Veröffentlicht: (2026)
Spatial As Deep: Spatial CNN for Traffic Scene Understanding
von: Pan, Xingang, et al.
Veröffentlicht: (2017)
von: Pan, Xingang, et al.
Veröffentlicht: (2017)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
von: Ma, Xueqi, et al.
Veröffentlicht: (2026)
von: Ma, Xueqi, et al.
Veröffentlicht: (2026)
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
von: Jeon, Yerim, et al.
Veröffentlicht: (2025)
von: Jeon, Yerim, et al.
Veröffentlicht: (2025)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026)
GFSR: Geometric Fidelity and Spatial Refinement for Reliable Lane Detection
von: Wang, Tiancheng, et al.
Veröffentlicht: (2026)
von: Wang, Tiancheng, et al.
Veröffentlicht: (2026)
Geometrically-Constrained Agent for Spatial Reasoning
von: Chen, Zeren, et al.
Veröffentlicht: (2025)
von: Chen, Zeren, et al.
Veröffentlicht: (2025)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
von: Fang, Jiading
Veröffentlicht: (2025)
von: Fang, Jiading
Veröffentlicht: (2025)
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
von: Zheng, Hongpei, et al.
Veröffentlicht: (2025)
von: Zheng, Hongpei, et al.
Veröffentlicht: (2025)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
von: Zhang, Yimeng, et al.
Veröffentlicht: (2025)
von: Zhang, Yimeng, et al.
Veröffentlicht: (2025)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
von: Zhang, Wanyue, et al.
Veröffentlicht: (2026)
von: Zhang, Wanyue, et al.
Veröffentlicht: (2026)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
von: Huang, Yihong, et al.
Veröffentlicht: (2026)
von: Huang, Yihong, et al.
Veröffentlicht: (2026)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
von: Shiri, Fatemeh, et al.
Veröffentlicht: (2024)
von: Shiri, Fatemeh, et al.
Veröffentlicht: (2024)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames
von: Chen, Chao, et al.
Veröffentlicht: (2023)
von: Chen, Chao, et al.
Veröffentlicht: (2023)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models
von: Bai, Maggie, et al.
Veröffentlicht: (2025)
von: Bai, Maggie, et al.
Veröffentlicht: (2025)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
von: Rizvi, Md Imbesat Hassan, et al.
Veröffentlicht: (2024)
von: Rizvi, Md Imbesat Hassan, et al.
Veröffentlicht: (2024)
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
von: Li, Keliang, et al.
Veröffentlicht: (2025)
von: Li, Keliang, et al.
Veröffentlicht: (2025)
MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
von: Hao, Jinkun, et al.
Veröffentlicht: (2025)
von: Hao, Jinkun, et al.
Veröffentlicht: (2025)
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
von: Ma, Chenyang, et al.
Veröffentlicht: (2024)
von: Ma, Chenyang, et al.
Veröffentlicht: (2024)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
von: Feng, Jie, et al.
Veröffentlicht: (2025) -
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024) -
RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale
von: Wang, Shengyuan, et al.
Veröffentlicht: (2025) -
CityRiSE: Reasoning Urban Socio-Economic Status in Vision-Language Models via Reinforcement Learning
von: Liu, Tianhui, et al.
Veröffentlicht: (2025) -
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
von: Zhang, Jian, et al.
Veröffentlicht: (2026)