From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhengshen, Li, Hao, Dai, Yalun, Zhu, Zhengbang, Zhou, Lei, Liu, Chenchen, Wang, Dong, Tay, Francis E. H., Chen, Sijin, Liu, Ziwei, Liu, Yuxiao, Li, Xinghang, Zhou, Pan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
World Guidance: World Modeling in Condition Space for Action Generation
by: Su, Yue, et al.
Published: (2026)
by: Su, Yue, et al.
Published: (2026)
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
by: Peng, Haosong, et al.
Published: (2026)
by: Peng, Haosong, et al.
Published: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025)
by: Peng, Haosong, et al.
Published: (2025)
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
by: Tu, Ruisen, et al.
Published: (2026)
by: Tu, Ruisen, et al.
Published: (2026)
DexGrasp-Diffusion: Diffusion-based Unified Functional Grasp Synthesis Method for Multi-Dexterous Robotic Hands
by: Zhang, Zhengshen, et al.
Published: (2024)
by: Zhang, Zhengshen, et al.
Published: (2024)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
What Matters in Building Vision-Language-Action Models for Generalist Robots
by: Li, Xinghang, et al.
Published: (2024)
by: Li, Xinghang, et al.
Published: (2024)
You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects
by: Zhou, Lei, et al.
Published: (2024)
by: Zhou, Lei, et al.
Published: (2024)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
by: Guo, Jun, et al.
Published: (2026)
by: Guo, Jun, et al.
Published: (2026)
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
by: Lin, Tao, et al.
Published: (2025)
by: Lin, Tao, et al.
Published: (2025)
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
by: Pan, Xu, et al.
Published: (2026)
by: Pan, Xu, et al.
Published: (2026)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
by: Li, Pengteng, et al.
Published: (2026)
by: Li, Pengteng, et al.
Published: (2026)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prior
by: Niu, Haochen, et al.
Published: (2026)
by: Niu, Haochen, et al.
Published: (2026)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
Advancing Vision Transformer with Enhanced Spatial Priors
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
Asymptotic Symmetries of the Holst Action at Spatial Infinity: Including Supertranslations
by: Bakhoda, Sepideh, et al.
Published: (2026)
by: Bakhoda, Sepideh, et al.
Published: (2026)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
by: Zhu, Juan, et al.
Published: (2026)
by: Zhu, Juan, et al.
Published: (2026)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
by: Wang, Hanzhen, et al.
Published: (2025)
by: Wang, Hanzhen, et al.
Published: (2025)
InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
by: Yuan, Tianyuan, et al.
Published: (2025)
by: Yuan, Tianyuan, et al.
Published: (2025)
Action-Prior Denoising for Smooth Real-Time Chunking
by: Liu, Dongyang, et al.
Published: (2026)
by: Liu, Dongyang, et al.
Published: (2026)
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
by: Babey, Nicholas, et al.
Published: (2025)
by: Babey, Nicholas, et al.
Published: (2025)
A New Spatial Coupling Model for Foundation Uplift and Its Application
by: Xuyang Duan, et al.
Published: (2026)
by: Xuyang Duan, et al.
Published: (2026)
Survey of Vision-Language-Action Models for Embodied Manipulation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
The Generalized Harmonic Mean for p‐Values: Combining Dependent and Independent Tests
by: Zhengbang Li, et al.
Published: (2026)
by: Zhengbang Li, et al.
Published: (2026)
When Spatial meets Temporal in Action Recognition
by: Chen, Huilin, et al.
Published: (2024)
by: Chen, Huilin, et al.
Published: (2024)
Dexbotic: Open-Source Vision-Language-Action Toolbox
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
by: Dai, Yuntao, et al.
Published: (2025)
by: Dai, Yuntao, et al.
Published: (2025)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation
by: Tian, Yuxuan, et al.
Published: (2026)
by: Tian, Yuxuan, et al.
Published: (2026)
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
by: Luo, Yulin, et al.
Published: (2026)
by: Luo, Yulin, et al.
Published: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
by: Li, Qixiu, et al.
Published: (2024)
by: Li, Qixiu, et al.
Published: (2024)
ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
by: Li, Yuhang, et al.
Published: (2025)
by: Li, Yuhang, et al.
Published: (2025)
Similar Items
-
World Guidance: World Modeling in Condition Space for Action Generation
by: Su, Yue, et al.
Published: (2026) -
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
by: Peng, Haosong, et al.
Published: (2026) -
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026) -
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025) -
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
by: Tu, Ruisen, et al.
Published: (2026)