From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Zhengshen, Li, Hao, Dai, Yalun, Zhu, Zhengbang, Zhou, Lei, Liu, Chenchen, Wang, Dong, Tay, Francis E. H., Chen, Sijin, Liu, Ziwei, Liu, Yuxiao, Li, Xinghang, Zhou, Pan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911499969101824
author Zhang, Zhengshen
Li, Hao
Dai, Yalun
Zhu, Zhengbang
Zhou, Lei
Liu, Chenchen
Wang, Dong
Tay, Francis E. H.
Chen, Sijin
Liu, Ziwei
Liu, Yuxiao
Li, Xinghang
Zhou, Pan
author_facet Zhang, Zhengshen
Li, Hao
Dai, Yalun
Zhu, Zhengbang
Zhou, Lei
Liu, Chenchen
Wang, Dong
Tay, Francis E. H.
Chen, Sijin
Liu, Ziwei
Liu, Yuxiao
Li, Xinghang
Zhou, Pan
contents Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, or inject weak cues that lack geometry and degrade vision-language alignment. In this work, we introduce FALCON (From Spatial to Action), a novel paradigm that injects rich 3D spatial tokens into the action head. FALCON leverages spatial foundation models to deliver strong geometric priors from RGB alone, and includes an Embodied Spatial Model that can optionally fuse depth, or pose for higher fidelity when available, without retraining or architectural changes. To preserve language reasoning, spatial tokens are consumed by a Spatial-Enhanced Action Head rather than being concatenated into the vision-language backbone. These designs enable FALCON to address limitations in spatial representation, modality transferability, and alignment. In comprehensive evaluations across three simulation benchmarks and eleven real-world tasks, our proposed FALCON achieves state-of-the-art performance, consistently surpasses competitive baselines, and remains robust under clutter, spatial-prompt conditioning, and variations in object scale and height.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
Zhang, Zhengshen
Li, Hao
Dai, Yalun
Zhu, Zhengbang
Zhou, Lei
Liu, Chenchen
Wang, Dong
Tay, Francis E. H.
Chen, Sijin
Liu, Ziwei
Liu, Yuxiao
Li, Xinghang
Zhou, Pan
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, or inject weak cues that lack geometry and degrade vision-language alignment. In this work, we introduce FALCON (From Spatial to Action), a novel paradigm that injects rich 3D spatial tokens into the action head. FALCON leverages spatial foundation models to deliver strong geometric priors from RGB alone, and includes an Embodied Spatial Model that can optionally fuse depth, or pose for higher fidelity when available, without retraining or architectural changes. To preserve language reasoning, spatial tokens are consumed by a Spatial-Enhanced Action Head rather than being concatenated into the vision-language backbone. These designs enable FALCON to address limitations in spatial representation, modality transferability, and alignment. In comprehensive evaluations across three simulation benchmarks and eleven real-world tasks, our proposed FALCON achieves state-of-the-art performance, consistently surpasses competitive baselines, and remains robust under clutter, spatial-prompt conditioning, and variations in object scale and height.
title From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.17439