Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Tao, Li, Gen, Zhong, Yilei, Zou, Yanwen, Du, Yuxin, Liu, Jiting, Gu, Encheng, Zhao, Bo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912725752348672
author Lin, Tao
Li, Gen
Zhong, Yilei
Zou, Yanwen
Du, Yuxin
Liu, Jiting
Gu, Encheng
Zhao, Bo
author_facet Lin, Tao
Li, Gen
Zhong, Yilei
Zou, Yanwen
Du, Yuxin
Liu, Jiting
Gu, Encheng
Zhao, Bo
contents Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models (VLMs), which excel at semantic understanding due to large-scale image and text pretraining. However, existing VLMs typically lack precise spatial understanding capabilities, as they are primarily tuned on 2D image-text pairs without 3D supervision. To address this limitation, recent approaches have incorporated explicit 3D inputs such as point clouds or depth maps, but this necessitates additional depth sensors or pre-trained depth estimation models, which may yield defective results. In contrast, our work introduces a plug-and-play module that implicitly incorporates 3D geometry features into VLA models by leveraging an off-the-shelf visual geometry foundation model. This integration provides the model with depth-aware visual representations, improving its ability to understand the geometric structure of the scene and the spatial relationships among objects from RGB images alone. We evaluate our method on a set of spatially challenging tasks in both simulation and the real world. Extensive evaluations show that our method significantly improves the performance of state-of-the-art VLA models across diverse scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00416
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
Lin, Tao
Li, Gen
Zhong, Yilei
Zou, Yanwen
Du, Yuxin
Liu, Jiting
Gu, Encheng
Zhao, Bo
Robotics
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models (VLMs), which excel at semantic understanding due to large-scale image and text pretraining. However, existing VLMs typically lack precise spatial understanding capabilities, as they are primarily tuned on 2D image-text pairs without 3D supervision. To address this limitation, recent approaches have incorporated explicit 3D inputs such as point clouds or depth maps, but this necessitates additional depth sensors or pre-trained depth estimation models, which may yield defective results. In contrast, our work introduces a plug-and-play module that implicitly incorporates 3D geometry features into VLA models by leveraging an off-the-shelf visual geometry foundation model. This integration provides the model with depth-aware visual representations, improving its ability to understand the geometric structure of the scene and the spatial relationships among objects from RGB images alone. We evaluate our method on a set of spatially challenging tasks in both simulation and the real world. Extensive evaluations show that our method significantly improves the performance of state-of-the-art VLA models across diverse scenarios.
title Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.00416