GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Lin, Xie, Bin, Liu, Yingfei, Shi, Hao, Wang, Tiancai, Cao, Jiale
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912536232722432
author Sun, Lin
Xie, Bin
Liu, Yingfei
Shi, Hao
Wang, Tiancai
Cao, Jiale
author_facet Sun, Lin
Xie, Bin
Liu, Yingfei
Shi, Hao
Wang, Tiancai
Cao, Jiale
contents Vision-Language-Action (VLA) models have emerged as a promising approach for enabling robots to follow language instructions and predict corresponding actions. However, current VLA models mainly rely on 2D visual inputs, neglecting the rich geometric information in the 3D physical world, which limits their spatial awareness and adaptability. In this paper, we present GeoVLA, a novel VLA framework that effectively integrates 3D information to advance robotic manipulation. It uses a vision-language model (VLM) to process images and language instructions,extracting fused vision-language embeddings. In parallel, it converts depth maps into point clouds and employs a customized point encoder, called Point Embedding Network, to generate 3D geometric embeddings independently. These produced embeddings are then concatenated and processed by our proposed spatial-aware action expert, called 3D-enhanced Action Expert, which combines information from different sensor modalities to produce precise action sequences. Through extensive experiments in both simulation and real-world environments, GeoVLA demonstrates superior performance and robustness. It achieves state-of-the-art results in the LIBERO and ManiSkill2 simulation benchmarks and shows remarkable robustness in real-world tasks requiring height adaptability, scale awareness and viewpoint invariance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09071
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
Sun, Lin
Xie, Bin
Liu, Yingfei
Shi, Hao
Wang, Tiancai
Cao, Jiale
Robotics
Vision-Language-Action (VLA) models have emerged as a promising approach for enabling robots to follow language instructions and predict corresponding actions. However, current VLA models mainly rely on 2D visual inputs, neglecting the rich geometric information in the 3D physical world, which limits their spatial awareness and adaptability. In this paper, we present GeoVLA, a novel VLA framework that effectively integrates 3D information to advance robotic manipulation. It uses a vision-language model (VLM) to process images and language instructions,extracting fused vision-language embeddings. In parallel, it converts depth maps into point clouds and employs a customized point encoder, called Point Embedding Network, to generate 3D geometric embeddings independently. These produced embeddings are then concatenated and processed by our proposed spatial-aware action expert, called 3D-enhanced Action Expert, which combines information from different sensor modalities to produce precise action sequences. Through extensive experiments in both simulation and real-world environments, GeoVLA demonstrates superior performance and robustness. It achieves state-of-the-art results in the LIBERO and ManiSkill2 simulation benchmarks and shows remarkable robustness in real-world tasks requiring height adaptability, scale awareness and viewpoint invariance.
title GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
topic Robotics
url https://arxiv.org/abs/2508.09071