Toward Embodiment Equivariant Vision-Language-Action Policy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Anzhe, Yang, Yifei, Zhu, Zhenjie, Xu, Kechun, Zhou, Zhongxiang, Xiong, Rong, Wang, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914044606152704
author Chen, Anzhe
Yang, Yifei
Zhu, Zhenjie
Xu, Kechun
Zhou, Zhongxiang
Xiong, Rong
Wang, Yue
author_facet Chen, Anzhe
Yang, Yifei
Zhu, Zhenjie
Xu, Kechun
Zhou, Zhongxiang
Xiong, Rong
Wang, Yue
contents Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches emphasize model size, dataset scale and diversity while paying less attention to the design of action spaces. This leads to the configuration generalization problem, which requires costly adaptation. We address this challenge by formulating cross-embodiment pre-training as designing policies equivariant to embodiment configuration transformations. Building on this principle, we propose a framework that (i) establishes a embodiment equivariance theory for action space and policy design, (ii) introduces an action decoder that enforces configuration equivariance, and (iii) incorporates a geometry-aware network architecture to enhance embodiment-agnostic spatial reasoning. Extensive experiments in both simulation and real-world settings demonstrate that our approach improves pre-training effectiveness and enables efficient fine-tuning on novel robot embodiments. Our code is available at https://github.com/hhcaz/e2vla
format Preprint
id arxiv_https___arxiv_org_abs_2509_14630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward Embodiment Equivariant Vision-Language-Action Policy
Chen, Anzhe
Yang, Yifei
Zhu, Zhenjie
Xu, Kechun
Zhou, Zhongxiang
Xiong, Rong
Wang, Yue
Robotics
Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches emphasize model size, dataset scale and diversity while paying less attention to the design of action spaces. This leads to the configuration generalization problem, which requires costly adaptation. We address this challenge by formulating cross-embodiment pre-training as designing policies equivariant to embodiment configuration transformations. Building on this principle, we propose a framework that (i) establishes a embodiment equivariance theory for action space and policy design, (ii) introduces an action decoder that enforces configuration equivariance, and (iii) incorporates a geometry-aware network architecture to enhance embodiment-agnostic spatial reasoning. Extensive experiments in both simulation and real-world settings demonstrate that our approach improves pre-training effectiveness and enables efficient fine-tuning on novel robot embodiments. Our code is available at https://github.com/hhcaz/e2vla
title Toward Embodiment Equivariant Vision-Language-Action Policy
topic Robotics
url https://arxiv.org/abs/2509.14630