LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zha, Lihan, Hancock, Asher J., Zhang, Mingtong, Yin, Tenny, Huang, Yixuan, Shah, Dhruv, Ren, Allen Z., Majumdar, Anirudha
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917275323334656
author Zha, Lihan
Hancock, Asher J.
Zhang, Mingtong
Yin, Tenny
Huang, Yixuan
Shah, Dhruv
Ren, Allen Z.
Majumdar, Anirudha
author_facet Zha, Lihan
Hancock, Asher J.
Zhang, Mingtong
Yin, Tenny
Huang, Yixuan
Shah, Dhruv
Ren, Allen Z.
Majumdar, Anirudha
contents A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10556
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
Zha, Lihan
Hancock, Asher J.
Zhang, Mingtong
Yin, Tenny
Huang, Yixuan
Shah, Dhruv
Ren, Allen Z.
Majumdar, Anirudha
Robotics
Artificial Intelligence
A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training.
title LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2602.10556