LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917275323334656 |
|---|---|
| author | Zha, Lihan Hancock, Asher J. Zhang, Mingtong Yin, Tenny Huang, Yixuan Shah, Dhruv Ren, Allen Z. Majumdar, Anirudha |
| author_facet | Zha, Lihan Hancock, Asher J. Zhang, Mingtong Yin, Tenny Huang, Yixuan Shah, Dhruv Ren, Allen Z. Majumdar, Anirudha |
| contents | A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_10556 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer Zha, Lihan Hancock, Asher J. Zhang, Mingtong Yin, Tenny Huang, Yixuan Shah, Dhruv Ren, Allen Z. Majumdar, Anirudha Robotics Artificial Intelligence A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training. |
| title | LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer |
| topic | Robotics Artificial Intelligence |
| url | https://arxiv.org/abs/2602.10556 |