Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910032804708352 |
|---|---|
| author | Luo, Hao Wang, Ye Zhang, Wanpeng Yuan, Haoqi Feng, Yicheng Xu, Haiweng Zheng, Sipeng Lu, Zongqing |
| author_facet | Luo, Hao Wang, Ye Zhang, Wanpeng Yuan, Haoqi Feng, Yicheng Xu, Haiweng Zheng, Sipeng Lu, Zongqing |
| contents | Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small, precisely-labeled datasets and vast in-the-wild footage with unreliable hand tracking labels. We present JALA, a pretraining framework that learns Jointly-Aligned Latent Actions. JALA bypasses full visual dynamic reconstruction, instead learns a predictive action embedding aligned with both inverse dynamics and real actions. This yields a transition-aware, behavior-centric latent space for learning from heterogeneous human data. We scale this approach with UniHand-Mix, a 7.5M video corpus (>2,000 hours) blending laboratory and in-the-wild footage. Experiments demonstrate that JALA generates more realistic hand motions in both controlled and unconstrained scenarios, significantly improving downstream robot manipulation performance in both simulation and real-world tasks. These results indicate that jointly-aligned latent actions offer a scalable pathway for VLA pretraining from human data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_21736 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild Luo, Hao Wang, Ye Zhang, Wanpeng Yuan, Haoqi Feng, Yicheng Xu, Haiweng Zheng, Sipeng Lu, Zongqing Robotics Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small, precisely-labeled datasets and vast in-the-wild footage with unreliable hand tracking labels. We present JALA, a pretraining framework that learns Jointly-Aligned Latent Actions. JALA bypasses full visual dynamic reconstruction, instead learns a predictive action embedding aligned with both inverse dynamics and real actions. This yields a transition-aware, behavior-centric latent space for learning from heterogeneous human data. We scale this approach with UniHand-Mix, a 7.5M video corpus (>2,000 hours) blending laboratory and in-the-wild footage. Experiments demonstrate that JALA generates more realistic hand motions in both controlled and unconstrained scenarios, significantly improving downstream robot manipulation performance in both simulation and real-world tasks. These results indicate that jointly-aligned latent actions offer a scalable pathway for VLA pretraining from human data. |
| title | Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild |
| topic | Robotics |
| url | https://arxiv.org/abs/2602.21736 |