Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Hao, Wang, Ye, Zhang, Wanpeng, Yuan, Haoqi, Feng, Yicheng, Xu, Haiweng, Zheng, Sipeng, Lu, Zongqing
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910032804708352
author Luo, Hao
Wang, Ye
Zhang, Wanpeng
Yuan, Haoqi
Feng, Yicheng
Xu, Haiweng
Zheng, Sipeng
Lu, Zongqing
author_facet Luo, Hao
Wang, Ye
Zhang, Wanpeng
Yuan, Haoqi
Feng, Yicheng
Xu, Haiweng
Zheng, Sipeng
Lu, Zongqing
contents Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small, precisely-labeled datasets and vast in-the-wild footage with unreliable hand tracking labels. We present JALA, a pretraining framework that learns Jointly-Aligned Latent Actions. JALA bypasses full visual dynamic reconstruction, instead learns a predictive action embedding aligned with both inverse dynamics and real actions. This yields a transition-aware, behavior-centric latent space for learning from heterogeneous human data. We scale this approach with UniHand-Mix, a 7.5M video corpus (>2,000 hours) blending laboratory and in-the-wild footage. Experiments demonstrate that JALA generates more realistic hand motions in both controlled and unconstrained scenarios, significantly improving downstream robot manipulation performance in both simulation and real-world tasks. These results indicate that jointly-aligned latent actions offer a scalable pathway for VLA pretraining from human data.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21736
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
Luo, Hao
Wang, Ye
Zhang, Wanpeng
Yuan, Haoqi
Feng, Yicheng
Xu, Haiweng
Zheng, Sipeng
Lu, Zongqing
Robotics
Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small, precisely-labeled datasets and vast in-the-wild footage with unreliable hand tracking labels. We present JALA, a pretraining framework that learns Jointly-Aligned Latent Actions. JALA bypasses full visual dynamic reconstruction, instead learns a predictive action embedding aligned with both inverse dynamics and real actions. This yields a transition-aware, behavior-centric latent space for learning from heterogeneous human data. We scale this approach with UniHand-Mix, a 7.5M video corpus (>2,000 hours) blending laboratory and in-the-wild footage. Experiments demonstrate that JALA generates more realistic hand motions in both controlled and unconstrained scenarios, significantly improving downstream robot manipulation performance in both simulation and real-world tasks. These results indicate that jointly-aligned latent actions offer a scalable pathway for VLA pretraining from human data.
title Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
topic Robotics
url https://arxiv.org/abs/2602.21736