VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Jingwen, Zhang, Wenyao, Qi, Zekun, Ren, Shaojie, Liu, Zezhi, Zhu, Hanxin, Sun, Guangzhong, Jin, Xin, Chen, Zhibo
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912904129806336
author Sun, Jingwen
Zhang, Wenyao
Qi, Zekun
Ren, Shaojie
Liu, Zezhi
Zhu, Hanxin
Sun, Guangzhong
Jin, Xin
Chen, Zhibo
author_facet Sun, Jingwen
Zhang, Wenyao
Qi, Zekun
Ren, Shaojie
Liu, Zezhi
Zhu, Hanxin
Sun, Guangzhong
Jin, Xin
Chen, Zhibo
contents Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions, making them vulnerable to appearance bias, nuisance motion, and information leakage. We introduce VLA-JEPA, a JEPA-style pretraining framework that sidesteps these pitfalls by design. The key idea is leakage-free state prediction: a target encoder produces latent representations from future frames, while the student pathway sees only the current observation -- future information is used solely as supervision targets, never as input. By predicting in latent space rather than pixel space, VLA-JEPA learns dynamics abstractions that are robust to camera motion and irrelevant background changes. This yields a simple two-stage recipe -- JEPA pretraining followed by action-head fine-tuning -- without the multi-stage complexity of prior latent-action pipelines. Experiments on LIBERO, LIBERO-Plus, SimplerEnv and real-world manipulation tasks show that VLA-JEPA achieves consistent gains in generalization and robustness over existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10098
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
Sun, Jingwen
Zhang, Wenyao
Qi, Zekun
Ren, Shaojie
Liu, Zezhi
Zhu, Hanxin
Sun, Guangzhong
Jin, Xin
Chen, Zhibo
Robotics
Computer Vision and Pattern Recognition
Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions, making them vulnerable to appearance bias, nuisance motion, and information leakage. We introduce VLA-JEPA, a JEPA-style pretraining framework that sidesteps these pitfalls by design. The key idea is leakage-free state prediction: a target encoder produces latent representations from future frames, while the student pathway sees only the current observation -- future information is used solely as supervision targets, never as input. By predicting in latent space rather than pixel space, VLA-JEPA learns dynamics abstractions that are robust to camera motion and irrelevant background changes. This yields a simple two-stage recipe -- JEPA pretraining followed by action-head fine-tuning -- without the multi-stage complexity of prior latent-action pipelines. Experiments on LIBERO, LIBERO-Plus, SimplerEnv and real-world manipulation tasks show that VLA-JEPA achieves consistent gains in generalization and robustness over existing methods.
title VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.10098