VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912904129806336 |
|---|---|
| author | Sun, Jingwen Zhang, Wenyao Qi, Zekun Ren, Shaojie Liu, Zezhi Zhu, Hanxin Sun, Guangzhong Jin, Xin Chen, Zhibo |
| author_facet | Sun, Jingwen Zhang, Wenyao Qi, Zekun Ren, Shaojie Liu, Zezhi Zhu, Hanxin Sun, Guangzhong Jin, Xin Chen, Zhibo |
| contents | Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions, making them vulnerable to appearance bias, nuisance motion, and information leakage. We introduce VLA-JEPA, a JEPA-style pretraining framework that sidesteps these pitfalls by design. The key idea is leakage-free state prediction: a target encoder produces latent representations from future frames, while the student pathway sees only the current observation -- future information is used solely as supervision targets, never as input. By predicting in latent space rather than pixel space, VLA-JEPA learns dynamics abstractions that are robust to camera motion and irrelevant background changes. This yields a simple two-stage recipe -- JEPA pretraining followed by action-head fine-tuning -- without the multi-stage complexity of prior latent-action pipelines. Experiments on LIBERO, LIBERO-Plus, SimplerEnv and real-world manipulation tasks show that VLA-JEPA achieves consistent gains in generalization and robustness over existing methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_10098 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model Sun, Jingwen Zhang, Wenyao Qi, Zekun Ren, Shaojie Liu, Zezhi Zhu, Hanxin Sun, Guangzhong Jin, Xin Chen, Zhibo Robotics Computer Vision and Pattern Recognition Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions, making them vulnerable to appearance bias, nuisance motion, and information leakage. We introduce VLA-JEPA, a JEPA-style pretraining framework that sidesteps these pitfalls by design. The key idea is leakage-free state prediction: a target encoder produces latent representations from future frames, while the student pathway sees only the current observation -- future information is used solely as supervision targets, never as input. By predicting in latent space rather than pixel space, VLA-JEPA learns dynamics abstractions that are robust to camera motion and irrelevant background changes. This yields a simple two-stage recipe -- JEPA pretraining followed by action-head fine-tuning -- without the multi-stage complexity of prior latent-action pipelines. Experiments on LIBERO, LIBERO-Plus, SimplerEnv and real-world manipulation tasks show that VLA-JEPA achieves consistent gains in generalization and robustness over existing methods. |
| title | VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.10098 |