Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jisoo, Cho, Jungbin, Chu, Sanghyeok, Bal, Ananya, Kim, Jinhyung, Lee, Gunhee, Lee, Sihaeng, Kim, Seung Hwan, Han, Bohyung, Lee, Hyunmin, Jeni, Laszlo A., Kim, Seungryong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910046944755712
author Kim, Jisoo
Cho, Jungbin
Chu, Sanghyeok
Bal, Ananya
Kim, Jinhyung
Lee, Gunhee
Lee, Sihaeng
Kim, Seung Hwan
Han, Bohyung
Lee, Hyunmin
Jeni, Laszlo A.
Kim, Seungryong
author_facet Kim, Jisoo
Cho, Jungbin
Chu, Sanghyeok
Bal, Ananya
Kim, Jinhyung
Lee, Gunhee
Lee, Sihaeng
Kim, Seung Hwan
Han, Bohyung
Lee, Hyunmin
Jeni, Laszlo A.
Kim, Seungryong
contents Humans learn not only how their bodies move, but also how the surrounding world responds to their actions. In contrast, while recent Vision-Language-Action (VLA) models exhibit impressive semantic understanding, they often fail to capture the spatiotemporal dynamics governing physical interaction. In this paper, we introduce Pri4R, a simple yet effective approach that endows VLA models with an implicit understanding of world dynamics by leveraging privileged 4D information during training. Specifically, Pri4R augments VLAs with a lightweight point track head that predicts 3D point tracks. By injecting VLA features into this head to jointly predict future 3D trajectories, the model learns to incorporate evolving scene geometry within its shared representation space, enabling more physically aware context for precise control. Due to its architectural simplicity, Pri4R is compatible with dominant VLA design patterns with minimal changes. During inference, we run the model using the original VLA architecture unchanged; Pri4R adds no extra inputs, outputs, or computational overhead. Across simulation and real-world evaluations, Pri4R significantly improves performance on challenging manipulation tasks, including a +10% gain on LIBERO-Long and a +40% gain on RoboCasa. We further show that 3D point track prediction is an effective supervision target for learning action-world dynamics, and validate our design choices through extensive ablations. Project page: https://jiiiisoo.github.io/Pri4R/
format Preprint
id arxiv_https___arxiv_org_abs_2603_01549
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
Kim, Jisoo
Cho, Jungbin
Chu, Sanghyeok
Bal, Ananya
Kim, Jinhyung
Lee, Gunhee
Lee, Sihaeng
Kim, Seung Hwan
Han, Bohyung
Lee, Hyunmin
Jeni, Laszlo A.
Kim, Seungryong
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Humans learn not only how their bodies move, but also how the surrounding world responds to their actions. In contrast, while recent Vision-Language-Action (VLA) models exhibit impressive semantic understanding, they often fail to capture the spatiotemporal dynamics governing physical interaction. In this paper, we introduce Pri4R, a simple yet effective approach that endows VLA models with an implicit understanding of world dynamics by leveraging privileged 4D information during training. Specifically, Pri4R augments VLAs with a lightweight point track head that predicts 3D point tracks. By injecting VLA features into this head to jointly predict future 3D trajectories, the model learns to incorporate evolving scene geometry within its shared representation space, enabling more physically aware context for precise control. Due to its architectural simplicity, Pri4R is compatible with dominant VLA design patterns with minimal changes. During inference, we run the model using the original VLA architecture unchanged; Pri4R adds no extra inputs, outputs, or computational overhead. Across simulation and real-world evaluations, Pri4R significantly improves performance on challenging manipulation tasks, including a +10% gain on LIBERO-Long and a +40% gain on RoboCasa. We further show that 3D point track prediction is an effective supervision target for learning action-world dynamics, and validate our design choices through extensive ablations. Project page: https://jiiiisoo.github.io/Pri4R/
title Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2603.01549