Physical Autoregressive Model for Robotic Manipulation without Action Pretraining

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Song, Zijian, Qin, Sihan, Chen, Tianshui, Lin, Liang, Wang, Guangrun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912575372918784
author Song, Zijian
Qin, Sihan
Chen, Tianshui
Lin, Liang
Wang, Guangrun
author_facet Song, Zijian
Qin, Sihan
Chen, Tianshui
Lin, Liang
Wang, Guangrun
contents The scarcity of manipulation data has motivated the use of pretrained large models from other modalities in robotics. In this work, we build upon autoregressive video generation models to propose a Physical Autoregressive Model (PAR), where physical tokens combine frames and actions to represent the joint evolution of the robot and its environment. PAR leverages the world knowledge embedded in video pretraining to understand physical dynamics without requiring action pretraining, enabling accurate video prediction and consistent action trajectories. It also adopts a DiT-based de-tokenizer to model frames and actions as continuous tokens, mitigating quantization errors and facilitating mutual enhancement. Furthermore, we incorporate a causal mask with inverse kinematics, parallel training, and the KV-cache mechanism to further improve performance and efficiency. Experiments on the ManiSkill benchmark show that PAR achieves a 100\% success rate on the PushCube task, matches the performance of action-pretrained baselines on other tasks, and accurately predicts future videos with tightly aligned action trajectories. These findings underscore a promising direction for robotic manipulation by transferring world knowledge from autoregressive video pretraining. The project page is here: https://hcplab-sysu.github.io/PhysicalAutoregressiveModel/
format Preprint
id arxiv_https___arxiv_org_abs_2508_09822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
Song, Zijian
Qin, Sihan
Chen, Tianshui
Lin, Liang
Wang, Guangrun
Computer Vision and Pattern Recognition
The scarcity of manipulation data has motivated the use of pretrained large models from other modalities in robotics. In this work, we build upon autoregressive video generation models to propose a Physical Autoregressive Model (PAR), where physical tokens combine frames and actions to represent the joint evolution of the robot and its environment. PAR leverages the world knowledge embedded in video pretraining to understand physical dynamics without requiring action pretraining, enabling accurate video prediction and consistent action trajectories. It also adopts a DiT-based de-tokenizer to model frames and actions as continuous tokens, mitigating quantization errors and facilitating mutual enhancement. Furthermore, we incorporate a causal mask with inverse kinematics, parallel training, and the KV-cache mechanism to further improve performance and efficiency. Experiments on the ManiSkill benchmark show that PAR achieves a 100\% success rate on the PushCube task, matches the performance of action-pretrained baselines on other tasks, and accurately predicts future videos with tightly aligned action trajectories. These findings underscore a promising direction for robotic manipulation by transferring world knowledge from autoregressive video pretraining. The project page is here: https://hcplab-sysu.github.io/PhysicalAutoregressiveModel/
title Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.09822