Pre-training Auto-regressive Robotic Models with 4D Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niu, Dantong, Sharma, Yuvan, Xue, Haoru, Biamby, Giscard, Zhang, Junyi, Ji, Ziteng, Darrell, Trevor, Herzig, Roei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913843519684608
author Niu, Dantong
Sharma, Yuvan
Xue, Haoru
Biamby, Giscard
Zhang, Junyi
Ji, Ziteng
Darrell, Trevor
Herzig, Roei
author_facet Niu, Dantong
Sharma, Yuvan
Xue, Haoru
Biamby, Giscard
Zhang, Junyi
Ji, Ziteng
Darrell, Trevor
Herzig, Roei
contents Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Yet, efforts in robotics have struggled to achieve similar success, limited by either the need for costly robotic annotations or the lack of representations that effectively model the physical world. In this paper, we introduce ARM4R, an Auto-regressive Robotic Model that leverages low-level 4D Representations learned from human video data to yield a better pre-trained robotic model. Specifically, we focus on utilizing 3D point tracking representations from videos derived by lifting 2D representations into 3D space via monocular depth estimation across time. These 4D representations maintain a shared geometric structure between the points and robot state representations up to a linear transformation, enabling efficient transfer learning from human video data to low-level robotic control. Our experiments show that ARM4R can transfer efficiently from human video data to robotics and consistently improves performance on tasks across various robot environments and configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pre-training Auto-regressive Robotic Models with 4D Representations
Niu, Dantong
Sharma, Yuvan
Xue, Haoru
Biamby, Giscard
Zhang, Junyi
Ji, Ziteng
Darrell, Trevor
Herzig, Roei
Robotics
Artificial Intelligence
Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Yet, efforts in robotics have struggled to achieve similar success, limited by either the need for costly robotic annotations or the lack of representations that effectively model the physical world. In this paper, we introduce ARM4R, an Auto-regressive Robotic Model that leverages low-level 4D Representations learned from human video data to yield a better pre-trained robotic model. Specifically, we focus on utilizing 3D point tracking representations from videos derived by lifting 2D representations into 3D space via monocular depth estimation across time. These 4D representations maintain a shared geometric structure between the points and robot state representations up to a linear transformation, enabling efficient transfer learning from human video data to low-level robotic control. Our experiments show that ARM4R can transfer efficiently from human video data to robotics and consistently improves performance on tasks across various robot environments and configurations.
title Pre-training Auto-regressive Robotic Models with 4D Representations
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2502.13142