Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Qiwei, Cai, Boyang, Lai, Minghao, Zhuang, Sitong, Lin, Tao, Qin, Yan, Ye, Yixuan, Liang, Jiaming, Xu, Renjing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915848126464000
author Liang, Qiwei
Cai, Boyang
Lai, Minghao
Zhuang, Sitong
Lin, Tao
Qin, Yan
Ye, Yixuan
Liang, Jiaming
Xu, Renjing
author_facet Liang, Qiwei
Cai, Boyang
Lai, Minghao
Zhuang, Sitong
Lin, Tao
Qin, Yan
Ye, Yixuan
Liang, Jiaming
Xu, Renjing
contents Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underperform on robotic manipulation. We attribute this gap to two factors: the lack of state-action-state dynamics modeling and the unnecessary redundancy of explicit geometric reconstruction. We introduce AFRO, a self-supervised framework that learns dynamics-aware 3D representations without action or reconstruction supervision. AFRO casts state prediction as a generative diffusion process and jointly models forward and inverse dynamics in a shared latent space to capture causal transition structure. To prevent feature leakage in action learning, we employ feature differencing and inverse-consistency supervision, improving the quality and stability of visual features. When combined with Diffusion Policy, AFRO substantially increases manipulation success rates across 16 simulated and 4 real-world tasks, outperforming existing pre-training approaches. The framework also scales favorably with data volume and task complexity. Qualitative visualizations indicate that AFRO learns semantically rich, discriminative features, offering an effective pre-training solution for 3D representation learning in robotics. Project page: https://kolakivy.github.io/AFRO/
format Preprint
id arxiv_https___arxiv_org_abs_2512_00074
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning
Liang, Qiwei
Cai, Boyang
Lai, Minghao
Zhuang, Sitong
Lin, Tao
Qin, Yan
Ye, Yixuan
Liang, Jiaming
Xu, Renjing
Robotics
Computer Vision and Pattern Recognition
Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underperform on robotic manipulation. We attribute this gap to two factors: the lack of state-action-state dynamics modeling and the unnecessary redundancy of explicit geometric reconstruction. We introduce AFRO, a self-supervised framework that learns dynamics-aware 3D representations without action or reconstruction supervision. AFRO casts state prediction as a generative diffusion process and jointly models forward and inverse dynamics in a shared latent space to capture causal transition structure. To prevent feature leakage in action learning, we employ feature differencing and inverse-consistency supervision, improving the quality and stability of visual features. When combined with Diffusion Policy, AFRO substantially increases manipulation success rates across 16 simulated and 4 real-world tasks, outperforming existing pre-training approaches. The framework also scales favorably with data volume and task complexity. Qualitative visualizations indicate that AFRO learns semantically rich, discriminative features, offering an effective pre-training solution for 3D representation learning in robotics. Project page: https://kolakivy.github.io/AFRO/
title Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.00074