Adapt3R: Adaptive 3D Scene Representation for Domain Transfer in Imitation Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wilcox, Albert, Ghanem, Mohamed, Moghani, Masoud, Barroso, Pierre, Joffe, Benjamin, Garg, Animesh
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916739148677120
author Wilcox, Albert
Ghanem, Mohamed
Moghani, Masoud
Barroso, Pierre
Joffe, Benjamin
Garg, Animesh
author_facet Wilcox, Albert
Ghanem, Mohamed
Moghani, Masoud
Barroso, Pierre
Joffe, Benjamin
Garg, Animesh
contents Imitation Learning can train robots to perform complex and diverse manipulation tasks, but learned policies are brittle with observations outside of the training distribution. 3D scene representations that incorporate observations from calibrated RGBD cameras have been proposed as a way to mitigate this, but in our evaluations with unseen embodiments and camera viewpoints they show only modest improvement. To address those challenges, we propose Adapt3R, a general-purpose 3D observation encoder which synthesizes data from calibrated RGBD cameras into a vector that can be used as conditioning for arbitrary IL algorithms. The key idea is to use a pretrained 2D backbone to extract semantic information, using 3D only as a medium to localize this information with respect to the end-effector. We show across 93 simulated and 6 real tasks that when trained end-to-end with a variety of IL algorithms, Adapt3R maintains these algorithms' learning capacity while enabling zero-shot transfer to novel embodiments and camera poses.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04877
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adapt3R: Adaptive 3D Scene Representation for Domain Transfer in Imitation Learning
Wilcox, Albert
Ghanem, Mohamed
Moghani, Masoud
Barroso, Pierre
Joffe, Benjamin
Garg, Animesh
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Imitation Learning can train robots to perform complex and diverse manipulation tasks, but learned policies are brittle with observations outside of the training distribution. 3D scene representations that incorporate observations from calibrated RGBD cameras have been proposed as a way to mitigate this, but in our evaluations with unseen embodiments and camera viewpoints they show only modest improvement. To address those challenges, we propose Adapt3R, a general-purpose 3D observation encoder which synthesizes data from calibrated RGBD cameras into a vector that can be used as conditioning for arbitrary IL algorithms. The key idea is to use a pretrained 2D backbone to extract semantic information, using 3D only as a medium to localize this information with respect to the end-effector. We show across 93 simulated and 6 real tasks that when trained end-to-end with a variety of IL algorithms, Adapt3R maintains these algorithms' learning capacity while enabling zero-shot transfer to novel embodiments and camera poses.
title Adapt3R: Adaptive 3D Scene Representation for Domain Transfer in Imitation Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2503.04877