Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pang, Xincheng, Xia, Wenke, Wang, Zhigang, Zhao, Bin, Hu, Di, Wang, Dong, Li, Xuelong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917744729915392
author Pang, Xincheng
Xia, Wenke
Wang, Zhigang
Zhao, Bin
Hu, Di
Wang, Dong
Li, Xuelong
author_facet Pang, Xincheng
Xia, Wenke
Wang, Zhigang
Zhao, Bin
Hu, Di
Wang, Dong
Li, Xuelong
contents 3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their effectiveness in fine-grained robotic manipulation tasks. To address these limitations, we propose a Depth Information Injection ($\bold{DI}^{\bold{2}}$) framework that leverages the RGB-Depth modality for policy fine-tuning, while relying solely on RGB images for robust and efficient deployment. Concretely, we introduce the Depth Completion Module (DCM) to extract the spatial prior knowledge related to depth information and generate virtual depth information from RGB inputs to aid policy deployment. Further, we propose the Depth-Aware Codebook (DAC) to eliminate noise and reduce the cumulative error from the depth prediction. In the inference phase, this framework employs RGB inputs and accurately predicted depth data to generate the manipulation action. We conduct experiments on simulated LIBERO environments and real-world scenarios, and the experiment results prove that our method could effectively enhance the pre-trained RGB-based policy with 3D perception ability for robotic manipulation. The website is released at https://gewu-lab.github.io/DepthHelps-IROS2024.
format Preprint
id arxiv_https___arxiv_org_abs_2408_05107
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection
Pang, Xincheng
Xia, Wenke
Wang, Zhigang
Zhao, Bin
Hu, Di
Wang, Dong
Li, Xuelong
Robotics
3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their effectiveness in fine-grained robotic manipulation tasks. To address these limitations, we propose a Depth Information Injection ($\bold{DI}^{\bold{2}}$) framework that leverages the RGB-Depth modality for policy fine-tuning, while relying solely on RGB images for robust and efficient deployment. Concretely, we introduce the Depth Completion Module (DCM) to extract the spatial prior knowledge related to depth information and generate virtual depth information from RGB inputs to aid policy deployment. Further, we propose the Depth-Aware Codebook (DAC) to eliminate noise and reduce the cumulative error from the depth prediction. In the inference phase, this framework employs RGB inputs and accurately predicted depth data to generate the manipulation action. We conduct experiments on simulated LIBERO environments and real-world scenarios, and the experiment results prove that our method could effectively enhance the pre-trained RGB-based policy with 3D perception ability for robotic manipulation. The website is released at https://gewu-lab.github.io/DepthHelps-IROS2024.
title Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection
topic Robotics
url https://arxiv.org/abs/2408.05107