Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Sizhe, Xu, Linning, Li, Hao, Mu, Juncheng, Zeng, Jia, Lin, Dahua, Pang, Jiangmiao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914525662412800
author Yang, Sizhe
Xu, Linning
Li, Hao
Mu, Juncheng
Zeng, Jia
Lin, Dahua
Pang, Jiangmiao
author_facet Yang, Sizhe
Xu, Linning
Li, Hao
Mu, Juncheng
Zeng, Jia
Lin, Dahua
Pang, Jiangmiao
contents 3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models lack the precision and metric consistency required for physical interaction. We introduce Robo3R, a feed-forward, manipulation-ready 3D reconstruction model that predicts accurate, metric-scale scene geometry directly from RGB images and robot states in real time. Robo3R jointly infers scale-invariant local geometry and relative camera poses, which are unified into the scene representation in the canonical robot frame via a learned global similarity transformation. To meet the precision demands of manipulation, Robo3R employs a masked point head for sharp, fine-grained point clouds, and a keypoint-based Perspective-n-Point (PnP) formulation to refine camera extrinsics and global alignment. Trained on Robo3R-4M, a curated large-scale synthetic dataset with four million high-fidelity annotated frames, Robo3R consistently outperforms state-of-the-art reconstruction methods and depth sensors. Across downstream tasks including imitation learning, sim-to-real transfer, grasp synthesis, and collision-free motion planning, we observe consistent gains in performance, suggesting the promise of this alternative 3D sensing module for robotic manipulation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10101
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
Yang, Sizhe
Xu, Linning
Li, Hao
Mu, Juncheng
Zeng, Jia
Lin, Dahua
Pang, Jiangmiao
Robotics
3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models lack the precision and metric consistency required for physical interaction. We introduce Robo3R, a feed-forward, manipulation-ready 3D reconstruction model that predicts accurate, metric-scale scene geometry directly from RGB images and robot states in real time. Robo3R jointly infers scale-invariant local geometry and relative camera poses, which are unified into the scene representation in the canonical robot frame via a learned global similarity transformation. To meet the precision demands of manipulation, Robo3R employs a masked point head for sharp, fine-grained point clouds, and a keypoint-based Perspective-n-Point (PnP) formulation to refine camera extrinsics and global alignment. Trained on Robo3R-4M, a curated large-scale synthetic dataset with four million high-fidelity annotated frames, Robo3R consistently outperforms state-of-the-art reconstruction methods and depth sensors. Across downstream tasks including imitation learning, sim-to-real transfer, grasp synthesis, and collision-free motion planning, we observe consistent gains in performance, suggesting the promise of this alternative 3D sensing module for robotic manipulation.
title Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
topic Robotics
url https://arxiv.org/abs/2602.10101