R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuhao, Dong, Wanxi, Shi, Yue, Liang, Yi, Gao, Jingnan, Yang, Qiaochu, Lyu, Yaxing, Liang, Zhixuan, Liu, Yibin, Xu, Congsheng, Guo, Xianda, Sui, Wei, Jin, Yaohui, Yang, Xiaokang, Xu, Yanyan, Mu, Yao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911549506977792
author Zhang, Yuhao
Dong, Wanxi
Shi, Yue
Liang, Yi
Gao, Jingnan
Yang, Qiaochu
Lyu, Yaxing
Liang, Zhixuan
Liu, Yibin
Xu, Congsheng
Guo, Xianda
Sui, Wei
Jin, Yaohui
Yang, Xiaokang
Xu, Yanyan
Mu, Yao
author_facet Zhang, Yuhao
Dong, Wanxi
Shi, Yue
Liang, Yi
Gao, Jingnan
Yang, Qiaochu
Lyu, Yaxing
Liang, Zhixuan
Liu, Yibin
Xu, Congsheng
Guo, Xianda
Sui, Wei
Jin, Yaohui
Yang, Xiaokang
Xu, Yanyan
Mu, Yao
contents Embodied manipulation requires accurate 3D understanding of objects and their spatial relations to plan and execute contact-rich actions. While large-scale 3D vision models provide strong priors, their computational cost incurs prohibitive latency for real-time control. We propose Real-time 3D-aware Policy (R3DP), which integrates powerful 3D priors into manipulation policies without sacrificing real-time performance. A core innovation of R3DP is the asynchronous fast-slow collaboration module, which seamlessly integrates large-scale 3D priors into the policy without compromising real-time performance. The system maintains real-time efficiency by querying the pre-trained slow system (VGGT) only on sparse key frames, while simultaneously employing a lightweight Temporal Feature Prediction Network (TFPNet) to predict features for all intermediate frames. By leveraging historical data to exploit temporal correlations, TFPNet explicitly improves task success rates through consistent feature estimation. Additionally, to enable more effective multi-view fusion, we introduce a Multi-View Feature Fuser (MVFF) that aggregates features across views by explicitly incorporating camera intrinsics and extrinsics. R3DP offers a plug-and-play solution for integrating large models into real-time inference systems. We evaluate R3DP against multiple baselines across different visual configurations. R3DP effectively harnesses large-scale 3D priors to achieve superior results, outperforming single-view and multi-view DP by 32.9% and 51.4% in average success rate, respectively. Furthermore, by decoupling heavy 3D reasoning from policy execution, R3DP achieves a 44.8% reduction in inference time compared to a naive DP+VGGT integration.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14498
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
Zhang, Yuhao
Dong, Wanxi
Shi, Yue
Liang, Yi
Gao, Jingnan
Yang, Qiaochu
Lyu, Yaxing
Liang, Zhixuan
Liu, Yibin
Xu, Congsheng
Guo, Xianda
Sui, Wei
Jin, Yaohui
Yang, Xiaokang
Xu, Yanyan
Mu, Yao
Robotics
Computer Vision and Pattern Recognition
Embodied manipulation requires accurate 3D understanding of objects and their spatial relations to plan and execute contact-rich actions. While large-scale 3D vision models provide strong priors, their computational cost incurs prohibitive latency for real-time control. We propose Real-time 3D-aware Policy (R3DP), which integrates powerful 3D priors into manipulation policies without sacrificing real-time performance. A core innovation of R3DP is the asynchronous fast-slow collaboration module, which seamlessly integrates large-scale 3D priors into the policy without compromising real-time performance. The system maintains real-time efficiency by querying the pre-trained slow system (VGGT) only on sparse key frames, while simultaneously employing a lightweight Temporal Feature Prediction Network (TFPNet) to predict features for all intermediate frames. By leveraging historical data to exploit temporal correlations, TFPNet explicitly improves task success rates through consistent feature estimation. Additionally, to enable more effective multi-view fusion, we introduce a Multi-View Feature Fuser (MVFF) that aggregates features across views by explicitly incorporating camera intrinsics and extrinsics. R3DP offers a plug-and-play solution for integrating large models into real-time inference systems. We evaluate R3DP against multiple baselines across different visual configurations. R3DP effectively harnesses large-scale 3D priors to achieve superior results, outperforming single-view and multi-view DP by 32.9% and 51.4% in average success rate, respectively. Furthermore, by decoupling heavy 3D reasoning from policy execution, R3DP achieves a 44.8% reduction in inference time compared to a naive DP+VGGT integration.
title R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.14498