RealD$^2$iff: Bridging Real-World Gap in Robot Manipulation via Depth Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liang, Xiujian, Liu, Jiacheng, Sun, Mingyang, He, Qichen, Lu, Cewu, Sun, Jianhua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912753832165376
author Liang, Xiujian
Liu, Jiacheng
Sun, Mingyang
He, Qichen
Lu, Cewu
Sun, Jianhua
author_facet Liang, Xiujian
Liu, Jiacheng
Sun, Mingyang
He, Qichen
Lu, Cewu
Sun, Jianhua
contents Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise patterns inherent to real sensors. In this work, inspired by the denoising capability of diffusion models, we invert the conventional perspective and propose a clean-to-noisy paradigm that learns to synthesize noisy depth, thereby bridging the visual sim2real gap through purely simulation-driven robotic learning. Building on this idea, we introduce RealD$^2$iff, a hierarchical coarse-to-fine diffusion framework that decomposes depth noise into global structural distortions and fine-grained local perturbations. To enable progressive learning of these components, we further develop two complementary strategies: Frequency-Guided Supervision (FGS) for global structure modeling and Discrepancy-Guided Optimization (DGO) for localized refinement. To integrate RealD$^2$iff seamlessly into imitation learning, we construct a pipeline that spans six stages. We provide comprehensive empirical and experimental validation demonstrating the effectiveness of this paradigm. RealD$^2$iff enables two key applications: (1) generating real-world-like depth to construct clean-noisy paired datasets without manual sensor data collection. (2) Achieving zero-shot sim2real robot manipulation, substantially improving real-world performance without additional fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealD$^2$iff: Bridging Real-World Gap in Robot Manipulation via Depth Diffusion
Liang, Xiujian
Liu, Jiacheng
Sun, Mingyang
He, Qichen
Lu, Cewu
Sun, Jianhua
Robotics
Computer Vision and Pattern Recognition
Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise patterns inherent to real sensors. In this work, inspired by the denoising capability of diffusion models, we invert the conventional perspective and propose a clean-to-noisy paradigm that learns to synthesize noisy depth, thereby bridging the visual sim2real gap through purely simulation-driven robotic learning. Building on this idea, we introduce RealD$^2$iff, a hierarchical coarse-to-fine diffusion framework that decomposes depth noise into global structural distortions and fine-grained local perturbations. To enable progressive learning of these components, we further develop two complementary strategies: Frequency-Guided Supervision (FGS) for global structure modeling and Discrepancy-Guided Optimization (DGO) for localized refinement. To integrate RealD$^2$iff seamlessly into imitation learning, we construct a pipeline that spans six stages. We provide comprehensive empirical and experimental validation demonstrating the effectiveness of this paradigm. RealD$^2$iff enables two key applications: (1) generating real-world-like depth to construct clean-noisy paired datasets without manual sensor data collection. (2) Achieving zero-shot sim2real robot manipulation, substantially improving real-world performance without additional fine-tuning.
title RealD$^2$iff: Bridging Real-World Gap in Robot Manipulation via Depth Diffusion
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22505