EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Boyuan, Meng, Xinpan, Wang, Xiaofeng, Zhu, Zheng, Ye, Angen, Wang, Yang, Yang, Zhiqin, Ni, Chaojun, Huang, Guan, Wang, Xingang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909677675085824
author Wang, Boyuan
Meng, Xinpan
Wang, Xiaofeng
Zhu, Zheng
Ye, Angen
Wang, Yang
Yang, Zhiqin
Ni, Chaojun
Huang, Guan
Wang, Xingang
author_facet Wang, Boyuan
Meng, Xinpan
Wang, Xiaofeng
Zhu, Zheng
Ye, Angen
Wang, Yang
Yang, Zhiqin
Ni, Chaojun
Huang, Guan
Wang, Xingang
contents The rapid advancement of Embodied AI has led to an increasing demand for large-scale, high-quality real-world data. However, collecting such embodied data remains costly and inefficient. As a result, simulation environments have become a crucial surrogate for training robot policies. Yet, the significant Real2Sim2Real gap remains a critical bottleneck, particularly in terms of physical dynamics and visual appearance. To address this challenge, we propose EmbodieDreamer, a novel framework that reduces the Real2Sim2Real gap from both the physics and appearance perspectives. Specifically, we propose PhysAligner, a differentiable physics module designed to reduce the Real2Sim physical gap. It jointly optimizes robot-specific parameters such as control gains and friction coefficients to better align simulated dynamics with real-world observations. In addition, we introduce VisAligner, which incorporates a conditional video diffusion model to bridge the Sim2Real appearance gap by translating low-fidelity simulated renderings into photorealistic videos conditioned on simulation states, enabling high-fidelity visual transfer. Extensive experiments validate the effectiveness of EmbodieDreamer. The proposed PhysAligner reduces physical parameter estimation error by 3.74% compared to simulated annealing methods while improving optimization speed by 89.91\%. Moreover, training robot policies in the generated photorealistic environment leads to a 29.17% improvement in the average task success rate across real-world tasks after reinforcement learning. Code, model and data will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05198
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
Wang, Boyuan
Meng, Xinpan
Wang, Xiaofeng
Zhu, Zheng
Ye, Angen
Wang, Yang
Yang, Zhiqin
Ni, Chaojun
Huang, Guan
Wang, Xingang
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
The rapid advancement of Embodied AI has led to an increasing demand for large-scale, high-quality real-world data. However, collecting such embodied data remains costly and inefficient. As a result, simulation environments have become a crucial surrogate for training robot policies. Yet, the significant Real2Sim2Real gap remains a critical bottleneck, particularly in terms of physical dynamics and visual appearance. To address this challenge, we propose EmbodieDreamer, a novel framework that reduces the Real2Sim2Real gap from both the physics and appearance perspectives. Specifically, we propose PhysAligner, a differentiable physics module designed to reduce the Real2Sim physical gap. It jointly optimizes robot-specific parameters such as control gains and friction coefficients to better align simulated dynamics with real-world observations. In addition, we introduce VisAligner, which incorporates a conditional video diffusion model to bridge the Sim2Real appearance gap by translating low-fidelity simulated renderings into photorealistic videos conditioned on simulation states, enabling high-fidelity visual transfer. Extensive experiments validate the effectiveness of EmbodieDreamer. The proposed PhysAligner reduces physical parameter estimation error by 3.74% compared to simulated annealing methods while improving optimization speed by 89.91\%. Moreover, training robot policies in the generated photorealistic environment leads to a 29.17% improvement in the average task success rate across real-world tasks after reinforcement learning. Code, model and data will be publicly available.
title EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.05198