ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Zifan, Gong, Ran, Minniti, Maria Vittoria, Sivakumar, Kausik, Gundogdu, Ahmet Salih, Rosen, Eric, Yan, Riedana, Kusnur, Tushar, Wang, Zixing, Deng, Di, Stone, Peter, Zhang, Xiaohan, Schmeckpeper, Karl
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910280758329344
author Xu, Zifan
Gong, Ran
Minniti, Maria Vittoria
Sivakumar, Kausik
Gundogdu, Ahmet Salih
Rosen, Eric
Yan, Riedana
Kusnur, Tushar
Wang, Zixing
Deng, Di
Stone, Peter
Zhang, Xiaohan
Schmeckpeper, Karl
author_facet Xu, Zifan
Gong, Ran
Minniti, Maria Vittoria
Sivakumar, Kausik
Gundogdu, Ahmet Salih
Rosen, Eric
Yan, Riedana
Kusnur, Tushar
Wang, Zixing
Deng, Di
Stone, Peter
Zhang, Xiaohan
Schmeckpeper, Karl
contents Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (e.g., through teleoperation) serve as the standard source for expert behaviors, acquiring such data at scale in the real world is prohibitively expensive. This paper introduces ExpertGen, a framework that automates expert policy learning in simulation to enable scalable sim-to-real transfer. ExpertGen first initializes a behavior prior using a diffusion policy trained on imperfect demonstrations, which may be synthesized by large language models or provided by humans. Reinforcement learning is then used to steer this prior toward high task success by optimizing the diffusion model's initial noise while keep original policy frozen. By keeping the pretrained diffusion policy frozen, ExpertGen regularizes exploration to remain within safe, human-like behavior manifolds, while also enabling effective learning with only sparse rewards. Empirical evaluations on challenging manipulation benchmarks demonstrate that ExpertGen reliably produces high-quality expert policies with no reward engineering. On industrial assembly tasks, ExpertGen achieves a 90.5% overall success rate, while on long-horizon manipulation tasks it attains 85% overall success, outperforming all baseline methods. The resulting policies exhibit dexterous control and remain robust across diverse initial configurations and failure states. To validate sim-to-real transfer, the learned state-based expert policies are further distilled into visuomotor policies via DAgger and successfully deployed on real robotic hardware.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15956
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors
Xu, Zifan
Gong, Ran
Minniti, Maria Vittoria
Sivakumar, Kausik
Gundogdu, Ahmet Salih
Rosen, Eric
Yan, Riedana
Kusnur, Tushar
Wang, Zixing
Deng, Di
Stone, Peter
Zhang, Xiaohan
Schmeckpeper, Karl
Robotics
Artificial Intelligence
Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (e.g., through teleoperation) serve as the standard source for expert behaviors, acquiring such data at scale in the real world is prohibitively expensive. This paper introduces ExpertGen, a framework that automates expert policy learning in simulation to enable scalable sim-to-real transfer. ExpertGen first initializes a behavior prior using a diffusion policy trained on imperfect demonstrations, which may be synthesized by large language models or provided by humans. Reinforcement learning is then used to steer this prior toward high task success by optimizing the diffusion model's initial noise while keep original policy frozen. By keeping the pretrained diffusion policy frozen, ExpertGen regularizes exploration to remain within safe, human-like behavior manifolds, while also enabling effective learning with only sparse rewards. Empirical evaluations on challenging manipulation benchmarks demonstrate that ExpertGen reliably produces high-quality expert policies with no reward engineering. On industrial assembly tasks, ExpertGen achieves a 90.5% overall success rate, while on long-horizon manipulation tasks it attains 85% overall success, outperforming all baseline methods. The resulting policies exhibit dexterous control and remain robust across diverse initial configurations and failure states. To validate sim-to-real transfer, the learned state-based expert policies are further distilled into visuomotor policies via DAgger and successfully deployed on real robotic hardware.
title ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2603.15956