Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xue, Haoru, He, Tairan, Wang, Zi, Ben, Qingwei, Xiao, Wenli, Luo, Zhengyi, Da, Xingye, Castañeda, Fernando, Shi, Guanya, Sastry, Shankar, Fan, Linxi "Jim", Zhu, Yuke
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912739971039232
author Xue, Haoru
He, Tairan
Wang, Zi
Ben, Qingwei
Xiao, Wenli
Luo, Zhengyi
Da, Xingye
Castañeda, Fernando
Shi, Guanya
Sastry, Shankar
Fan, Linxi "Jim"
Zhu, Yuke
author_facet Xue, Haoru
He, Tairan
Wang, Zi
Ben, Qingwei
Xiao, Wenli
Luo, Zhengyi
Da, Xingye
Castañeda, Fernando
Shi, Guanya
Sastry, Shankar
Fan, Linxi "Jim"
Zhu, Yuke
contents Recent progress in GPU-accelerated, photorealistic simulation has opened a scalable data-generation path for robot learning, where massive physics and visual randomization allow policies to generalize beyond curated environments. Building on these advances, we develop a teacher-student-bootstrap learning framework for vision-based humanoid loco-manipulation, using articulated-object interaction as a representative high-difficulty benchmark. Our approach introduces a staged-reset exploration strategy that stabilizes long-horizon privileged-policy training, and a GRPO-based fine-tuning procedure that mitigates partial observability and improves closed-loop consistency in sim-to-real RL. Trained entirely on simulation data, the resulting policy achieves robust zero-shot performance across diverse door types and outperforms human teleoperators by up to 31.7% in task completion time under the same whole-body control stack. This represents the first humanoid sim-to-real policy capable of diverse articulated loco-manipulation using pure RGB perception.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01061
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
Xue, Haoru
He, Tairan
Wang, Zi
Ben, Qingwei
Xiao, Wenli
Luo, Zhengyi
Da, Xingye
Castañeda, Fernando
Shi, Guanya
Sastry, Shankar
Fan, Linxi "Jim"
Zhu, Yuke
Robotics
Computer Vision and Pattern Recognition
Recent progress in GPU-accelerated, photorealistic simulation has opened a scalable data-generation path for robot learning, where massive physics and visual randomization allow policies to generalize beyond curated environments. Building on these advances, we develop a teacher-student-bootstrap learning framework for vision-based humanoid loco-manipulation, using articulated-object interaction as a representative high-difficulty benchmark. Our approach introduces a staged-reset exploration strategy that stabilizes long-horizon privileged-policy training, and a GRPO-based fine-tuning procedure that mitigates partial observability and improves closed-loop consistency in sim-to-real RL. Trained entirely on simulation data, the resulting policy achieves robust zero-shot performance across diverse door types and outperforms human teleoperators by up to 31.7% in task completion time under the same whole-body control stack. This represents the first humanoid sim-to-real policy capable of diverse articulated loco-manipulation using pure RGB perception.
title Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.01061