Learning Sim-to-Real Humanoid Locomotion in 15 Minutes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seo, Younggyo, Sferrazza, Carmelo, Chen, Juyue, Shi, Guanya, Duan, Rocky, Abbeel, Pieter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917116545859584
author Seo, Younggyo
Sferrazza, Carmelo
Chen, Juyue
Shi, Guanya
Duan, Rocky
Abbeel, Pieter
author_facet Seo, Younggyo
Sferrazza, Carmelo
Chen, Juyue
Shi, Guanya
Duan, Rocky
Abbeel, Pieter
contents Massively parallel simulation has reduced reinforcement learning (RL) training time for robots from days to minutes. However, achieving fast and reliable sim-to-real RL for humanoid control remains difficult due to the challenges introduced by factors such as high dimensionality and domain randomization. In this work, we introduce a simple and practical recipe based on off-policy RL algorithms, i.e., FastSAC and FastTD3, that enables rapid training of humanoid locomotion policies in just 15 minutes with a single RTX 4090 GPU. Our simple recipe stabilizes off-policy RL algorithms at massive scale with thousands of parallel environments through carefully tuned design choices and minimalist reward functions. We demonstrate rapid end-to-end learning of humanoid locomotion controllers on Unitree G1 and Booster T1 robots under strong domain randomization, e.g., randomized dynamics, rough terrain, and push perturbations, as well as fast training of whole-body human-motion tracking policies. We provide videos and open-source implementation at: https://younggyo.me/fastsac-humanoid.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01996
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
Seo, Younggyo
Sferrazza, Carmelo
Chen, Juyue
Shi, Guanya
Duan, Rocky
Abbeel, Pieter
Robotics
Artificial Intelligence
Machine Learning
Massively parallel simulation has reduced reinforcement learning (RL) training time for robots from days to minutes. However, achieving fast and reliable sim-to-real RL for humanoid control remains difficult due to the challenges introduced by factors such as high dimensionality and domain randomization. In this work, we introduce a simple and practical recipe based on off-policy RL algorithms, i.e., FastSAC and FastTD3, that enables rapid training of humanoid locomotion policies in just 15 minutes with a single RTX 4090 GPU. Our simple recipe stabilizes off-policy RL algorithms at massive scale with thousands of parallel environments through carefully tuned design choices and minimalist reward functions. We demonstrate rapid end-to-end learning of humanoid locomotion controllers on Unitree G1 and Booster T1 robots under strong domain randomization, e.g., randomized dynamics, rough terrain, and push perturbations, as well as fast training of whole-body human-motion tracking policies. We provide videos and open-source implementation at: https://younggyo.me/fastsac-humanoid.
title Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.01996