FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915317025865728 |
|---|---|
| author | Seo, Younggyo Sferrazza, Carmelo Geng, Haoran Nauman, Michal Yin, Zhao-Heng Abbeel, Pieter |
| author_facet | Seo, Younggyo Sferrazza, Carmelo Geng, Haoran Nauman, Michal Yin, Zhao-Heng Abbeel, Pieter |
| contents | Reinforcement learning (RL) has driven significant progress in robotics, but its complexity and long training times remain major bottlenecks. In this report, we introduce FastTD3, a simple, fast, and capable RL algorithm that significantly speeds up training for humanoid robots in popular suites such as HumanoidBench, IsaacLab, and MuJoCo Playground. Our recipe is remarkably simple: we train an off-policy TD3 agent with several modifications -- parallel simulation, large-batch updates, a distributional critic, and carefully tuned hyperparameters. FastTD3 solves a range of HumanoidBench tasks in under 3 hours on a single A100 GPU, while remaining stable during training. We also provide a lightweight and easy-to-use implementation of FastTD3 to accelerate RL research in robotics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_22642 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control Seo, Younggyo Sferrazza, Carmelo Geng, Haoran Nauman, Michal Yin, Zhao-Heng Abbeel, Pieter Robotics Artificial Intelligence Machine Learning Reinforcement learning (RL) has driven significant progress in robotics, but its complexity and long training times remain major bottlenecks. In this report, we introduce FastTD3, a simple, fast, and capable RL algorithm that significantly speeds up training for humanoid robots in popular suites such as HumanoidBench, IsaacLab, and MuJoCo Playground. Our recipe is remarkably simple: we train an off-policy TD3 agent with several modifications -- parallel simulation, large-batch updates, a distributional critic, and carefully tuned hyperparameters. FastTD3 solves a range of HumanoidBench tasks in under 3 hours on a single A100 GPU, while remaining stable during training. We also provide a lightweight and easy-to-use implementation of FastTD3 to accelerate RL research in robotics. |
| title | FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control |
| topic | Robotics Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2505.22642 |