FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seo, Younggyo, Sferrazza, Carmelo, Geng, Haoran, Nauman, Michal, Yin, Zhao-Heng, Abbeel, Pieter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915317025865728
author Seo, Younggyo
Sferrazza, Carmelo
Geng, Haoran
Nauman, Michal
Yin, Zhao-Heng
Abbeel, Pieter
author_facet Seo, Younggyo
Sferrazza, Carmelo
Geng, Haoran
Nauman, Michal
Yin, Zhao-Heng
Abbeel, Pieter
contents Reinforcement learning (RL) has driven significant progress in robotics, but its complexity and long training times remain major bottlenecks. In this report, we introduce FastTD3, a simple, fast, and capable RL algorithm that significantly speeds up training for humanoid robots in popular suites such as HumanoidBench, IsaacLab, and MuJoCo Playground. Our recipe is remarkably simple: we train an off-policy TD3 agent with several modifications -- parallel simulation, large-batch updates, a distributional critic, and carefully tuned hyperparameters. FastTD3 solves a range of HumanoidBench tasks in under 3 hours on a single A100 GPU, while remaining stable during training. We also provide a lightweight and easy-to-use implementation of FastTD3 to accelerate RL research in robotics.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22642
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
Seo, Younggyo
Sferrazza, Carmelo
Geng, Haoran
Nauman, Michal
Yin, Zhao-Heng
Abbeel, Pieter
Robotics
Artificial Intelligence
Machine Learning
Reinforcement learning (RL) has driven significant progress in robotics, but its complexity and long training times remain major bottlenecks. In this report, we introduce FastTD3, a simple, fast, and capable RL algorithm that significantly speeds up training for humanoid robots in popular suites such as HumanoidBench, IsaacLab, and MuJoCo Playground. Our recipe is remarkably simple: we train an off-policy TD3 agent with several modifications -- parallel simulation, large-batch updates, a distributional critic, and carefully tuned hyperparameters. FastTD3 solves a range of HumanoidBench tasks in under 3 hours on a single A100 GPU, while remaining stable during training. We also provide a lightweight and easy-to-use implementation of FastTD3 to accelerate RL research in robotics.
title FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.22642