FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Donghu, Lee, Youngdo, Park, Minho, Kim, Kinam, Nahendra, I Made Aswin, Seno, Takuma, Min, Sehee, Palenicek, Daniel, Vogt, Florian, Kragic, Danica, Peters, Jan, Choo, Jaegul, Lee, Hojoon
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913131676041216
author Kim, Donghu
Lee, Youngdo
Park, Minho
Kim, Kinam
Nahendra, I Made Aswin
Seno, Takuma
Min, Sehee
Palenicek, Daniel
Vogt, Florian
Kragic, Danica
Peters, Jan
Choo, Jaegul
Lee, Hojoon
author_facet Kim, Donghu
Lee, Youngdo
Park, Minho
Kim, Kinam
Nahendra, I Made Aswin
Seno, Takuma
Min, Sehee
Palenicek, Daniel
Vogt, Florian
Kragic, Danica
Peters, Jan
Choo, Jaegul
Lee, Hojoon
contents Reinforcement learning (RL) is a core approach for robot control when expert demonstrations are unavailable. On-policy methods such as Proximal Policy Optimization (PPO) are widely used for their stability, but their reliance on narrowly distributed on-policy data limits accurate policy evaluation in high-dimensional state and action spaces. Off-policy methods can overcome this limitation by learning from a broader state-action distribution, yet suffer from slow convergence and instability, as fitting a value function over diverse data requires many gradient updates, causing critic errors to accumulate through bootstrapping. We present FlashSAC, a fast and stable off-policy RL algorithm built on Soft Actor-Critic. Motivated by scaling laws observed in supervised learning, FlashSAC sharply reduces gradient updates while compensating with larger models and higher data throughput. To maintain stability at increased scale, FlashSAC explicitly bounds weight, feature, and gradient norms, curbing critic error accumulation. Across over 60 tasks in 10 simulators, FlashSAC consistently outperforms PPO and strong off-policy baselines in both final performance and training efficiency, with the largest gains on high-dimensional tasks such as dexterous manipulation. In sim-to-real humanoid locomotion, FlashSAC reduces training time from hours to minutes, demonstrating the promise of off-policy RL for sim-to-real transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04539
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
Kim, Donghu
Lee, Youngdo
Park, Minho
Kim, Kinam
Nahendra, I Made Aswin
Seno, Takuma
Min, Sehee
Palenicek, Daniel
Vogt, Florian
Kragic, Danica
Peters, Jan
Choo, Jaegul
Lee, Hojoon
Machine Learning
Robotics
Reinforcement learning (RL) is a core approach for robot control when expert demonstrations are unavailable. On-policy methods such as Proximal Policy Optimization (PPO) are widely used for their stability, but their reliance on narrowly distributed on-policy data limits accurate policy evaluation in high-dimensional state and action spaces. Off-policy methods can overcome this limitation by learning from a broader state-action distribution, yet suffer from slow convergence and instability, as fitting a value function over diverse data requires many gradient updates, causing critic errors to accumulate through bootstrapping. We present FlashSAC, a fast and stable off-policy RL algorithm built on Soft Actor-Critic. Motivated by scaling laws observed in supervised learning, FlashSAC sharply reduces gradient updates while compensating with larger models and higher data throughput. To maintain stability at increased scale, FlashSAC explicitly bounds weight, feature, and gradient norms, curbing critic error accumulation. Across over 60 tasks in 10 simulators, FlashSAC consistently outperforms PPO and strong off-policy baselines in both final performance and training efficiency, with the largest gains on high-dimensional tasks such as dexterous manipulation. In sim-to-real humanoid locomotion, FlashSAC reduces training time from hours to minutes, demonstrating the promise of off-policy RL for sim-to-real transfer.
title FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
topic Machine Learning
Robotics
url https://arxiv.org/abs/2604.04539