Hyperspherical Normalization for Scalable Deep Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Hojoon, Lee, Youngdo, Seno, Takuma, Kim, Donghu, Stone, Peter, Choo, Jaegul
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909626797129728
author Lee, Hojoon
Lee, Youngdo
Seno, Takuma
Kim, Donghu
Stone, Peter
Choo, Jaegul
author_facet Lee, Hojoon
Lee, Youngdo
Seno, Takuma
Kim, Donghu
Stone, Peter
Choo, Jaegul
contents Scaling up the model size and computation has brought consistent performance improvements in supervised learning. However, this lesson often fails to apply to reinforcement learning (RL) because training the model on non-stationary data easily leads to overfitting and unstable optimization. In response, we introduce SimbaV2, a novel RL architecture designed to stabilize optimization by (i) constraining the growth of weight and feature norm by hyperspherical normalization; and (ii) using a distributional value estimation with reward scaling to maintain stable gradients under varying reward magnitudes. Using the soft actor-critic as a base algorithm, SimbaV2 scales up effectively with larger models and greater compute, achieving state-of-the-art performance on 57 continuous control tasks across 4 domains. The code is available at https://dojeon-ai.github.io/SimbaV2.
format Preprint
id arxiv_https___arxiv_org_abs_2502_15280
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hyperspherical Normalization for Scalable Deep Reinforcement Learning
Lee, Hojoon
Lee, Youngdo
Seno, Takuma
Kim, Donghu
Stone, Peter
Choo, Jaegul
Machine Learning
Scaling up the model size and computation has brought consistent performance improvements in supervised learning. However, this lesson often fails to apply to reinforcement learning (RL) because training the model on non-stationary data easily leads to overfitting and unstable optimization. In response, we introduce SimbaV2, a novel RL architecture designed to stabilize optimization by (i) constraining the growth of weight and feature norm by hyperspherical normalization; and (ii) using a distributional value estimation with reward scaling to maintain stable gradients under varying reward magnitudes. Using the soft actor-critic as a base algorithm, SimbaV2 scales up effectively with larger models and greater compute, achieving state-of-the-art performance on 57 continuous control tasks across 4 domains. The code is available at https://dojeon-ai.github.io/SimbaV2.
title Hyperspherical Normalization for Scalable Deep Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2502.15280