Understanding and Improving Hyperbolic Deep Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Klein, Timo, Lang, Thomas, Shkabrii, Andrii, Sturm, Alexander, Sidak, Kevin, Miklautz, Lukas, Plant, Claudia, Velaj, Yllka, Tschiatschek, Sebastian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911491782868992
author Klein, Timo
Lang, Thomas
Shkabrii, Andrii
Sturm, Alexander
Sidak, Kevin
Miklautz, Lukas
Plant, Claudia
Velaj, Yllka
Tschiatschek, Sebastian
author_facet Klein, Timo
Lang, Thomas
Shkabrii, Andrii
Sturm, Alexander
Sidak, Kevin
Miklautz, Lukas
Plant, Claudia
Velaj, Yllka
Tschiatschek, Sebastian
contents The exponential volume growth of hyperbolic geometry can embed the hierarchical relationships between states in reinforcement learning (RL) with far less distortion than Euclidean space. However, hyperbolic deep RL faces severe optimization challenges, and formal analysis of why optimization fails is lacking. We identify key factors that determine the success and failure of training hyperbolic deep RL agents. By analyzing the gradients of core operations in the Poincaré Ball and Hyperboloid models of hyperbolic geometry, we show that large-norm embeddings destabilize gradient-based training, leading to trust-region violations in proximal policy optimization (PPO). Based on these insights, we introduce Hyper++, a new hyperbolic deep RL agent that consists of three components: (1) feature regularization guaranteeing bounded norms while avoiding the curse of dimensionality from clipping; (2) a categorical value loss for stable critic training; and (3) a more optimization-friendly formulation of hyperbolic network layers. On ProcGen, we show that Hyper++ guarantees stable learning, outperforms prior hyperbolic agents, and reduces wall-clock time by approximately 30%. On Atari-5 with Double DQN, Hyper++ strongly outperforms Euclidean and hyperbolic baselines. We release our code at https://github.com/Probabilistic-and-Interactive-ML/hyper-rl.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14202
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding and Improving Hyperbolic Deep Reinforcement Learning
Klein, Timo
Lang, Thomas
Shkabrii, Andrii
Sturm, Alexander
Sidak, Kevin
Miklautz, Lukas
Plant, Claudia
Velaj, Yllka
Tschiatschek, Sebastian
Machine Learning
Artificial Intelligence
The exponential volume growth of hyperbolic geometry can embed the hierarchical relationships between states in reinforcement learning (RL) with far less distortion than Euclidean space. However, hyperbolic deep RL faces severe optimization challenges, and formal analysis of why optimization fails is lacking. We identify key factors that determine the success and failure of training hyperbolic deep RL agents. By analyzing the gradients of core operations in the Poincaré Ball and Hyperboloid models of hyperbolic geometry, we show that large-norm embeddings destabilize gradient-based training, leading to trust-region violations in proximal policy optimization (PPO). Based on these insights, we introduce Hyper++, a new hyperbolic deep RL agent that consists of three components: (1) feature regularization guaranteeing bounded norms while avoiding the curse of dimensionality from clipping; (2) a categorical value loss for stable critic training; and (3) a more optimization-friendly formulation of hyperbolic network layers. On ProcGen, we show that Hyper++ guarantees stable learning, outperforms prior hyperbolic agents, and reduces wall-clock time by approximately 30%. On Atari-5 with Double DQN, Hyper++ strongly outperforms Euclidean and hyperbolic baselines. We release our code at https://github.com/Probabilistic-and-Interactive-ML/hyper-rl.
title Understanding and Improving Hyperbolic Deep Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.14202