Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917476271390720 |
|---|---|
| author | Hoppe, Heiko Akkerman, Fabian van Heeswijk, Wouter Schiffer, Maximilian |
| author_facet | Hoppe, Heiko Akkerman, Fabian van Heeswijk, Wouter Schiffer, Maximilian |
| contents | Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender systems, but existing algorithms struggle with the curse of dimensionality in such large discrete action spaces. We propose Distance-Guided Reinforcement Learning (DGRL), combining Sampled Dynamic Neighborhoods and Distance-Based Updates to enable efficient RL in problems with up to $10^{20}$ actions. Unlike prior methods, DGRL performs stochastic volumetric exploration and transforms policy optimization into a stable regression task, decoupling gradient variance from action space cardinality. On structured tasks, DGRL provably guarantees local value improvement. DGRL naturally generalizes to hybrid continuous-discrete action spaces. We demonstrate performance improvements of up to 66% against state-of-the-art benchmarks across regularly and irregularly structured environments, while simultaneously improving convergence speed and computational complexity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_08616 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces Hoppe, Heiko Akkerman, Fabian van Heeswijk, Wouter Schiffer, Maximilian Machine Learning Artificial Intelligence Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender systems, but existing algorithms struggle with the curse of dimensionality in such large discrete action spaces. We propose Distance-Guided Reinforcement Learning (DGRL), combining Sampled Dynamic Neighborhoods and Distance-Based Updates to enable efficient RL in problems with up to $10^{20}$ actions. Unlike prior methods, DGRL performs stochastic volumetric exploration and transforms policy optimization into a stable regression task, decoupling gradient variance from action space cardinality. On structured tasks, DGRL provably guarantees local value improvement. DGRL naturally generalizes to hybrid continuous-discrete action spaces. We demonstrate performance improvements of up to 66% against state-of-the-art benchmarks across regularly and irregularly structured environments, while simultaneously improving convergence speed and computational complexity. |
| title | Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2602.08616 |