Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hoppe, Heiko, Akkerman, Fabian, van Heeswijk, Wouter, Schiffer, Maximilian
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917476271390720
author Hoppe, Heiko
Akkerman, Fabian
van Heeswijk, Wouter
Schiffer, Maximilian
author_facet Hoppe, Heiko
Akkerman, Fabian
van Heeswijk, Wouter
Schiffer, Maximilian
contents Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender systems, but existing algorithms struggle with the curse of dimensionality in such large discrete action spaces. We propose Distance-Guided Reinforcement Learning (DGRL), combining Sampled Dynamic Neighborhoods and Distance-Based Updates to enable efficient RL in problems with up to $10^{20}$ actions. Unlike prior methods, DGRL performs stochastic volumetric exploration and transforms policy optimization into a stable regression task, decoupling gradient variance from action space cardinality. On structured tasks, DGRL provably guarantees local value improvement. DGRL naturally generalizes to hybrid continuous-discrete action spaces. We demonstrate performance improvements of up to 66% against state-of-the-art benchmarks across regularly and irregularly structured environments, while simultaneously improving convergence speed and computational complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08616
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces
Hoppe, Heiko
Akkerman, Fabian
van Heeswijk, Wouter
Schiffer, Maximilian
Machine Learning
Artificial Intelligence
Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender systems, but existing algorithms struggle with the curse of dimensionality in such large discrete action spaces. We propose Distance-Guided Reinforcement Learning (DGRL), combining Sampled Dynamic Neighborhoods and Distance-Based Updates to enable efficient RL in problems with up to $10^{20}$ actions. Unlike prior methods, DGRL performs stochastic volumetric exploration and transforms policy optimization into a stable regression task, decoupling gradient variance from action space cardinality. On structured tasks, DGRL provably guarantees local value improvement. DGRL naturally generalizes to hybrid continuous-discrete action spaces. We demonstrate performance improvements of up to 66% against state-of-the-art benchmarks across regularly and irregularly structured environments, while simultaneously improving convergence speed and computational complexity.
title Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.08616