Structured Reinforcement Learning for Combinatorial Decision-Making
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918174264393728 |
|---|---|
| author | Hoppe, Heiko Baty, Léo Bouvier, Louis Parmentier, Axel Schiffer, Maximilian |
| author_facet | Hoppe, Heiko Baty, Léo Bouvier, Louis Parmentier, Axel Schiffer, Maximilian |
| contents | Reinforcement learning (RL) is increasingly applied to real-world problems involving complex and structured decisions, such as routing, scheduling, and assortment planning. These settings challenge standard RL algorithms, which struggle to scale, generalize, and exploit structure in the presence of combinatorial action spaces. We propose Structured Reinforcement Learning (SRL), a novel actor-critic paradigm that embeds combinatorial optimization-layers into the actor neural network. We enable end-to-end learning of the actor via Fenchel-Young losses and provide a geometric interpretation of SRL as a primal-dual algorithm in the dual of the moment polytope. Across six environments with exogenous and endogenous uncertainty, SRL matches or surpasses the performance of unstructured RL and imitation learning on static tasks and improves over these baselines by up to 92% on dynamic problems, with improved stability and convergence speed. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_19053 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Structured Reinforcement Learning for Combinatorial Decision-Making Hoppe, Heiko Baty, Léo Bouvier, Louis Parmentier, Axel Schiffer, Maximilian Machine Learning Optimization and Control Reinforcement learning (RL) is increasingly applied to real-world problems involving complex and structured decisions, such as routing, scheduling, and assortment planning. These settings challenge standard RL algorithms, which struggle to scale, generalize, and exploit structure in the presence of combinatorial action spaces. We propose Structured Reinforcement Learning (SRL), a novel actor-critic paradigm that embeds combinatorial optimization-layers into the actor neural network. We enable end-to-end learning of the actor via Fenchel-Young losses and provide a geometric interpretation of SRL as a primal-dual algorithm in the dual of the moment polytope. Across six environments with exogenous and endogenous uncertainty, SRL matches or surpasses the performance of unstructured RL and imitation learning on static tasks and improves over these baselines by up to 92% on dynamic problems, with improved stability and convergence speed. |
| title | Structured Reinforcement Learning for Combinatorial Decision-Making |
| topic | Machine Learning Optimization and Control |
| url | https://arxiv.org/abs/2505.19053 |