Structured Reinforcement Learning for Combinatorial Decision-Making

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hoppe, Heiko, Baty, Léo, Bouvier, Louis, Parmentier, Axel, Schiffer, Maximilian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918174264393728
author Hoppe, Heiko
Baty, Léo
Bouvier, Louis
Parmentier, Axel
Schiffer, Maximilian
author_facet Hoppe, Heiko
Baty, Léo
Bouvier, Louis
Parmentier, Axel
Schiffer, Maximilian
contents Reinforcement learning (RL) is increasingly applied to real-world problems involving complex and structured decisions, such as routing, scheduling, and assortment planning. These settings challenge standard RL algorithms, which struggle to scale, generalize, and exploit structure in the presence of combinatorial action spaces. We propose Structured Reinforcement Learning (SRL), a novel actor-critic paradigm that embeds combinatorial optimization-layers into the actor neural network. We enable end-to-end learning of the actor via Fenchel-Young losses and provide a geometric interpretation of SRL as a primal-dual algorithm in the dual of the moment polytope. Across six environments with exogenous and endogenous uncertainty, SRL matches or surpasses the performance of unstructured RL and imitation learning on static tasks and improves over these baselines by up to 92% on dynamic problems, with improved stability and convergence speed.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19053
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structured Reinforcement Learning for Combinatorial Decision-Making
Hoppe, Heiko
Baty, Léo
Bouvier, Louis
Parmentier, Axel
Schiffer, Maximilian
Machine Learning
Optimization and Control
Reinforcement learning (RL) is increasingly applied to real-world problems involving complex and structured decisions, such as routing, scheduling, and assortment planning. These settings challenge standard RL algorithms, which struggle to scale, generalize, and exploit structure in the presence of combinatorial action spaces. We propose Structured Reinforcement Learning (SRL), a novel actor-critic paradigm that embeds combinatorial optimization-layers into the actor neural network. We enable end-to-end learning of the actor via Fenchel-Young losses and provide a geometric interpretation of SRL as a primal-dual algorithm in the dual of the moment polytope. Across six environments with exogenous and endogenous uncertainty, SRL matches or surpasses the performance of unstructured RL and imitation learning on static tasks and improves over these baselines by up to 92% on dynamic problems, with improved stability and convergence speed.
title Structured Reinforcement Learning for Combinatorial Decision-Making
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2505.19053