Projecting Latent RL Actions: Towards Generalizable and Scalable Graph Combinatorial Optimization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Terranova, Franco, Bernardez, Guillermo, Cabellos-Aparicio, Albert, Miolane, Nina, Lahmadi, Abdelkader
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913146242859008
author Terranova, Franco
Bernardez, Guillermo
Cabellos-Aparicio, Albert
Miolane, Nina
Lahmadi, Abdelkader
author_facet Terranova, Franco
Bernardez, Guillermo
Cabellos-Aparicio, Albert
Miolane, Nina
Lahmadi, Abdelkader
contents Graph combinatorial optimization (GCO) has attracted growing interest, as many NP-hard problems naturally admit graph formulations, yet their combinatorial explosion renders exact methods computationally intractable. Recent advances in Reinforcement Learning (RL) combined with Graph Neural Networks (GNNs) have significantly improved learning-based GCO solvers. However, existing approaches face limitations in both generalization across diverse graph instances and computational scalability as action spaces grow. To address both challenges, we introduce projection agents, a novel RL-GCO approach that operates directly in a continuous GNN-based action embedding space, predicting a desired latent action in a single forward pass and subsequently decoding it into a valid discrete action. Additionally, we enable fair comparison across RL methods through a shared embedding space for both observations and actions. Across diverse benchmarks, our approach achieves up to 16.2x faster inference and up to 40% better generalization than existing solutions using only simple nearest-neighbor decoding, while opening the door to strong RL performance in super-linear decision spaces with multiple interdependent variables. Finally, we release LaGCO-RL, a Python library that automates latent action-space construction and supports existing RL-GCO solutions, promoting reproducibility and adaptation to new GCO benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19721
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Projecting Latent RL Actions: Towards Generalizable and Scalable Graph Combinatorial Optimization
Terranova, Franco
Bernardez, Guillermo
Cabellos-Aparicio, Albert
Miolane, Nina
Lahmadi, Abdelkader
Artificial Intelligence
Machine Learning
Networking and Internet Architecture
Graph combinatorial optimization (GCO) has attracted growing interest, as many NP-hard problems naturally admit graph formulations, yet their combinatorial explosion renders exact methods computationally intractable. Recent advances in Reinforcement Learning (RL) combined with Graph Neural Networks (GNNs) have significantly improved learning-based GCO solvers. However, existing approaches face limitations in both generalization across diverse graph instances and computational scalability as action spaces grow. To address both challenges, we introduce projection agents, a novel RL-GCO approach that operates directly in a continuous GNN-based action embedding space, predicting a desired latent action in a single forward pass and subsequently decoding it into a valid discrete action. Additionally, we enable fair comparison across RL methods through a shared embedding space for both observations and actions. Across diverse benchmarks, our approach achieves up to 16.2x faster inference and up to 40% better generalization than existing solutions using only simple nearest-neighbor decoding, while opening the door to strong RL performance in super-linear decision spaces with multiple interdependent variables. Finally, we release LaGCO-RL, a Python library that automates latent action-space construction and supports existing RL-GCO solutions, promoting reproducibility and adaptation to new GCO benchmarks.
title Projecting Latent RL Actions: Towards Generalizable and Scalable Graph Combinatorial Optimization
topic Artificial Intelligence
Machine Learning
Networking and Internet Architecture
url https://arxiv.org/abs/2605.19721