Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Yuxin, Cao, Zhiguang, Gu, Chengyang, Liu, Liu, Zhao, Peilin, Chen, Yize, Lin, Fangzhen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911229761552384
author Pan, Yuxin
Cao, Zhiguang
Gu, Chengyang
Liu, Liu
Zhao, Peilin
Chen, Yize
Lin, Fangzhen
author_facet Pan, Yuxin
Cao, Zhiguang
Gu, Chengyang
Liu, Liu
Zhao, Peilin
Chen, Yize
Lin, Fangzhen
contents Existing neural methods for multi-task vehicle routing problems (VRPs) typically learn unified solvers to handle multiple constraints simultaneously. However, they often underutilize the compositional structure of VRP variants, each derivable from a common set of basis VRP variants. This critical oversight causes unified solvers to miss out the potential benefits of basis solvers, each specialized for a basis VRP variant. To overcome this limitation, we propose a framework that enables unified solvers to perceive the shared-component nature across VRP variants by proactively reusing basis solvers, while mitigating the exponential growth of trained neural solvers. Specifically, we introduce a State-Decomposable MDP (SDMDP) that reformulates VRPs by expressing the state space as the Cartesian product of basis state spaces associated with basis VRP variants. More crucially, this formulation inherently yields the optimal basis policy for each basis VRP variant. Furthermore, a Latent Space-based SDMDP extension is developed by incorporating both the optimal basis policies and a learnable mixture function to enable the policy reuse in the latent space. Under mild assumptions, this extension provably recovers the optimal unified policy of SDMDP through the mixture function that computes the state embedding as a mapping from the basis state embeddings generated by optimal basis policies. For practical implementation, we introduce the Mixture-of-Specialized-Experts Solver (MoSES), which realizes basis policies through specialized Low-Rank Adaptation (LoRA) experts, and implements the mixture function via an adaptive gating mechanism. Extensive experiments conducted across VRP variants showcase the superiority of MoSES over prior methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21453
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDP
Pan, Yuxin
Cao, Zhiguang
Gu, Chengyang
Liu, Liu
Zhao, Peilin
Chen, Yize
Lin, Fangzhen
Artificial Intelligence
Machine Learning
Existing neural methods for multi-task vehicle routing problems (VRPs) typically learn unified solvers to handle multiple constraints simultaneously. However, they often underutilize the compositional structure of VRP variants, each derivable from a common set of basis VRP variants. This critical oversight causes unified solvers to miss out the potential benefits of basis solvers, each specialized for a basis VRP variant. To overcome this limitation, we propose a framework that enables unified solvers to perceive the shared-component nature across VRP variants by proactively reusing basis solvers, while mitigating the exponential growth of trained neural solvers. Specifically, we introduce a State-Decomposable MDP (SDMDP) that reformulates VRPs by expressing the state space as the Cartesian product of basis state spaces associated with basis VRP variants. More crucially, this formulation inherently yields the optimal basis policy for each basis VRP variant. Furthermore, a Latent Space-based SDMDP extension is developed by incorporating both the optimal basis policies and a learnable mixture function to enable the policy reuse in the latent space. Under mild assumptions, this extension provably recovers the optimal unified policy of SDMDP through the mixture function that computes the state embedding as a mapping from the basis state embeddings generated by optimal basis policies. For practical implementation, we introduce the Mixture-of-Specialized-Experts Solver (MoSES), which realizes basis policies through specialized Low-Rank Adaptation (LoRA) experts, and implements the mixture function via an adaptive gating mechanism. Extensive experiments conducted across VRP variants showcase the superiority of MoSES over prior methods.
title Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDP
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.21453