KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pavel, Monirul Islam, Hu, Siyi, Masum, Muhammad Anwar, Pratama, Mahardhika, Kowalczyk, Ryszard, Cao, Zehong Jimmy
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908947342950400
author Pavel, Monirul Islam
Hu, Siyi
Masum, Muhammad Anwar
Pratama, Mahardhika
Kowalczyk, Ryszard
Cao, Zehong Jimmy
author_facet Pavel, Monirul Islam
Hu, Siyi
Masum, Muhammad Anwar
Pratama, Mahardhika
Kowalczyk, Ryszard
Cao, Zehong Jimmy
contents Real world deployment of multi agent reinforcement learning MARL systems is fundamentally constrained by limited compute memory and inference time. While expert policies achieve high performance they rely on costly decision cycles and large scale models that are impractical for edge devices or embedded platforms. Knowledge distillation KD offers a promising path toward resource aware execution but existing KD methods in MARL focus narrowly on action imitation often neglecting coordination structure and assuming uniform agent capabilities. We propose resource aware Knowledge Distillation for Multi Agent Reinforcement Learning KD MARL a two stage framework that transfers coordinated behavior from a centralized expert to lightweight decentralized student agents. The student policies are trained without a critic relying instead on distilled advantage signals and structured policy supervision to preserve coordination under heterogeneous and limited observations. Our approach transfers both action level behavior and structural coordination patterns from expert policies while supporting heterogeneous student architectures allowing each agent model capacity to match its observation complexity which is crucial for efficient execution under partial or limited observability and limited onboard resources. Extensive experiments on SMAC and MPE benchmarks demonstrate that KD MARL achieves high performance retention while substantially reducing computational cost. Across standard multi agent benchmarks KD MARL retains over 90 percent of expert performance while reducing computational cost by up to 28.6 times FLOPs. The proposed approach achieves expert level coordination and preserves it through structured distillation enabling practical MARL deployment across resource constrained onboard platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06691
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning
Pavel, Monirul Islam
Hu, Siyi
Masum, Muhammad Anwar
Pratama, Mahardhika
Kowalczyk, Ryszard
Cao, Zehong Jimmy
Artificial Intelligence
Real world deployment of multi agent reinforcement learning MARL systems is fundamentally constrained by limited compute memory and inference time. While expert policies achieve high performance they rely on costly decision cycles and large scale models that are impractical for edge devices or embedded platforms. Knowledge distillation KD offers a promising path toward resource aware execution but existing KD methods in MARL focus narrowly on action imitation often neglecting coordination structure and assuming uniform agent capabilities. We propose resource aware Knowledge Distillation for Multi Agent Reinforcement Learning KD MARL a two stage framework that transfers coordinated behavior from a centralized expert to lightweight decentralized student agents. The student policies are trained without a critic relying instead on distilled advantage signals and structured policy supervision to preserve coordination under heterogeneous and limited observations. Our approach transfers both action level behavior and structural coordination patterns from expert policies while supporting heterogeneous student architectures allowing each agent model capacity to match its observation complexity which is crucial for efficient execution under partial or limited observability and limited onboard resources. Extensive experiments on SMAC and MPE benchmarks demonstrate that KD MARL achieves high performance retention while substantially reducing computational cost. Across standard multi agent benchmarks KD MARL retains over 90 percent of expert performance while reducing computational cost by up to 28.6 times FLOPs. The proposed approach achieves expert level coordination and preserves it through structured distillation enabling practical MARL deployment across resource constrained onboard platforms.
title KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2604.06691