PC3D: Zero-Shot Cooperation Across Variable Rosters via Personalized Context Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akman, Ahmet Onur, Kucharski, Rafał
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916000708952064
author Akman, Ahmet Onur
Kucharski, Rafał
author_facet Akman, Ahmet Onur
Kucharski, Rafał
contents Cooperative multi-agent reinforcement learning often assumes a fixed execution team, yet many decentralized systems must operate with varying numbers of active agents during deployment. We study this setting under episodic roster variation: each episode is executed by a set of homogeneous agents, with the team size varying across episodes. Agents act only from local histories, without execution-time communication, privileged coordinators, or online retraining. Therefore, effective cooperation requires each agent to recover relevant context about the active team and adapt its behavior accordingly. To this end, we propose PC3D (Personalized Central Coordination Context Distillation), a method for training decentralized policies to recover and use personalized coordination context from local interaction histories. During training, a set-structured centralized teacher compresses the active team into coordination tokens and personalizes them into agent-specific contexts, which are distilled into decentralized policies. At execution, each agent predicts its own context from local history and adaptively uses it to condition decision-making. Across three cooperative MARL benchmarks, PC3D achieves higher returns than the evaluated baselines with both seen and unseen roster sizes, and ablations attribute these gains to both context distillation and adaptive context use.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10377
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PC3D: Zero-Shot Cooperation Across Variable Rosters via Personalized Context Distillation
Akman, Ahmet Onur
Kucharski, Rafał
Machine Learning
Multiagent Systems
Cooperative multi-agent reinforcement learning often assumes a fixed execution team, yet many decentralized systems must operate with varying numbers of active agents during deployment. We study this setting under episodic roster variation: each episode is executed by a set of homogeneous agents, with the team size varying across episodes. Agents act only from local histories, without execution-time communication, privileged coordinators, or online retraining. Therefore, effective cooperation requires each agent to recover relevant context about the active team and adapt its behavior accordingly. To this end, we propose PC3D (Personalized Central Coordination Context Distillation), a method for training decentralized policies to recover and use personalized coordination context from local interaction histories. During training, a set-structured centralized teacher compresses the active team into coordination tokens and personalizes them into agent-specific contexts, which are distilled into decentralized policies. At execution, each agent predicts its own context from local history and adaptively uses it to condition decision-making. Across three cooperative MARL benchmarks, PC3D achieves higher returns than the evaluated baselines with both seen and unseen roster sizes, and ablations attribute these gains to both context distillation and adaptive context use.
title PC3D: Zero-Shot Cooperation Across Variable Rosters via Personalized Context Distillation
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2605.10377