From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xun, Li, Zhuoran, Lin, Yanshan, Zhong, Hai, Huang, Longbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911248699883520
author Wang, Xun
Li, Zhuoran
Lin, Yanshan
Zhong, Hai
Huang, Longbo
author_facet Wang, Xun
Li, Zhuoran
Lin, Yanshan
Zhong, Hai
Huang, Longbo
contents Training a team of agents from scratch in multi-agent reinforcement learning (MARL) is highly inefficient, much like asking beginners to play a symphony together without first practicing solo. Existing methods, such as offline or transferable MARL, can ease this burden, but they still rely on costly multi-agent data, which often becomes the bottleneck. In contrast, solo experiences are far easier to obtain in many important scenarios, e.g., collaborative coding, household cooperation, and search-and-rescue. To unlock their potential, we propose Solo-to-Collaborative RL (SoCo), a framework that transfers solo knowledge into cooperative learning. SoCo first pretrains a shared solo policy from solo demonstrations, then adapts it for cooperation during multi-agent training through a policy fusion mechanism that combines an MoE-like gating selector and an action editor. Experiments across diverse cooperative tasks show that SoCo significantly boosts the training efficiency and performance of backbone algorithms. These results demonstrate that solo demonstrations provide a scalable and effective complement to multi-agent data, making cooperative learning more practical and broadly applicable.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02762
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos
Wang, Xun
Li, Zhuoran
Lin, Yanshan
Zhong, Hai
Huang, Longbo
Machine Learning
Multiagent Systems
Training a team of agents from scratch in multi-agent reinforcement learning (MARL) is highly inefficient, much like asking beginners to play a symphony together without first practicing solo. Existing methods, such as offline or transferable MARL, can ease this burden, but they still rely on costly multi-agent data, which often becomes the bottleneck. In contrast, solo experiences are far easier to obtain in many important scenarios, e.g., collaborative coding, household cooperation, and search-and-rescue. To unlock their potential, we propose Solo-to-Collaborative RL (SoCo), a framework that transfers solo knowledge into cooperative learning. SoCo first pretrains a shared solo policy from solo demonstrations, then adapts it for cooperation during multi-agent training through a policy fusion mechanism that combines an MoE-like gating selector and an action editor. Experiments across diverse cooperative tasks show that SoCo significantly boosts the training efficiency and performance of backbone algorithms. These results demonstrate that solo demonstrations provide a scalable and effective complement to multi-agent data, making cooperative learning more practical and broadly applicable.
title From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2511.02762