Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruhdorfer, Constantin, Bortoletto, Matteo, Oei, Victor, Penzkofer, Anna, Bulling, Andreas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911098722058240
author Ruhdorfer, Constantin
Bortoletto, Matteo
Oei, Victor
Penzkofer, Anna
Bulling, Andreas
author_facet Ruhdorfer, Constantin
Bortoletto, Matteo
Oei, Victor
Penzkofer, Anna
Bulling, Andreas
contents We introduce Unsupervised Partner Design (UPD) - a population-free, multi-agent reinforcement learning framework for robust ad-hoc teamwork that adaptively generates training partners without requiring pretrained partners or manual parameter tuning. UPD constructs diverse partners by stochastically mixing an ego agent's policy with biased random behaviours and scores them using a variance-based learnability metric that prioritises partners near the ego agent's current learning frontier. We show that UPD can be integrated with unsupervised environment design, resulting in the first method enabling fully unsupervised curricula over both level and partner distributions in a cooperative setting. Through extensive evaluations on Overcooked-AI and the Overcooked Generalisation Challenge, we demonstrate that this dynamic partner curriculum is highly effective: UPD consistently outperforms both population-based and population-free baselines as well as ablations. In a user study, we further show that UPD achieves higher returns than all baselines and was perceived as significantly more adaptive, more human-like, a better collaborator, and less frustrating.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06336
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Partner Design Enables Robust Ad-hoc Teamwork
Ruhdorfer, Constantin
Bortoletto, Matteo
Oei, Victor
Penzkofer, Anna
Bulling, Andreas
Machine Learning
Artificial Intelligence
Human-Computer Interaction
Multiagent Systems
We introduce Unsupervised Partner Design (UPD) - a population-free, multi-agent reinforcement learning framework for robust ad-hoc teamwork that adaptively generates training partners without requiring pretrained partners or manual parameter tuning. UPD constructs diverse partners by stochastically mixing an ego agent's policy with biased random behaviours and scores them using a variance-based learnability metric that prioritises partners near the ego agent's current learning frontier. We show that UPD can be integrated with unsupervised environment design, resulting in the first method enabling fully unsupervised curricula over both level and partner distributions in a cooperative setting. Through extensive evaluations on Overcooked-AI and the Overcooked Generalisation Challenge, we demonstrate that this dynamic partner curriculum is highly effective: UPD consistently outperforms both population-based and population-free baselines as well as ablations. In a user study, we further show that UPD achieves higher returns than all baselines and was perceived as significantly more adaptive, more human-like, a better collaborator, and less frustrating.
title Unsupervised Partner Design Enables Robust Ad-hoc Teamwork
topic Machine Learning
Artificial Intelligence
Human-Computer Interaction
Multiagent Systems
url https://arxiv.org/abs/2508.06336