Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jing, Yuheng, Li, Kai, Zhang, Ziwen, Zhang, Jiajun, Ma, Zeyao, Yang, Jiaxi, Zhang, Lei, Wu, Zhe, He, Jinmin, Xing, Junliang, Cheng, Jian
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914594784542720
author Jing, Yuheng
Li, Kai
Zhang, Ziwen
Zhang, Jiajun
Ma, Zeyao
Yang, Jiaxi
Zhang, Lei
Wu, Zhe
He, Jinmin
Xing, Junliang
Cheng, Jian
author_facet Jing, Yuheng
Li, Kai
Zhang, Ziwen
Zhang, Jiajun
Ma, Zeyao
Yang, Jiaxi
Zhang, Lei
Wu, Zhe
He, Jinmin
Xing, Junliang
Cheng, Jian
contents In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT)-where coordination with unknown partners is required-remains unexplored. To rigorously evaluate this, we introduce a large-scale benchmark ICRL4AHT, built upon a high-throughput JAX implementation of Overcooked-V2. Our benchmark includes a large, diverse teammate suite spanning both RL and heuristic policies, enabling controlled train-test shifts, and provides a reproducible end-to-end pipeline for teammate generation, learning-history collection, dataset construction, and online multi-episode evaluation. We evaluate representative history-conditioned ICRL algorithms, including Algorithm Distillation (AD) and Decision-Pretrained Transformer (DPT), across millions of transitions. Results reveal notable limitations: contrary to their success in single-agent domains, these baselines fail to exhibit robust test-time adaptation in multi-agent settings. Specifically, these methods frequently underperform random baselines across both unseen teammate and unseen layout tracks, with no clear in-context improvement over long horizons. These findings highlight the challenges of strategic inference under partial observability within the OvercookedV2 AHT protocol, establishing our benchmark as a critical testbed for next-generation coordination algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24423
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
Jing, Yuheng
Li, Kai
Zhang, Ziwen
Zhang, Jiajun
Ma, Zeyao
Yang, Jiaxi
Zhang, Lei
Wu, Zhe
He, Jinmin
Xing, Junliang
Cheng, Jian
Artificial Intelligence
68T05, 68T07, 93A16
I.2.11; I.2.6
In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT)-where coordination with unknown partners is required-remains unexplored. To rigorously evaluate this, we introduce a large-scale benchmark ICRL4AHT, built upon a high-throughput JAX implementation of Overcooked-V2. Our benchmark includes a large, diverse teammate suite spanning both RL and heuristic policies, enabling controlled train-test shifts, and provides a reproducible end-to-end pipeline for teammate generation, learning-history collection, dataset construction, and online multi-episode evaluation. We evaluate representative history-conditioned ICRL algorithms, including Algorithm Distillation (AD) and Decision-Pretrained Transformer (DPT), across millions of transitions. Results reveal notable limitations: contrary to their success in single-agent domains, these baselines fail to exhibit robust test-time adaptation in multi-agent settings. Specifically, these methods frequently underperform random baselines across both unseen teammate and unseen layout tracks, with no clear in-context improvement over long horizons. These findings highlight the challenges of strategic inference under partial observability within the OvercookedV2 AHT protocol, establishing our benchmark as a critical testbed for next-generation coordination algorithms.
title Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
topic Artificial Intelligence
68T05, 68T07, 93A16
I.2.11; I.2.6
url https://arxiv.org/abs/2605.24423