CooperBench: Why Coding Agents Cannot be Your Teammates Yet

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khatua, Arpandeep, Zhu, Hao, Tran, Peter, Prabhudesai, Arya, Sadrieh, Frederic, Lieberwirth, Johann K., Yu, Xinkai, Fu, Yicheng, Ryan, Michael J., Pei, Jiaxin, Yang, Diyi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908788307525632
author Khatua, Arpandeep
Zhu, Hao
Tran, Peter
Prabhudesai, Arya
Sadrieh, Frederic
Lieberwirth, Johann K.
Yu, Xinkai
Fu, Yicheng
Ryan, Michael J.
Pei, Jiaxin
Yang, Diyi
author_facet Khatua, Arpandeep
Zhu, Hao
Tran, Peter
Prabhudesai, Arya
Sadrieh, Frederic
Lieberwirth, Johann K.
Yu, Xinkai
Fu, Yicheng
Ryan, Michael J.
Pei, Jiaxin
Yang, Diyi
contents Resolving team conflicts requires not only task-specific competence, but also social intelligence to find common ground and build consensus. As AI agents increasingly collaborate on complex work, they must develop coordination capabilities to function as effective teammates. Yet we hypothesize that current agents lack these capabilities. To test this, we introduce CooperBench, a benchmark of over 600 collaborative coding tasks across 12 libraries in 4 programming languages. Each task assigns two agents different features that can be implemented independently but may conflict without proper coordination. Tasks are grounded in real open-source repositories with expert-written tests. Evaluating state-of-the-art coding agents, we observe the curse of coordination: agents achieve on average 30% lower success rates when working together compared to performing both tasks individually. This contrasts sharply with human teams, where adding teammates typically improves productivity. Our analysis reveals three key issues: (1) communication channels become jammed with vague, ill-timed, and inaccurate messages; (2) even with effective communication, agents deviate from their commitments; and (3) agents often hold incorrect expectations about others' plans and communication. Through large-scale simulation, we also observe rare but interesting emergent coordination behavior including role division, resource division, and negotiation. Our research presents a novel benchmark for collaborative coding and calls for a shift from pursuing individual agent capability to developing social intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13295
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CooperBench: Why Coding Agents Cannot be Your Teammates Yet
Khatua, Arpandeep
Zhu, Hao
Tran, Peter
Prabhudesai, Arya
Sadrieh, Frederic
Lieberwirth, Johann K.
Yu, Xinkai
Fu, Yicheng
Ryan, Michael J.
Pei, Jiaxin
Yang, Diyi
Machine Learning
Artificial Intelligence
Computation and Language
Multiagent Systems
Social and Information Networks
Resolving team conflicts requires not only task-specific competence, but also social intelligence to find common ground and build consensus. As AI agents increasingly collaborate on complex work, they must develop coordination capabilities to function as effective teammates. Yet we hypothesize that current agents lack these capabilities. To test this, we introduce CooperBench, a benchmark of over 600 collaborative coding tasks across 12 libraries in 4 programming languages. Each task assigns two agents different features that can be implemented independently but may conflict without proper coordination. Tasks are grounded in real open-source repositories with expert-written tests. Evaluating state-of-the-art coding agents, we observe the curse of coordination: agents achieve on average 30% lower success rates when working together compared to performing both tasks individually. This contrasts sharply with human teams, where adding teammates typically improves productivity. Our analysis reveals three key issues: (1) communication channels become jammed with vague, ill-timed, and inaccurate messages; (2) even with effective communication, agents deviate from their commitments; and (3) agents often hold incorrect expectations about others' plans and communication. Through large-scale simulation, we also observe rare but interesting emergent coordination behavior including role division, resource division, and negotiation. Our research presents a novel benchmark for collaborative coding and calls for a shift from pursuing individual agent capability to developing social intelligence.
title CooperBench: Why Coding Agents Cannot be Your Teammates Yet
topic Machine Learning
Artificial Intelligence
Computation and Language
Multiagent Systems
Social and Information Networks
url https://arxiv.org/abs/2601.13295