CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Lingyue, Ding, Xin, Pan, Linyue, Zhu, Yaoming, Zhang, Shao, Qiu, Lin, Liu, Weiwen, Zhang, Weinan, Cao, Xuezhi, Cai, Xunliang, Ding, Jiaxin, Yu, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915763315539968
author Fu, Lingyue
Ding, Xin
Pan, Linyue
Zhu, Yaoming
Zhang, Shao
Qiu, Lin
Liu, Weiwen
Zhang, Weinan
Cao, Xuezhi
Cai, Xunliang
Ding, Jiaxin
Yu, Yong
author_facet Fu, Lingyue
Ding, Xin
Pan, Linyue
Zhu, Yaoming
Zhang, Shao
Qiu, Lin
Liu, Weiwen
Zhang, Weinan
Cao, Xuezhi
Cai, Xunliang
Ding, Jiaxin
Yu, Yong
contents Current evaluation for Large Language Model (LLM) code agents predominantly focus on generating functional code in single-turn scenarios, which fails to evaluate the agent's capability for continuous code optimization and multi-turn iterative development. To bridge this gap, we introduce CATArena, a framework designed to evaluate the evolutionary capabilities of code agents via iterative tournaments. Agents engage in multi-turn tournaments and continuously refine their code through self-reflection and peer-learning based on comprehensive execution feedback. For evaluation, we propose a dual-metric system to decouple static generation proficiency from evolutionary potential. Extensive experiments reveal that an agent's evolutionary potential is not strictly correlated with its initial proficiency. Our analysis further reveals that current agents struggle to concurrently leverage both peer-learning and self-reflection for effective performance gains. Furthermore, the results validate CATArena's high extensibility and resistance to variance tasks, establishing it as a continuous and reliable standard for assessing the evolutionary capability of LLM code agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
Fu, Lingyue
Ding, Xin
Pan, Linyue
Zhu, Yaoming
Zhang, Shao
Qiu, Lin
Liu, Weiwen
Zhang, Weinan
Cao, Xuezhi
Cai, Xunliang
Ding, Jiaxin
Yu, Yong
Artificial Intelligence
Computation and Language
Current evaluation for Large Language Model (LLM) code agents predominantly focus on generating functional code in single-turn scenarios, which fails to evaluate the agent's capability for continuous code optimization and multi-turn iterative development. To bridge this gap, we introduce CATArena, a framework designed to evaluate the evolutionary capabilities of code agents via iterative tournaments. Agents engage in multi-turn tournaments and continuously refine their code through self-reflection and peer-learning based on comprehensive execution feedback. For evaluation, we propose a dual-metric system to decouple static generation proficiency from evolutionary potential. Extensive experiments reveal that an agent's evolutionary potential is not strictly correlated with its initial proficiency. Our analysis further reveals that current agents struggle to concurrently leverage both peer-learning and self-reflection for effective performance gains. Furthermore, the results validate CATArena's high extensibility and resistance to variance tasks, establishing it as a continuous and reliable standard for assessing the evolutionary capability of LLM code agents.
title CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.26852