An Empirical Study of Automating Agent Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Kang, Woo, Sangmin, Ding, Haibo, Ramnath, Kiran, Chidambaram, Subramanian, Feng, Aosong, Arannil, Vinayak, Kim, Muhyun, Singh, Ishan, Wang, Darren, Xu, Zhichao, Gandhi, Megha, Prabhu, Nirmal, Mishra, Soumya Smruti, Singh, Vivek, Pandeshwar, Gouri, Cheong, Lin Lee
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909035109810176
author Zhou, Kang
Woo, Sangmin
Ding, Haibo
Ramnath, Kiran
Chidambaram, Subramanian
Feng, Aosong
Arannil, Vinayak
Kim, Muhyun
Singh, Ishan
Wang, Darren
Xu, Zhichao
Gandhi, Megha
Prabhu, Nirmal
Mishra, Soumya Smruti
Singh, Vivek
Pandeshwar, Gouri
Cheong, Lin Lee
author_facet Zhou, Kang
Woo, Sangmin
Ding, Haibo
Ramnath, Kiran
Chidambaram, Subramanian
Feng, Aosong
Arannil, Vinayak
Kim, Muhyun
Singh, Ishan
Wang, Darren
Xu, Zhichao
Gandhi, Megha
Prabhu, Nirmal
Mishra, Soumya Smruti
Singh, Vivek
Pandeshwar, Gouri
Cheong, Lin Lee
contents Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive. A natural question arises: can frontier coding assistants reliably automate this evaluation process? Our study shows that simply prompting coding assistants is insufficient for this task. Without domain-specific evaluation knowledge, frontier coding assistants achieve only a 30% execution success rate and produce over-engineered evaluations averaging 12+ metrics per agent, indicating that strong coding ability does not automatically translate to reliable agent evaluation. We introduce EvalAgent, an AI assistant that automates the end-to-end agent evaluation pipeline. EvalAgent encodes evaluation domain expertise as evaluation skills (procedural instructions, reusable code and templates, and dynamically retrieved API documentation) that compose into a trace-based pipeline producing complete evaluation artifacts including metrics, executable code, and reports. To systematically assess generated evaluations, we introduce a meta-evaluation framework alongside AgentEvalBench, a benchmark comprising 20 agents, each paired with evaluation requirements and test scenarios. We further propose the Eval@1 metric to measure whether generated evaluation code both executes and yields meaningful results on the first run. Our experiments show that EvalAgent produces focused evaluations, improving Eval@1 from 17.5% to 65%, and achieving 79.5% human expert preference over baseline approaches. Further ablation studies show that evaluation skills are critical for handling complex evaluation: removing them causes Eval@1 to drop significantly from 65% to 30%.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11378
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle An Empirical Study of Automating Agent Evaluation
Zhou, Kang
Woo, Sangmin
Ding, Haibo
Ramnath, Kiran
Chidambaram, Subramanian
Feng, Aosong
Arannil, Vinayak
Kim, Muhyun
Singh, Ishan
Wang, Darren
Xu, Zhichao
Gandhi, Megha
Prabhu, Nirmal
Mishra, Soumya Smruti
Singh, Vivek
Pandeshwar, Gouri
Cheong, Lin Lee
Computation and Language
Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive. A natural question arises: can frontier coding assistants reliably automate this evaluation process? Our study shows that simply prompting coding assistants is insufficient for this task. Without domain-specific evaluation knowledge, frontier coding assistants achieve only a 30% execution success rate and produce over-engineered evaluations averaging 12+ metrics per agent, indicating that strong coding ability does not automatically translate to reliable agent evaluation. We introduce EvalAgent, an AI assistant that automates the end-to-end agent evaluation pipeline. EvalAgent encodes evaluation domain expertise as evaluation skills (procedural instructions, reusable code and templates, and dynamically retrieved API documentation) that compose into a trace-based pipeline producing complete evaluation artifacts including metrics, executable code, and reports. To systematically assess generated evaluations, we introduce a meta-evaluation framework alongside AgentEvalBench, a benchmark comprising 20 agents, each paired with evaluation requirements and test scenarios. We further propose the Eval@1 metric to measure whether generated evaluation code both executes and yields meaningful results on the first run. Our experiments show that EvalAgent produces focused evaluations, improving Eval@1 from 17.5% to 65%, and achieving 79.5% human expert preference over baseline approaches. Further ablation studies show that evaluation skills are critical for handling complex evaluation: removing them causes Eval@1 to drop significantly from 65% to 30%.
title An Empirical Study of Automating Agent Evaluation
topic Computation and Language
url https://arxiv.org/abs/2605.11378