Agent-as-a-Judge: Evaluate Agents with Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuge, Mingchen, Zhao, Changsheng, Ashley, Dylan, Wang, Wenyi, Khizbullin, Dmitrii, Xiong, Yunyang, Liu, Zechun, Chang, Ernie, Krishnamoorthi, Raghuraman, Tian, Yuandong, Shi, Yangyang, Chandra, Vikas, Schmidhuber, Jürgen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910653821747200
author Zhuge, Mingchen
Zhao, Changsheng
Ashley, Dylan
Wang, Wenyi
Khizbullin, Dmitrii
Xiong, Yunyang
Liu, Zechun
Chang, Ernie
Krishnamoorthi, Raghuraman
Tian, Yuandong
Shi, Yangyang
Chandra, Vikas
Schmidhuber, Jürgen
author_facet Zhuge, Mingchen
Zhao, Changsheng
Ashley, Dylan
Wang, Wenyi
Khizbullin, Dmitrii
Xiong, Yunyang
Liu, Zechun
Chang, Ernie
Krishnamoorthi, Raghuraman
Tian, Yuandong
Shi, Yangyang
Chandra, Vikas
Schmidhuber, Jürgen
contents Contemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes -- ignoring the step-by-step nature of agentic systems, or require excessive manual labour. To address this, we introduce the Agent-as-a-Judge framework, wherein agentic systems are used to evaluate agentic systems. This is an organic extension of the LLM-as-a-Judge framework, incorporating agentic features that enable intermediate feedback for the entire task-solving process. We apply the Agent-as-a-Judge to the task of code generation. To overcome issues with existing benchmarks and provide a proof-of-concept testbed for Agent-as-a-Judge, we present DevAI, a new benchmark of 55 realistic automated AI development tasks. It includes rich manual annotations, like a total of 365 hierarchical user requirements. We benchmark three of the popular agentic systems using Agent-as-a-Judge and find it dramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline. Altogether, we believe that Agent-as-a-Judge marks a concrete step forward for modern agentic systems -- by providing rich and reliable reward signals necessary for dynamic and scalable self-improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10934
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Agent-as-a-Judge: Evaluate Agents with Agents
Zhuge, Mingchen
Zhao, Changsheng
Ashley, Dylan
Wang, Wenyi
Khizbullin, Dmitrii
Xiong, Yunyang
Liu, Zechun
Chang, Ernie
Krishnamoorthi, Raghuraman
Tian, Yuandong
Shi, Yangyang
Chandra, Vikas
Schmidhuber, Jürgen
Artificial Intelligence
Contemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes -- ignoring the step-by-step nature of agentic systems, or require excessive manual labour. To address this, we introduce the Agent-as-a-Judge framework, wherein agentic systems are used to evaluate agentic systems. This is an organic extension of the LLM-as-a-Judge framework, incorporating agentic features that enable intermediate feedback for the entire task-solving process. We apply the Agent-as-a-Judge to the task of code generation. To overcome issues with existing benchmarks and provide a proof-of-concept testbed for Agent-as-a-Judge, we present DevAI, a new benchmark of 55 realistic automated AI development tasks. It includes rich manual annotations, like a total of 365 hierarchical user requirements. We benchmark three of the popular agentic systems using Agent-as-a-Judge and find it dramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline. Altogether, we believe that Agent-as-a-Judge marks a concrete step forward for modern agentic systems -- by providing rich and reliable reward signals necessary for dynamic and scalable self-improvement.
title Agent-as-a-Judge: Evaluate Agents with Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2410.10934