AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Wentao, Wang, Yu, Zhao, Yuyang, Chen, Yuxin, Feng, Fuli, Hao, Xueyuan, Su, Xi, Gu, Qi, Su, Hui, Cai, Xunliang, He, Xiangnan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910148749950976
author Shi, Wentao
Wang, Yu
Zhao, Yuyang
Chen, Yuxin
Feng, Fuli
Hao, Xueyuan
Su, Xi
Gu, Qi
Su, Hui
Cai, Xunliang
He, Xiangnan
author_facet Shi, Wentao
Wang, Yu
Zhao, Yuyang
Chen, Yuxin
Feng, Fuli
Hao, Xueyuan
Su, Xi
Gu, Qi
Su, Hui
Cai, Xunliang
He, Xiangnan
contents As reinforcement learning continues to scale the training of large language model-based agents, reliably verifying agent behaviors in complex environments has become increasingly challenging. Existing approaches rely on rule-based verifiers or LLM-as-a-Judge models, which struggle to generalize beyond narrow domains. Agent-as-a-Judge addresses this limitation by actively interacting with environments and tools to acquire verifiable evidence, yet its capabilities remain underexplored. We introduce a benchmark AJ-Bench to systematically evaluate Agent-as-a-Judge across three domains-search, data systems, and graphical user interfaces-comprising 155 tasks and 516 annotated trajectories. The benchmark comprehensively assesses judge agents' abilities in information acquisition, state verification, and process verification. Experiments demonstrate consistent performance gains over LLM-as-a-Judge baselines, while also revealing substantial open challenges in agent-based verification. Our data and code are available at https://aj-bench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18240
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
Shi, Wentao
Wang, Yu
Zhao, Yuyang
Chen, Yuxin
Feng, Fuli
Hao, Xueyuan
Su, Xi
Gu, Qi
Su, Hui
Cai, Xunliang
He, Xiangnan
Artificial Intelligence
As reinforcement learning continues to scale the training of large language model-based agents, reliably verifying agent behaviors in complex environments has become increasingly challenging. Existing approaches rely on rule-based verifiers or LLM-as-a-Judge models, which struggle to generalize beyond narrow domains. Agent-as-a-Judge addresses this limitation by actively interacting with environments and tools to acquire verifiable evidence, yet its capabilities remain underexplored. We introduce a benchmark AJ-Bench to systematically evaluate Agent-as-a-Judge across three domains-search, data systems, and graphical user interfaces-comprising 155 tasks and 516 annotated trajectories. The benchmark comprehensively assesses judge agents' abilities in information acquisition, state verification, and process verification. Experiments demonstrate consistent performance gains over LLM-as-a-Judge baselines, while also revealing substantial open challenges in agent-based verification. Our data and code are available at https://aj-bench.github.io/.
title AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
topic Artificial Intelligence
url https://arxiv.org/abs/2604.18240