Saved in:
Bibliographic Details
Main Authors: Mohammadi, Mahmoud, Li, Yipeng, Lo, Jane, Yip, Wendy
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.21504
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908470872113152
author Mohammadi, Mahmoud
Li, Yipeng
Lo, Jane
Yip, Wendy
author_facet Mohammadi, Mahmoud
Li, Yipeng
Lo, Jane
Yip, Wendy
contents The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a two-dimensional taxonomy that organizes existing work along (1) evaluation objectives -- what to evaluate, such as agent behavior, capabilities, reliability, and safety -- and (2) evaluation process -- how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling. In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and long-horizon interactions, and compliance, which are often overlooked in current research. We also identify future research directions, including holistic, more realistic, and scalable evaluation. This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21504
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation and Benchmarking of LLM Agents: A Survey
Mohammadi, Mahmoud
Li, Yipeng
Lo, Jane
Yip, Wendy
Machine Learning
Artificial Intelligence
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a two-dimensional taxonomy that organizes existing work along (1) evaluation objectives -- what to evaluate, such as agent behavior, capabilities, reliability, and safety -- and (2) evaluation process -- how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling. In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and long-horizon interactions, and compliance, which are often overlooked in current research. We also identify future research directions, including holistic, more realistic, and scalable evaluation. This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.
title Evaluation and Benchmarking of LLM Agents: A Survey
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.21504