AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Chang, Zhang, Junlei, Zhu, Zhihao, Yang, Cheng, Yang, Yujiu, Jin, Yaohui, Lan, Zhenzhong, Kong, Lingpeng, He, Junxian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916540831498240
author Ma, Chang
Zhang, Junlei
Zhu, Zhihao
Yang, Cheng
Yang, Yujiu
Jin, Yaohui
Lan, Zhenzhong
Kong, Lingpeng
He, Junxian
author_facet Ma, Chang
Zhang, Junlei
Zhu, Zhihao
Yang, Cheng
Yang, Yujiu
Jin, Yaohui
Lan, Zhenzhong
Kong, Lingpeng
He, Junxian
contents Evaluating Large Language Models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial challenges. A primary obstacle is the benchmarking of agent performance across diverse scenarios within a unified framework, especially in maintaining partially-observable environments and ensuring multi-round interactions. Moreover, current evaluation frameworks mostly focus on the final success rate, revealing few insights during the process and failing to provide a deep understanding of the model abilities. To address these challenges, we introduce AgentBoard, a pioneering comprehensive benchmark and accompanied open-source evaluation framework tailored to analytical evaluation of LLM agents. AgentBoard offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis. This not only sheds light on the capabilities and limitations of LLM agents but also propels the interpretability of their performance to the forefront. Ultimately, AgentBoard serves as a step towards demystifying agent behaviors and accelerating the development of stronger LLM agents.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13178
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Ma, Chang
Zhang, Junlei
Zhu, Zhihao
Yang, Cheng
Yang, Yujiu
Jin, Yaohui
Lan, Zhenzhong
Kong, Lingpeng
He, Junxian
Computation and Language
Artificial Intelligence
Machine Learning
Evaluating Large Language Models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial challenges. A primary obstacle is the benchmarking of agent performance across diverse scenarios within a unified framework, especially in maintaining partially-observable environments and ensuring multi-round interactions. Moreover, current evaluation frameworks mostly focus on the final success rate, revealing few insights during the process and failing to provide a deep understanding of the model abilities. To address these challenges, we introduce AgentBoard, a pioneering comprehensive benchmark and accompanied open-source evaluation framework tailored to analytical evaluation of LLM agents. AgentBoard offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis. This not only sheds light on the capabilities and limitations of LLM agents but also propels the interpretability of their performance to the forefront. Ultimately, AgentBoard serves as a step towards demystifying agent behaviors and accelerating the development of stronger LLM agents.
title AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2401.13178