AgentBench: Evaluating LLMs as Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Xiao, Yu, Hao, Zhang, Hanchen, Xu, Yifan, Lei, Xuanyu, Lai, Hanyu, Gu, Yu, Ding, Hangliang, Men, Kaiwen, Yang, Kejuan, Zhang, Shudan, Deng, Xiang, Zeng, Aohan, Du, Zhengxiao, Zhang, Chenhui, Shen, Sheng, Zhang, Tianjun, Su, Yu, Sun, Huan, Huang, Minlie, Dong, Yuxiao, Tang, Jie
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911191298736128
author Liu, Xiao
Yu, Hao
Zhang, Hanchen
Xu, Yifan
Lei, Xuanyu
Lai, Hanyu
Gu, Yu
Ding, Hangliang
Men, Kaiwen
Yang, Kejuan
Zhang, Shudan
Deng, Xiang
Zeng, Aohan
Du, Zhengxiao
Zhang, Chenhui
Shen, Sheng
Zhang, Tianjun
Su, Yu
Sun, Huan
Huang, Minlie
Dong, Yuxiao
Tang, Jie
author_facet Liu, Xiao
Yu, Hao
Zhang, Hanchen
Xu, Yifan
Lei, Xuanyu
Lai, Hanyu
Gu, Yu
Ding, Hangliang
Men, Kaiwen
Yang, Kejuan
Zhang, Shudan
Deng, Xiang
Zeng, Aohan
Du, Zhengxiao
Zhang, Chenhui
Shen, Sheng
Zhang, Tianjun
Su, Yu
Sun, Huan
Huang, Minlie
Dong, Yuxiao
Tang, Jie
contents The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively \textit{evaluate LLMs as agents} on challenging tasks in interactive environments. We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities. Our extensive test over \num API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B. We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents. Improving instruction following and training on high quality multi-round alignment data could improve agent performance. And different from existing assumptions, training on code present ambivalent impacts on different agent tasks. Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench.
format Preprint
id arxiv_https___arxiv_org_abs_2308_03688
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AgentBench: Evaluating LLMs as Agents
Liu, Xiao
Yu, Hao
Zhang, Hanchen
Xu, Yifan
Lei, Xuanyu
Lai, Hanyu
Gu, Yu
Ding, Hangliang
Men, Kaiwen
Yang, Kejuan
Zhang, Shudan
Deng, Xiang
Zeng, Aohan
Du, Zhengxiao
Zhang, Chenhui
Shen, Sheng
Zhang, Tianjun
Su, Yu
Sun, Huan
Huang, Minlie
Dong, Yuxiao
Tang, Jie
Artificial Intelligence
Computation and Language
Machine Learning
The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively \textit{evaluate LLMs as agents} on challenging tasks in interactive environments. We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities. Our extensive test over \num API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B. We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents. Improving instruction following and training on high quality multi-round alignment data could improve agent performance. And different from existing assumptions, training on code present ambivalent impacts on different agent tasks. Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench.
title AgentBench: Evaluating LLMs as Agents
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2308.03688