LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junzhe, Hu, Xuming, Liu, Shuodi, Huang, Shiyu, Tu, Wei-Wei, He, Zhaofeng, Wen, Lijie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917597957586944
author Chen, Junzhe
Hu, Xuming
Liu, Shuodi
Huang, Shiyu
Tu, Wei-Wei
He, Zhaofeng
Wen, Lijie
author_facet Chen, Junzhe
Hu, Xuming
Liu, Shuodi
Huang, Shiyu
Tu, Wei-Wei
He, Zhaofeng
Wen, Lijie
contents Recent advancements in large language models (LLMs) have revealed their potential for achieving autonomous agents possessing human-level intelligence. However, existing benchmarks for evaluating LLM Agents either use static datasets, potentially leading to data leakage or focus only on single-agent scenarios, overlooking the complexities of multi-agent interactions. There is a lack of a benchmark that evaluates the diverse capabilities of LLM agents in multi-agent, dynamic environments. To this end, we introduce LLMArena, a novel and easily extensible framework for evaluating the diverse capabilities of LLM in multi-agent dynamic environments. LLMArena encompasses seven distinct gaming environments, employing Trueskill scoring to assess crucial abilities in LLM agents, including spatial reasoning, strategic planning, numerical reasoning, risk assessment, communication, opponent modeling, and team collaboration. We conduct an extensive experiment and human evaluation among different sizes and types of LLMs, showing that LLMs still have a significant journey ahead in their development towards becoming fully autonomous agents, especially in opponent modeling and team collaboration. We hope LLMArena could guide future research towards enhancing these capabilities in LLMs, ultimately leading to more sophisticated and practical applications in dynamic, multi-agent settings. The code and data will be available.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16499
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments
Chen, Junzhe
Hu, Xuming
Liu, Shuodi
Huang, Shiyu
Tu, Wei-Wei
He, Zhaofeng
Wen, Lijie
Computation and Language
Recent advancements in large language models (LLMs) have revealed their potential for achieving autonomous agents possessing human-level intelligence. However, existing benchmarks for evaluating LLM Agents either use static datasets, potentially leading to data leakage or focus only on single-agent scenarios, overlooking the complexities of multi-agent interactions. There is a lack of a benchmark that evaluates the diverse capabilities of LLM agents in multi-agent, dynamic environments. To this end, we introduce LLMArena, a novel and easily extensible framework for evaluating the diverse capabilities of LLM in multi-agent dynamic environments. LLMArena encompasses seven distinct gaming environments, employing Trueskill scoring to assess crucial abilities in LLM agents, including spatial reasoning, strategic planning, numerical reasoning, risk assessment, communication, opponent modeling, and team collaboration. We conduct an extensive experiment and human evaluation among different sizes and types of LLMs, showing that LLMs still have a significant journey ahead in their development towards becoming fully autonomous agents, especially in opponent modeling and team collaboration. We hope LLMArena could guide future research towards enhancing these capabilities in LLMs, ultimately leading to more sophisticated and practical applications in dynamic, multi-agent settings. The code and data will be available.
title LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments
topic Computation and Language
url https://arxiv.org/abs/2402.16499