Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Jiachen, Zhu, Menghui, Rui, Renting, Shan, Rong, Zheng, Congmin, Chen, Bo, Xi, Yunjia, Lin, Jianghao, Liu, Weiwen, Tang, Ruiming, Yu, Yong, Zhang, Weinan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918057329295360
author Zhu, Jiachen
Zhu, Menghui
Rui, Renting
Shan, Rong
Zheng, Congmin
Chen, Bo
Xi, Yunjia
Lin, Jianghao
Liu, Weiwen
Tang, Ruiming
Yu, Yong
Zhang, Weinan
author_facet Zhu, Jiachen
Zhu, Menghui
Rui, Renting
Shan, Rong
Zheng, Congmin
Chen, Bo
Xi, Yunjia
Lin, Jianghao
Liu, Weiwen
Tang, Ruiming
Yu, Yong
Zhang, Weinan
contents The advent of large language models (LLMs), such as GPT, Gemini, and DeepSeek, has significantly advanced natural language processing, giving rise to sophisticated chatbots capable of diverse language-related tasks. The transition from these traditional LLM chatbots to more advanced AI agents represents a pivotal evolutionary step. However, existing evaluation frameworks often blur the distinctions between LLM chatbots and AI agents, leading to confusion among researchers selecting appropriate benchmarks. To bridge this gap, this paper introduces a systematic analysis of current evaluation approaches, grounded in an evolutionary perspective. We provide a detailed analytical framework that clearly differentiates AI agents from LLM chatbots along five key aspects: complex environment, multi-source instructor, dynamic feedback, multi-modal perception, and advanced capability. Further, we categorize existing evaluation benchmarks based on external environments driving forces, and resulting advanced internal capabilities. For each category, we delineate relevant evaluation attributes, presented comprehensively in practical reference tables. Finally, we synthesize current trends and outline future evaluation methodologies through four critical lenses: environment, agent, evaluator, and metrics. Our findings offer actionable guidance for researchers, facilitating the informed selection and application of benchmarks in AI agent evaluation, thus fostering continued advancement in this rapidly evolving research domain.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11102
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
Zhu, Jiachen
Zhu, Menghui
Rui, Renting
Shan, Rong
Zheng, Congmin
Chen, Bo
Xi, Yunjia
Lin, Jianghao
Liu, Weiwen
Tang, Ruiming
Yu, Yong
Zhang, Weinan
Computation and Language
Artificial Intelligence
The advent of large language models (LLMs), such as GPT, Gemini, and DeepSeek, has significantly advanced natural language processing, giving rise to sophisticated chatbots capable of diverse language-related tasks. The transition from these traditional LLM chatbots to more advanced AI agents represents a pivotal evolutionary step. However, existing evaluation frameworks often blur the distinctions between LLM chatbots and AI agents, leading to confusion among researchers selecting appropriate benchmarks. To bridge this gap, this paper introduces a systematic analysis of current evaluation approaches, grounded in an evolutionary perspective. We provide a detailed analytical framework that clearly differentiates AI agents from LLM chatbots along five key aspects: complex environment, multi-source instructor, dynamic feedback, multi-modal perception, and advanced capability. Further, we categorize existing evaluation benchmarks based on external environments driving forces, and resulting advanced internal capabilities. For each category, we delineate relevant evaluation attributes, presented comprehensively in practical reference tables. Finally, we synthesize current trends and outline future evaluation methodologies through four critical lenses: environment, agent, evaluator, and metrics. Our findings offer actionable guidance for researchers, facilitating the informed selection and application of benchmarks in AI agent evaluation, thus fostering continued advancement in this rapidly evolving research domain.
title Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.11102