AgentBench: Evaluating LLMs as Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiao, Yu, Hao, Zhang, Hanchen, Xu, Yifan, Lei, Xuanyu, Lai, Hanyu, Gu, Yu, Ding, Hangliang, Men, Kaiwen, Yang, Kejuan, Zhang, Shudan, Deng, Xiang, Zeng, Aohan, Du, Zhengxiao, Zhang, Chenhui, Shen, Sheng, Zhang, Tianjun, Su, Yu, Sun, Huan, Huang, Minlie, Dong, Yuxiao, Tang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
AndroidGen: Building an Android Language Agent under Data Scarcity
by: Lai, Hanyu, et al.
Published: (2025)
by: Lai, Hanyu, et al.
Published: (2025)
AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025)
by: Lee, Gyubok, et al.
Published: (2025)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
by: Xu, Yifan, et al.
Published: (2025)
by: Xu, Yifan, et al.
Published: (2025)
Understanding Emergent Abilities of Language Models from the Loss Perspective
by: Du, Zhengxiao, et al.
Published: (2024)
by: Du, Zhengxiao, et al.
Published: (2024)
AutoWebGLM: A Large Language Model-based Web Navigating Agent
by: Lai, Hanyu, et al.
Published: (2024)
by: Lai, Hanyu, et al.
Published: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
by: Qu, Yuxiao, et al.
Published: (2024)
by: Qu, Yuxiao, et al.
Published: (2024)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
by: Shang, Yu, et al.
Published: (2025)
by: Shang, Yu, et al.
Published: (2025)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
by: Lai, Hanyu, et al.
Published: (2025)
by: Lai, Hanyu, et al.
Published: (2025)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
by: Hou, Zhenyu, et al.
Published: (2024)
by: Hou, Zhenyu, et al.
Published: (2024)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
by: Zhang, Shudan, et al.
Published: (2024)
by: Zhang, Shudan, et al.
Published: (2024)
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
by: Gu, Yu, et al.
Published: (2024)
by: Gu, Yu, et al.
Published: (2024)
UniShield: An Adaptive Multi-Agent Framework for Unified Forgery Image Detection and Localization
by: Huang, Qing, et al.
Published: (2025)
by: Huang, Qing, et al.
Published: (2025)
Feature‐Driven DEM Generation With Enhanced Detail Preservation and Noise Mitigation Using Conditional GANs
by: Chenhui Wu, et al.
Published: (2026)
by: Chenhui Wu, et al.
Published: (2026)
$C^3$-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
by: Yu, Peijie, et al.
Published: (2025)
by: Yu, Peijie, et al.
Published: (2025)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
by: Hou, Zhenyu, et al.
Published: (2024)
by: Hou, Zhenyu, et al.
Published: (2024)
APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decoding
by: Liu, Mingdao, et al.
Published: (2024)
by: Liu, Mingdao, et al.
Published: (2024)
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
by: Yu, Peijie, et al.
Published: (2025)
by: Yu, Peijie, et al.
Published: (2025)
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
by: Bai, Yushi, et al.
Published: (2023)
by: Bai, Yushi, et al.
Published: (2023)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025)
by: Jiang, Yilin, et al.
Published: (2025)
AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
by: Zhang, Hanchen, et al.
Published: (2025)
by: Zhang, Hanchen, et al.
Published: (2025)
Exploration of cardiopulmonary resuscitation teamwork training for maternal cardiac arrest using the SimMan intelligent simulation platform: A simulation teaching study
by: Ruirui Zhang, et al.
Published: (2024)
by: Ruirui Zhang, et al.
Published: (2024)
AutoGLM: Autonomous Foundation Agents for GUIs
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
$\texttt{PatentAgent}$: Intelligent Agent for Automated Pharmaceutical Patent Analysis
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
DCA-Bench: A Benchmark for Dataset Curation Agents
by: Huang, Benhao, et al.
Published: (2024)
by: Huang, Benhao, et al.
Published: (2024)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
AgentAsk: Multi-Agent Systems Need to Ask
by: Lin, Bohan, et al.
Published: (2025)
by: Lin, Bohan, et al.
Published: (2025)
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
Constraint-Guided Multi-Agent Decompilation for Executable Binary Recovery
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
EXG: Self-Evolving Agents with Experience Graphs
by: Jin, Yuxin, et al.
Published: (2026)
by: Jin, Yuxin, et al.
Published: (2026)
Towards Exception Safety Code Generation with Intermediate Representation Agents Framework
by: Zhang, Xuanming, et al.
Published: (2024)
by: Zhang, Xuanming, et al.
Published: (2024)
GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents
by: Wang, Yanxi, et al.
Published: (2026)
by: Wang, Yanxi, et al.
Published: (2026)
MapCraft: Dissecting and Designing Custom Geo-Infographics
by: Zhang, Xinyuan, et al.
Published: (2024)
by: Zhang, Xinyuan, et al.
Published: (2024)
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
How do Visual Attributes Influence Web Agents? A Comprehensive Evaluation of User Interface Design Factors
by: Yu, Kuai, et al.
Published: (2026)
by: Yu, Kuai, et al.
Published: (2026)
Similar Items
-
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
by: Liu, Xiao, et al.
Published: (2024) -
AndroidGen: Building an Android Language Agent under Data Scarcity
by: Lai, Hanyu, et al.
Published: (2025) -
AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents
by: Xu, Yifan, et al.
Published: (2024) -
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025) -
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)