LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Hao, Wang, Huan, Gu, Jinjie, Wang, Wenjie, Zhuang, Chenyi, Bian, Sikang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
von: He, Kaiwen, et al.
Veröffentlicht: (2025)
von: He, Kaiwen, et al.
Veröffentlicht: (2025)
Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld
von: Xie, Zhitian, et al.
Veröffentlicht: (2025)
von: Xie, Zhitian, et al.
Veröffentlicht: (2025)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)
CharPoet: A Chinese Classical Poetry Generation System Based on Token-free LLM
von: Yu, Chengyue, et al.
Veröffentlicht: (2024)
von: Yu, Chengyue, et al.
Veröffentlicht: (2024)
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
von: He, Hang, et al.
Veröffentlicht: (2025)
von: He, Hang, et al.
Veröffentlicht: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
von: Long, Xiang, et al.
Veröffentlicht: (2026)
von: Long, Xiang, et al.
Veröffentlicht: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026)
von: Yang, Jie, et al.
Veröffentlicht: (2026)
Don't Just Fine-tune the Agent, Tune the Environment
von: Lu, Siyuan, et al.
Veröffentlicht: (2025)
von: Lu, Siyuan, et al.
Veröffentlicht: (2025)
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
von: Li, Keyu, et al.
Veröffentlicht: (2026)
von: Li, Keyu, et al.
Veröffentlicht: (2026)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
von: Li, Chenxin, et al.
Veröffentlicht: (2026)
von: Li, Chenxin, et al.
Veröffentlicht: (2026)
LiveClin: A Live Clinical Benchmark without Leakage
von: Wang, Xidong, et al.
Veröffentlicht: (2026)
von: Wang, Xidong, et al.
Veröffentlicht: (2026)
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
von: Chen, Jingxuan, et al.
Veröffentlicht: (2024)
von: Chen, Jingxuan, et al.
Veröffentlicht: (2024)
Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy
von: Zhao, Yao, et al.
Veröffentlicht: (2023)
von: Zhao, Yao, et al.
Veröffentlicht: (2023)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
von: Song, Zhiheng, et al.
Veröffentlicht: (2026)
von: Song, Zhiheng, et al.
Veröffentlicht: (2026)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment
von: Zheng, Jingnan, et al.
Veröffentlicht: (2026)
von: Zheng, Jingnan, et al.
Veröffentlicht: (2026)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
von: Zhang, Qiaohong, et al.
Veröffentlicht: (2026)
von: Zhang, Qiaohong, et al.
Veröffentlicht: (2026)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
von: Cui, Fan, et al.
Veröffentlicht: (2026)
von: Cui, Fan, et al.
Veröffentlicht: (2026)
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
von: Chen, Jikai, et al.
Veröffentlicht: (2026)
von: Chen, Jikai, et al.
Veröffentlicht: (2026)
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
von: Chen, Jikai, et al.
Veröffentlicht: (2025)
von: Chen, Jikai, et al.
Veröffentlicht: (2025)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
von: Chen, Wanyi, et al.
Veröffentlicht: (2026)
von: Chen, Wanyi, et al.
Veröffentlicht: (2026)
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
von: Wang, Maolin, et al.
Veröffentlicht: (2023)
von: Wang, Maolin, et al.
Veröffentlicht: (2023)
AWorld: Orchestrating the Training Recipe for Agentic AI
von: Yu, Chengyue, et al.
Veröffentlicht: (2025)
von: Yu, Chengyue, et al.
Veröffentlicht: (2025)
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
von: Li, Xiangyi, et al.
Veröffentlicht: (2026)
von: Li, Xiangyi, et al.
Veröffentlicht: (2026)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
von: Dong, Haonan, et al.
Veröffentlicht: (2026)
von: Dong, Haonan, et al.
Veröffentlicht: (2026)
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
von: Liu, Yibing, et al.
Veröffentlicht: (2026)
von: Liu, Yibing, et al.
Veröffentlicht: (2026)
DNS-Rec: Data-aware Neural Architecture Search for Recommender Systems
von: Zhang, Sheng, et al.
Veröffentlicht: (2024)
von: Zhang, Sheng, et al.
Veröffentlicht: (2024)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
von: Fu, Yuchuan, et al.
Veröffentlicht: (2025)
von: Fu, Yuchuan, et al.
Veröffentlicht: (2025)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
von: Zhou, Qixing, et al.
Veröffentlicht: (2026)
von: Zhou, Qixing, et al.
Veröffentlicht: (2026)
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
von: Tan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Tan, Zhiwen, et al.
Veröffentlicht: (2025)
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
PPU-Bench:Real World Benchmark for Personalized Partial Unlearning in Vision Language Models
von: Guang, Jiahui, et al.
Veröffentlicht: (2026)
von: Guang, Jiahui, et al.
Veröffentlicht: (2026)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
von: He, Wei, et al.
Veröffentlicht: (2025)
von: He, Wei, et al.
Veröffentlicht: (2025)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
NEWSAGENT: Benchmarking Multimodal Agents as Journalists with Real-World Newswriting Tasks
von: Chien, Yen-Che, et al.
Veröffentlicht: (2025)
von: Chien, Yen-Che, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
von: He, Kaiwen, et al.
Veröffentlicht: (2025) -
Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld
von: Xie, Zhitian, et al.
Veröffentlicht: (2025) -
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
von: Gao, Zeyu, et al.
Veröffentlicht: (2025) -
CharPoet: A Chinese Classical Poetry Generation System Based on Token-free LLM
von: Yu, Chengyue, et al.
Veröffentlicht: (2024) -
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
von: He, Hang, et al.
Veröffentlicht: (2025)