$C^3$-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
Fuente:
arXiv
Guardado en:
| Autores principales: | Yu, Peijie, Yang, Yifan, Li, Jinjian, Zhang, Zelong, Wang, Haorui, Feng, Xiao, Zhang, Feng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
por: Yu, Peijie, et al.
Publicado: (2025)
por: Yu, Peijie, et al.
Publicado: (2025)
Benchmarking LLM Tool-Use in the Wild
por: Yu, Peijie, et al.
Publicado: (2026)
por: Yu, Peijie, et al.
Publicado: (2026)
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
por: Shang, Yu, et al.
Publicado: (2025)
por: Shang, Yu, et al.
Publicado: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
por: Liu, Wenrui, et al.
Publicado: (2025)
por: Liu, Wenrui, et al.
Publicado: (2025)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
por: He, Wei, et al.
Publicado: (2025)
por: He, Wei, et al.
Publicado: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
por: Xiao, Jianfei, et al.
Publicado: (2026)
por: Xiao, Jianfei, et al.
Publicado: (2026)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
por: Cui, Fan, et al.
Publicado: (2026)
por: Cui, Fan, et al.
Publicado: (2026)
PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?
por: Hua, Dongdong, et al.
Publicado: (2026)
por: Hua, Dongdong, et al.
Publicado: (2026)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
por: Zhou, Yifan, et al.
Publicado: (2026)
por: Zhou, Yifan, et al.
Publicado: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
por: Long, Xiang, et al.
Publicado: (2026)
por: Long, Xiang, et al.
Publicado: (2026)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
por: Lu, Jiaxuan, et al.
Publicado: (2026)
por: Lu, Jiaxuan, et al.
Publicado: (2026)
TrustAgent: Towards Safe and Trustworthy LLM-based Agents
por: Hua, Wenyue, et al.
Publicado: (2024)
por: Hua, Wenyue, et al.
Publicado: (2024)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
por: Zhu, Jie, et al.
Publicado: (2026)
por: Zhu, Jie, et al.
Publicado: (2026)
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
por: Li, Zelong, et al.
Publicado: (2024)
por: Li, Zelong, et al.
Publicado: (2024)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
por: Zheng, Junhao, et al.
Publicado: (2025)
por: Zheng, Junhao, et al.
Publicado: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
por: Zhang, Hanrong, et al.
Publicado: (2024)
por: Zhang, Hanrong, et al.
Publicado: (2024)
TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System
por: Fan, Wenzhe, et al.
Publicado: (2026)
por: Fan, Wenzhe, et al.
Publicado: (2026)
Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration Framework
por: Xu, Jiasheng, et al.
Publicado: (2025)
por: Xu, Jiasheng, et al.
Publicado: (2025)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
por: Chen, Wanyi, et al.
Publicado: (2026)
por: Chen, Wanyi, et al.
Publicado: (2026)
LLM-based Multi-Agent Systems: Techniques and Business Perspectives
por: Yang, Yingxuan, et al.
Publicado: (2024)
por: Yang, Yingxuan, et al.
Publicado: (2024)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
por: Ni, Ziyi, et al.
Publicado: (2025)
por: Ni, Ziyi, et al.
Publicado: (2025)
Multi-Agent Evolve: LLM Self-Improve through Co-evolution
por: Chen, Yixing, et al.
Publicado: (2025)
por: Chen, Yixing, et al.
Publicado: (2025)
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
por: Liu, Ruoqi, et al.
Publicado: (2026)
por: Liu, Ruoqi, et al.
Publicado: (2026)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
por: Dong, Haonan, et al.
Publicado: (2026)
por: Dong, Haonan, et al.
Publicado: (2026)
CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments
por: Yu, Yi, et al.
Publicado: (2026)
por: Yu, Yi, et al.
Publicado: (2026)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
por: Feng, Yunhao, et al.
Publicado: (2026)
por: Feng, Yunhao, et al.
Publicado: (2026)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
por: Chen, Haorui, et al.
Publicado: (2025)
por: Chen, Haorui, et al.
Publicado: (2025)
WorkstreamBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
por: Yen, Thomson, et al.
Publicado: (2026)
por: Yen, Thomson, et al.
Publicado: (2026)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
por: Schmidt, Jan-Philipp
Publicado: (2026)
por: Schmidt, Jan-Philipp
Publicado: (2026)
AgentBench: Evaluating LLMs as Agents
por: Liu, Xiao, et al.
Publicado: (2023)
por: Liu, Xiao, et al.
Publicado: (2023)
AIOS Compiler: LLM as Interpreter for Natural Language Programming and Flow Programming of AI Agents
por: Xu, Shuyuan, et al.
Publicado: (2024)
por: Xu, Shuyuan, et al.
Publicado: (2024)
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
por: Shen, Haiyang, et al.
Publicado: (2024)
por: Shen, Haiyang, et al.
Publicado: (2024)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
por: Liu, Zhou, et al.
Publicado: (2025)
por: Liu, Zhou, et al.
Publicado: (2025)
Cross-Task Experiential Learning on LLM-based Multi-Agent Collaboration
por: Li, Yilong, et al.
Publicado: (2025)
por: Li, Yilong, et al.
Publicado: (2025)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
por: Deng, Shihan, et al.
Publicado: (2024)
por: Deng, Shihan, et al.
Publicado: (2024)
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
por: Liu, Yi, et al.
Publicado: (2026)
por: Liu, Yi, et al.
Publicado: (2026)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
por: Shen, Yuanzhe, et al.
Publicado: (2026)
por: Shen, Yuanzhe, et al.
Publicado: (2026)
InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
por: Du, Yaxin, et al.
Publicado: (2025)
por: Du, Yaxin, et al.
Publicado: (2025)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
por: Zhang, Yuzhe, et al.
Publicado: (2026)
por: Zhang, Yuzhe, et al.
Publicado: (2026)
Efficient Prompting for LLM-based Generative Internet of Things
por: Xiao, Bin, et al.
Publicado: (2024)
por: Xiao, Bin, et al.
Publicado: (2024)
Ejemplares similares
-
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
por: Yu, Peijie, et al.
Publicado: (2025) -
Benchmarking LLM Tool-Use in the Wild
por: Yu, Peijie, et al.
Publicado: (2026) -
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
por: Shang, Yu, et al.
Publicado: (2025) -
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
por: Liu, Wenrui, et al.
Publicado: (2025) -
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
por: He, Wei, et al.
Publicado: (2025)