WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
Fuente:
arXiv
Saved in:
| Main Authors: | Styles, Olly, Miller, Sam, Cerda-Mardini, Patricio, Guha, Tanaya, Sanchez, Victor, Vidgen, Bertie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
by: Jiang, Yixing, et al.
Published: (2025)
by: Jiang, Yixing, et al.
Published: (2025)
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
by: Zhong, Shanshan, et al.
Published: (2026)
by: Zhong, Shanshan, et al.
Published: (2026)
Scaling Lifelong Multi-Agent Path Finding to More Realistic Settings: Research Challenges and Opportunities
by: Jiang, He, et al.
Published: (2024)
by: Jiang, He, et al.
Published: (2024)
Achieving Realistic Cyclist Behavior in SUMO using the SimRa Dataset
by: Karakaya, Ahmet-Serdar, et al.
Published: (2023)
by: Karakaya, Ahmet-Serdar, et al.
Published: (2023)
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
BenchMARL: Benchmarking Multi-Agent Reinforcement Learning
by: Bettini, Matteo, et al.
Published: (2023)
by: Bettini, Matteo, et al.
Published: (2023)
AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
by: Patel, Dhaval, et al.
Published: (2025)
by: Patel, Dhaval, et al.
Published: (2025)
SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
by: Hamblin, Grant, et al.
Published: (2026)
by: Hamblin, Grant, et al.
Published: (2026)
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
by: Xu, Shihao, et al.
Published: (2026)
by: Xu, Shihao, et al.
Published: (2026)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
by: Siegel, Zachary S., et al.
Published: (2024)
by: Siegel, Zachary S., et al.
Published: (2024)
A Multi-Agent Retrieval-Augmented Framework for Work-in-Progress Predictio
by: Bibalan, Yousef Mehrdad, et al.
Published: (2025)
by: Bibalan, Yousef Mehrdad, et al.
Published: (2025)
Policy Optimization in Multi-Agent Settings under Partially Observable Environments
by: Zhaikhan, Ainur, et al.
Published: (2025)
by: Zhaikhan, Ainur, et al.
Published: (2025)
HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
by: Yang, Yiqian, et al.
Published: (2025)
by: Yang, Yiqian, et al.
Published: (2025)
Is AI Ready for Multimodal Hate Speech Detection? A Comprehensive Dataset and Benchmark Evaluation
by: Xing, Rui, et al.
Published: (2026)
by: Xing, Rui, et al.
Published: (2026)
FinDeepForecast: A Live Multi-Agent System for Benchmarking Deep Research Agents in Financial Forecasting
by: Li, Xiangyu, et al.
Published: (2026)
by: Li, Xiangyu, et al.
Published: (2026)
Task Capability Improvement Algorithm for Collaborative Manipulators
by: Patra, Keshab, et al.
Published: (2026)
by: Patra, Keshab, et al.
Published: (2026)
Agents on the Bench: Large Language Model Based Multi Agent Framework for Trustworthy Digital Justice
by: Jiang, Cong, et al.
Published: (2024)
by: Jiang, Cong, et al.
Published: (2024)
Crisis-Bench: Benchmarking Strategic Ambiguity and Reputation Management in Large Language Models
by: Lin, Cooper, et al.
Published: (2026)
by: Lin, Cooper, et al.
Published: (2026)
Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
by: Chen, Mengzhuo, et al.
Published: (2026)
by: Chen, Mengzhuo, et al.
Published: (2026)
FedGUI: Benchmarking Federated GUI Agents across Heterogeneous Platforms, Devices, and Operating Systems
by: Wang, Wenhao, et al.
Published: (2026)
by: Wang, Wenhao, et al.
Published: (2026)
MARLadona -- Towards Cooperative Team Play Using Multi-Agent Reinforcement Learning
by: Li, Zichong, et al.
Published: (2024)
by: Li, Zichong, et al.
Published: (2024)
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
by: Zou, Chelsea, et al.
Published: (2026)
by: Zou, Chelsea, et al.
Published: (2026)
FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic Topics
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2026)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2026)
Toward a Safe Internet of Agents
by: Wibowo, Juan A., et al.
Published: (2025)
by: Wibowo, Juan A., et al.
Published: (2025)
Designing for Accountable Agents: a Viewpoint
by: Cranefield, Stephen, et al.
Published: (2026)
by: Cranefield, Stephen, et al.
Published: (2026)
Improvisational Games as a Benchmark for Social Intelligence of AI Agents: The Case of Connections
by: Parikh, Gaurav Rajesh, et al.
Published: (2026)
by: Parikh, Gaurav Rajesh, et al.
Published: (2026)
TeamFusion: Supporting Open-ended Teamwork with Multi-Agent Systems
by: Liu, Jiale, et al.
Published: (2026)
by: Liu, Jiale, et al.
Published: (2026)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
by: Saxena, Siddhant, et al.
Published: (2026)
by: Saxena, Siddhant, et al.
Published: (2026)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
by: Zhang, Yuzhe, et al.
Published: (2026)
by: Zhang, Yuzhe, et al.
Published: (2026)
Set-Rationalizable Choice and Self-Stability
by: Brandt, Felix, et al.
Published: (2009)
by: Brandt, Felix, et al.
Published: (2009)
Multi-Agent DRL for V2X Resource Allocation: Disentangling Challenges and Benchmarking Solutions
by: Wang, Siyuan, et al.
Published: (2026)
by: Wang, Siyuan, et al.
Published: (2026)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
by: Feng, Yukang, et al.
Published: (2026)
by: Feng, Yukang, et al.
Published: (2026)
Retrieval-Augmented Multi-Agent System for Rapid Statement of Work Generation
by: Suravarjhula, Amulya, et al.
Published: (2025)
by: Suravarjhula, Amulya, et al.
Published: (2025)
Multi-Agent Reinforcement Learning for Autonomous Multi-Satellite Earth Observation: A Realistic Case Study
by: Hady, Mohamad A., et al.
Published: (2025)
by: Hady, Mohamad A., et al.
Published: (2025)
Multi-Agent Craftax: Benchmarking Open-Ended Multi-Agent Reinforcement Learning at the Hyperscale
by: Omari, Bassel Al, et al.
Published: (2025)
by: Omari, Bassel Al, et al.
Published: (2025)
From Grounding to Planning: Benchmarking Bottlenecks in Web Agents
by: Shlomov, Segev, et al.
Published: (2024)
by: Shlomov, Segev, et al.
Published: (2024)
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering
by: Guan, Shaowei, et al.
Published: (2026)
by: Guan, Shaowei, et al.
Published: (2026)
CREW-WILDFIRE: Benchmarking Agentic Multi-Agent Collaborations at Scale
by: Hyun, Jonathan, et al.
Published: (2025)
by: Hyun, Jonathan, et al.
Published: (2025)
Similar Items
-
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
by: Jiang, Yixing, et al.
Published: (2025) -
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
by: Zhong, Shanshan, et al.
Published: (2026) -
Scaling Lifelong Multi-Agent Path Finding to More Realistic Settings: Research Challenges and Opportunities
by: Jiang, He, et al.
Published: (2024) -
Achieving Realistic Cyclist Behavior in SUMO using the SimRa Dataset
by: Karakaya, Ahmet-Serdar, et al.
Published: (2023) -
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
by: Wu, Bin, et al.
Published: (2026)