WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Styles, Olly, Miller, Sam, Cerda-Mardini, Patricio, Guha, Tanaya, Sanchez, Victor, Vidgen, Bertie
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917740364693504
author Styles, Olly
Miller, Sam
Cerda-Mardini, Patricio
Guha, Tanaya
Sanchez, Victor
Vidgen, Bertie
author_facet Styles, Olly
Miller, Sam
Cerda-Mardini, Patricio
Guha, Tanaya
Sanchez, Victor
Vidgen, Bertie
contents We introduce WorkBench: a benchmark dataset for evaluating agents' ability to execute tasks in a workplace setting. WorkBench contains a sandbox environment with five databases, 26 tools, and 690 tasks. These tasks represent common business activities, such as sending emails and scheduling meetings. The tasks in WorkBench are challenging as they require planning, tool selection, and often multiple actions. If a task has been successfully executed, one (or more) of the database values may change. The correct outcome for each task is unique and unambiguous, which allows for robust, automated evaluation. We call this key contribution outcome-centric evaluation. We evaluate five existing ReAct agents on WorkBench, finding they successfully complete as few as 3% of tasks (Llama2-70B), and just 43% for the best-performing (GPT-4). We further find that agents' errors can result in the wrong action being taken, such as an email being sent to the wrong person. WorkBench reveals weaknesses in agents' ability to undertake common business activities, raising questions about their use in high-stakes workplace settings. WorkBench is publicly available as a free resource at https://github.com/olly-styles/WorkBench.
format Preprint
id arxiv_https___arxiv_org_abs_2405_00823
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
Styles, Olly
Miller, Sam
Cerda-Mardini, Patricio
Guha, Tanaya
Sanchez, Victor
Vidgen, Bertie
Computation and Language
Artificial Intelligence
Multiagent Systems
We introduce WorkBench: a benchmark dataset for evaluating agents' ability to execute tasks in a workplace setting. WorkBench contains a sandbox environment with five databases, 26 tools, and 690 tasks. These tasks represent common business activities, such as sending emails and scheduling meetings. The tasks in WorkBench are challenging as they require planning, tool selection, and often multiple actions. If a task has been successfully executed, one (or more) of the database values may change. The correct outcome for each task is unique and unambiguous, which allows for robust, automated evaluation. We call this key contribution outcome-centric evaluation. We evaluate five existing ReAct agents on WorkBench, finding they successfully complete as few as 3% of tasks (Llama2-70B), and just 43% for the best-performing (GPT-4). We further find that agents' errors can result in the wrong action being taken, such as an email being sent to the wrong person. WorkBench reveals weaknesses in agents' ability to undertake common business activities, raising questions about their use in high-stakes workplace settings. WorkBench is publicly available as a free resource at https://github.com/olly-styles/WorkBench.
title WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
topic Computation and Language
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2405.00823