TaskBench: Benchmarking Large Language Models for Task Automation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shen, Yongliang, Song, Kaitao, Tan, Xu, Zhang, Wenqi, Ren, Kan, Yuan, Siyu, Lu, Weiming, Li, Dongsheng, Zhuang, Yueting
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915001272369152
author Shen, Yongliang
Song, Kaitao
Tan, Xu
Zhang, Wenqi
Ren, Kan
Yuan, Siyu
Lu, Weiming
Li, Dongsheng
Zhuang, Yueting
author_facet Shen, Yongliang
Song, Kaitao
Tan, Xu
Zhang, Wenqi
Ren, Kan
Yuan, Siyu
Lu, Weiming
Li, Dongsheng
Zhuang, Yueting
contents In recent years, the remarkable progress of large language models (LLMs) has sparked interest in task automation, which involves decomposing complex tasks described by user instructions into sub-tasks and invoking external tools to execute them, playing a central role in autonomous agents. However, there is a lack of systematic and standardized benchmarks to promote the development of LLMs in task automation. To address this, we introduce TaskBench, a comprehensive framework to evaluate the capability of LLMs in task automation. Specifically, task automation can be divided into three critical stages: task decomposition, tool selection, and parameter prediction. To tackle the complexities inherent in these stages, we introduce the concept of Tool Graph to represent decomposed tasks and adopt a back-instruct method to generate high-quality user instructions. We propose TaskEval, a multi-faceted evaluation methodology that assesses LLM performance across these three stages. Our approach combines automated construction with rigorous human verification, ensuring high consistency with human evaluation. Experimental results demonstrate that TaskBench effectively reflects the capabilities of various LLMs in task automation. It provides insights into model performance across different task complexities and domains, pushing the boundaries of what current models can achieve. TaskBench offers a scalable, adaptable, and reliable benchmark for advancing LLM-based autonomous agents.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18760
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TaskBench: Benchmarking Large Language Models for Task Automation
Shen, Yongliang
Song, Kaitao
Tan, Xu
Zhang, Wenqi
Ren, Kan
Yuan, Siyu
Lu, Weiming
Li, Dongsheng
Zhuang, Yueting
Computation and Language
Artificial Intelligence
In recent years, the remarkable progress of large language models (LLMs) has sparked interest in task automation, which involves decomposing complex tasks described by user instructions into sub-tasks and invoking external tools to execute them, playing a central role in autonomous agents. However, there is a lack of systematic and standardized benchmarks to promote the development of LLMs in task automation. To address this, we introduce TaskBench, a comprehensive framework to evaluate the capability of LLMs in task automation. Specifically, task automation can be divided into three critical stages: task decomposition, tool selection, and parameter prediction. To tackle the complexities inherent in these stages, we introduce the concept of Tool Graph to represent decomposed tasks and adopt a back-instruct method to generate high-quality user instructions. We propose TaskEval, a multi-faceted evaluation methodology that assesses LLM performance across these three stages. Our approach combines automated construction with rigorous human verification, ensuring high consistency with human evaluation. Experimental results demonstrate that TaskBench effectively reflects the capabilities of various LLMs in task automation. It provides insights into model performance across different task complexities and domains, pushing the boundaries of what current models can achieve. TaskBench offers a scalable, adaptable, and reliable benchmark for advancing LLM-based autonomous agents.
title TaskBench: Benchmarking Large Language Models for Task Automation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.18760