Benchmarking Agentic Workflow Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiao, Shuofei, Fang, Runnan, Qiu, Zhisong, Wang, Xiaobin, Zhang, Ningyu, Jiang, Yong, Xie, Pengjun, Huang, Fei, Chen, Huajun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913702893060096
author Qiao, Shuofei
Fang, Runnan
Qiu, Zhisong
Wang, Xiaobin
Zhang, Ningyu
Jiang, Yong
Xie, Pengjun
Huang, Fei
Chen, Huajun
author_facet Qiao, Shuofei
Fang, Runnan
Qiu, Zhisong
Wang, Xiaobin
Zhang, Ningyu
Jiang, Yong
Xie, Pengjun
Huang, Fei
Chen, Huajun
contents Large Language Models (LLMs), with their exceptional ability to handle a wide range of tasks, have driven significant advancements in tackling reasoning and planning tasks, wherein decomposing complex problems into executable workflows is a crucial step in this process. Existing workflow evaluation frameworks either focus solely on holistic performance or suffer from limitations such as restricted scenario coverage, simplistic workflow structures, and lax evaluation standards. To this end, we introduce WorfBench, a unified workflow generation benchmark with multi-faceted scenarios and intricate graph workflow structures. Additionally, we present WorfEval, a systemic evaluation protocol utilizing subsequence and subgraph matching algorithms to accurately quantify the LLM agent's workflow generation capabilities. Through comprehensive evaluations across different types of LLMs, we discover distinct gaps between the sequence planning capabilities and graph planning capabilities of LLM agents, with even GPT-4 exhibiting a gap of around 15%. We also train two open-source models and evaluate their generalization abilities on held-out tasks. Furthermore, we observe that the generated workflows can enhance downstream tasks, enabling them to achieve superior performance with less time during inference. Code and dataset are available at https://github.com/zjunlp/WorfBench.
format Preprint
id arxiv_https___arxiv_org_abs_2410_07869
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking Agentic Workflow Generation
Qiao, Shuofei
Fang, Runnan
Qiu, Zhisong
Wang, Xiaobin
Zhang, Ningyu
Jiang, Yong
Xie, Pengjun
Huang, Fei
Chen, Huajun
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Multiagent Systems
Large Language Models (LLMs), with their exceptional ability to handle a wide range of tasks, have driven significant advancements in tackling reasoning and planning tasks, wherein decomposing complex problems into executable workflows is a crucial step in this process. Existing workflow evaluation frameworks either focus solely on holistic performance or suffer from limitations such as restricted scenario coverage, simplistic workflow structures, and lax evaluation standards. To this end, we introduce WorfBench, a unified workflow generation benchmark with multi-faceted scenarios and intricate graph workflow structures. Additionally, we present WorfEval, a systemic evaluation protocol utilizing subsequence and subgraph matching algorithms to accurately quantify the LLM agent's workflow generation capabilities. Through comprehensive evaluations across different types of LLMs, we discover distinct gaps between the sequence planning capabilities and graph planning capabilities of LLM agents, with even GPT-4 exhibiting a gap of around 15%. We also train two open-source models and evaluate their generalization abilities on held-out tasks. Furthermore, we observe that the generated workflows can enhance downstream tasks, enabling them to achieve superior performance with less time during inference. Code and dataset are available at https://github.com/zjunlp/WorfBench.
title Benchmarking Agentic Workflow Generation
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2410.07869