InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Xueyu, Zhao, Ziyu, Wei, Shuang, Chai, Ziwei, Ma, Qianli, Wang, Guoyin, Wang, Xuwu, Su, Jing, Xu, Jingjing, Zhu, Ming, Cheng, Yao, Yuan, Jianbo, Li, Jiwei, Kuang, Kun, Yang, Yang, Yang, Hongxia, Wu, Fei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913259814125568
author Hu, Xueyu
Zhao, Ziyu
Wei, Shuang
Chai, Ziwei
Ma, Qianli
Wang, Guoyin
Wang, Xuwu
Su, Jing
Xu, Jingjing
Zhu, Ming
Cheng, Yao
Yuan, Jianbo
Li, Jiwei
Kuang, Kun
Yang, Yang
Yang, Hongxia
Wu, Fei
author_facet Hu, Xueyu
Zhao, Ziyu
Wei, Shuang
Chai, Ziwei
Ma, Qianli
Wang, Guoyin
Wang, Xuwu
Su, Jing
Xu, Jingjing
Zhu, Ming
Cheng, Yao
Yuan, Jianbo
Li, Jiwei
Kuang, Kun
Yang, Yang
Yang, Hongxia
Wu, Fei
contents In this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. These tasks require agents to end-to-end solving complex tasks by interacting with an execution environment. This benchmark contains DAEval, a dataset consisting of 257 data analysis questions derived from 52 CSV files, and an agent framework which incorporates LLMs to serve as data analysis agents for both serving and evaluation. Since data analysis questions are often open-ended and hard to evaluate without human supervision, we adopt a format-prompting technique to convert each question into a closed-form format so that they can be automatically evaluated. Our extensive benchmarking of 34 LLMs uncovers the current challenges encountered in data analysis tasks. In addition, building on top of our agent framework, we develop a specialized agent, DAAgent, which surpasses GPT-3.5 by 3.9% on DABench. Evaluation datasets and toolkits for InfiAgent-DABench are released at https://github.com/InfiAgent/InfiAgent .
format Preprint
id arxiv_https___arxiv_org_abs_2401_05507
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
Hu, Xueyu
Zhao, Ziyu
Wei, Shuang
Chai, Ziwei
Ma, Qianli
Wang, Guoyin
Wang, Xuwu
Su, Jing
Xu, Jingjing
Zhu, Ming
Cheng, Yao
Yuan, Jianbo
Li, Jiwei
Kuang, Kun
Yang, Yang
Yang, Hongxia
Wu, Fei
Computation and Language
Artificial Intelligence
In this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. These tasks require agents to end-to-end solving complex tasks by interacting with an execution environment. This benchmark contains DAEval, a dataset consisting of 257 data analysis questions derived from 52 CSV files, and an agent framework which incorporates LLMs to serve as data analysis agents for both serving and evaluation. Since data analysis questions are often open-ended and hard to evaluate without human supervision, we adopt a format-prompting technique to convert each question into a closed-form format so that they can be automatically evaluated. Our extensive benchmarking of 34 LLMs uncovers the current challenges encountered in data analysis tasks. In addition, building on top of our agent framework, we develop a specialized agent, DAAgent, which surpasses GPT-3.5 by 3.9% on DABench. Evaluation datasets and toolkits for InfiAgent-DABench are released at https://github.com/InfiAgent/InfiAgent .
title InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.05507