InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913259814125568 |
|---|---|
| author | Hu, Xueyu Zhao, Ziyu Wei, Shuang Chai, Ziwei Ma, Qianli Wang, Guoyin Wang, Xuwu Su, Jing Xu, Jingjing Zhu, Ming Cheng, Yao Yuan, Jianbo Li, Jiwei Kuang, Kun Yang, Yang Yang, Hongxia Wu, Fei |
| author_facet | Hu, Xueyu Zhao, Ziyu Wei, Shuang Chai, Ziwei Ma, Qianli Wang, Guoyin Wang, Xuwu Su, Jing Xu, Jingjing Zhu, Ming Cheng, Yao Yuan, Jianbo Li, Jiwei Kuang, Kun Yang, Yang Yang, Hongxia Wu, Fei |
| contents | In this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. These tasks require agents to end-to-end solving complex tasks by interacting with an execution environment. This benchmark contains DAEval, a dataset consisting of 257 data analysis questions derived from 52 CSV files, and an agent framework which incorporates LLMs to serve as data analysis agents for both serving and evaluation. Since data analysis questions are often open-ended and hard to evaluate without human supervision, we adopt a format-prompting technique to convert each question into a closed-form format so that they can be automatically evaluated. Our extensive benchmarking of 34 LLMs uncovers the current challenges encountered in data analysis tasks. In addition, building on top of our agent framework, we develop a specialized agent, DAAgent, which surpasses GPT-3.5 by 3.9% on DABench. Evaluation datasets and toolkits for InfiAgent-DABench are released at https://github.com/InfiAgent/InfiAgent . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_05507 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks Hu, Xueyu Zhao, Ziyu Wei, Shuang Chai, Ziwei Ma, Qianli Wang, Guoyin Wang, Xuwu Su, Jing Xu, Jingjing Zhu, Ming Cheng, Yao Yuan, Jianbo Li, Jiwei Kuang, Kun Yang, Yang Yang, Hongxia Wu, Fei Computation and Language Artificial Intelligence In this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. These tasks require agents to end-to-end solving complex tasks by interacting with an execution environment. This benchmark contains DAEval, a dataset consisting of 257 data analysis questions derived from 52 CSV files, and an agent framework which incorporates LLMs to serve as data analysis agents for both serving and evaluation. Since data analysis questions are often open-ended and hard to evaluate without human supervision, we adopt a format-prompting technique to convert each question into a closed-form format so that they can be automatically evaluated. Our extensive benchmarking of 34 LLMs uncovers the current challenges encountered in data analysis tasks. In addition, building on top of our agent framework, we develop a specialized agent, DAAgent, which surpasses GPT-3.5 by 3.9% on DABench. Evaluation datasets and toolkits for InfiAgent-DABench are released at https://github.com/InfiAgent/InfiAgent . |
| title | InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2401.05507 |