Data Interpreter: An LLM Agent For Data Science

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hong, Sirui, Lin, Yizhang, Liu, Bang, Liu, Bangbang, Wu, Binhao, Zhang, Ceyao, Wei, Chenxing, Li, Danyang, Chen, Jiaqi, Zhang, Jiayi, Wang, Jinlin, Zhang, Li, Zhang, Lingyao, Yang, Min, Zhuge, Mingchen, Guo, Taicheng, Zhou, Tuo, Tao, Wei, Tang, Xiangru, Lu, Xiangtao, Zheng, Xiawu, Liang, Xinbing, Fei, Yaying, Cheng, Yuheng, Gou, Zhibin, Xu, Zongze, Wu, Chenglin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917803035983872
author Hong, Sirui
Lin, Yizhang
Liu, Bang
Liu, Bangbang
Wu, Binhao
Zhang, Ceyao
Wei, Chenxing
Li, Danyang
Chen, Jiaqi
Zhang, Jiayi
Wang, Jinlin
Zhang, Li
Zhang, Lingyao
Yang, Min
Zhuge, Mingchen
Guo, Taicheng
Zhou, Tuo
Tao, Wei
Tang, Xiangru
Lu, Xiangtao
Zheng, Xiawu
Liang, Xinbing
Fei, Yaying
Cheng, Yuheng
Gou, Zhibin
Xu, Zongze
Wu, Chenglin
author_facet Hong, Sirui
Lin, Yizhang
Liu, Bang
Liu, Bangbang
Wu, Binhao
Zhang, Ceyao
Wei, Chenxing
Li, Danyang
Chen, Jiaqi
Zhang, Jiayi
Wang, Jinlin
Zhang, Li
Zhang, Lingyao
Yang, Min
Zhuge, Mingchen
Guo, Taicheng
Zhou, Tuo
Tao, Wei
Tang, Xiangru
Lu, Xiangtao
Zheng, Xiawu
Liang, Xinbing
Fei, Yaying
Cheng, Yuheng
Gou, Zhibin
Xu, Zongze
Wu, Chenglin
contents Large Language Model (LLM)-based agents have shown effectiveness across many applications. However, their use in data science scenarios requiring solving long-term interconnected tasks, dynamic data adjustments and domain expertise remains challenging. Previous approaches primarily focus on individual tasks, making it difficult to assess the complete data science workflow. Moreover, they struggle to handle real-time changes in intermediate data and fail to adapt dynamically to evolving task dependencies inherent to data science problems. In this paper, we present Data Interpreter, an LLM-based agent designed to automatically solve various data science problems end-to-end. Our Data Interpreter incorporates two key modules: 1) Hierarchical Graph Modeling, which breaks down complex problems into manageable subproblems, enabling dynamic node generation and graph optimization; and 2) Programmable Node Generation, a technique that refines and verifies each subproblem to iteratively improve code generation results and robustness. Extensive experiments consistently demonstrate the superiority of Data Interpreter. On InfiAgent-DABench, it achieves a 25% performance boost, raising accuracy from 75.9% to 94.9%. For machine learning and open-ended tasks, it improves performance from 88% to 95%, and from 60% to 97%, respectively. Moreover, on the MATH dataset, Data Interpreter achieves remarkable performance with a 26% improvement compared to state-of-the-art baselines. The code is available at https://github.com/geekan/MetaGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18679
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Interpreter: An LLM Agent For Data Science
Hong, Sirui
Lin, Yizhang
Liu, Bang
Liu, Bangbang
Wu, Binhao
Zhang, Ceyao
Wei, Chenxing
Li, Danyang
Chen, Jiaqi
Zhang, Jiayi
Wang, Jinlin
Zhang, Li
Zhang, Lingyao
Yang, Min
Zhuge, Mingchen
Guo, Taicheng
Zhou, Tuo
Tao, Wei
Tang, Xiangru
Lu, Xiangtao
Zheng, Xiawu
Liang, Xinbing
Fei, Yaying
Cheng, Yuheng
Gou, Zhibin
Xu, Zongze
Wu, Chenglin
Artificial Intelligence
Machine Learning
Large Language Model (LLM)-based agents have shown effectiveness across many applications. However, their use in data science scenarios requiring solving long-term interconnected tasks, dynamic data adjustments and domain expertise remains challenging. Previous approaches primarily focus on individual tasks, making it difficult to assess the complete data science workflow. Moreover, they struggle to handle real-time changes in intermediate data and fail to adapt dynamically to evolving task dependencies inherent to data science problems. In this paper, we present Data Interpreter, an LLM-based agent designed to automatically solve various data science problems end-to-end. Our Data Interpreter incorporates two key modules: 1) Hierarchical Graph Modeling, which breaks down complex problems into manageable subproblems, enabling dynamic node generation and graph optimization; and 2) Programmable Node Generation, a technique that refines and verifies each subproblem to iteratively improve code generation results and robustness. Extensive experiments consistently demonstrate the superiority of Data Interpreter. On InfiAgent-DABench, it achieves a 25% performance boost, raising accuracy from 75.9% to 94.9%. For machine learning and open-ended tasks, it improves performance from 88% to 95%, and from 60% to 97%, respectively. Moreover, on the MATH dataset, Data Interpreter achieves remarkable performance with a 26% improvement compared to state-of-the-art baselines. The code is available at https://github.com/geekan/MetaGPT.
title Data Interpreter: An LLM Agent For Data Science
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2402.18679