Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yuqi, Zhong, Yi, Zhang, Jintian, Zhang, Ziheng, Qiao, Shuofei, Luo, Yujie, Du, Lun, Zheng, Da, Zhang, Ningyu, Chen, Huajun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917077322825728
author Zhu, Yuqi
Zhong, Yi
Zhang, Jintian
Zhang, Ziheng
Qiao, Shuofei
Luo, Yujie
Du, Lun
Zheng, Da
Zhang, Ningyu
Chen, Huajun
author_facet Zhu, Yuqi
Zhong, Yi
Zhang, Jintian
Zhang, Ziheng
Qiao, Shuofei
Luo, Yujie
Du, Lun
Zheng, Da
Zhang, Ningyu
Chen, Huajun
contents Large Language Models (LLMs) hold promise in automating data analysis tasks, yet open-source models face significant limitations in these kinds of reasoning-intensive scenarios. In this work, we investigate strategies to enhance the data analysis capabilities of open-source LLMs. By curating a seed dataset of diverse, realistic scenarios, we evaluate model behavior across three core dimensions: data understanding, code generation, and strategic planning. Our analysis reveals three key findings: (1) Strategic planning quality serves as the primary determinant of model performance; (2) Interaction design and task complexity significantly influence reasoning capabilities; (3) Data quality demonstrates a greater impact than diversity in achieving optimal performance. We leverage these insights to develop a data synthesis methodology, demonstrating significant improvements in open-source LLMs' analytical reasoning capabilities. Code is available at https://github.com/zjunlp/DataMind.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
Zhu, Yuqi
Zhong, Yi
Zhang, Jintian
Zhang, Ziheng
Qiao, Shuofei
Luo, Yujie
Du, Lun
Zheng, Da
Zhang, Ningyu
Chen, Huajun
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
Multiagent Systems
Large Language Models (LLMs) hold promise in automating data analysis tasks, yet open-source models face significant limitations in these kinds of reasoning-intensive scenarios. In this work, we investigate strategies to enhance the data analysis capabilities of open-source LLMs. By curating a seed dataset of diverse, realistic scenarios, we evaluate model behavior across three core dimensions: data understanding, code generation, and strategic planning. Our analysis reveals three key findings: (1) Strategic planning quality serves as the primary determinant of model performance; (2) Interaction design and task complexity significantly influence reasoning capabilities; (3) Data quality demonstrates a greater impact than diversity in achieving optimal performance. We leverage these insights to develop a data synthesis methodology, demonstrating significant improvements in open-source LLMs' analytical reasoning capabilities. Code is available at https://github.com/zjunlp/DataMind.
title Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2506.19794