Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917077322825728 |
|---|---|
| author | Zhu, Yuqi Zhong, Yi Zhang, Jintian Zhang, Ziheng Qiao, Shuofei Luo, Yujie Du, Lun Zheng, Da Zhang, Ningyu Chen, Huajun |
| author_facet | Zhu, Yuqi Zhong, Yi Zhang, Jintian Zhang, Ziheng Qiao, Shuofei Luo, Yujie Du, Lun Zheng, Da Zhang, Ningyu Chen, Huajun |
| contents | Large Language Models (LLMs) hold promise in automating data analysis tasks, yet open-source models face significant limitations in these kinds of reasoning-intensive scenarios. In this work, we investigate strategies to enhance the data analysis capabilities of open-source LLMs. By curating a seed dataset of diverse, realistic scenarios, we evaluate model behavior across three core dimensions: data understanding, code generation, and strategic planning. Our analysis reveals three key findings: (1) Strategic planning quality serves as the primary determinant of model performance; (2) Interaction design and task complexity significantly influence reasoning capabilities; (3) Data quality demonstrates a greater impact than diversity in achieving optimal performance. We leverage these insights to develop a data synthesis methodology, demonstrating significant improvements in open-source LLMs' analytical reasoning capabilities. Code is available at https://github.com/zjunlp/DataMind. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_19794 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study Zhu, Yuqi Zhong, Yi Zhang, Jintian Zhang, Ziheng Qiao, Shuofei Luo, Yujie Du, Lun Zheng, Da Zhang, Ningyu Chen, Huajun Computation and Language Artificial Intelligence Information Retrieval Machine Learning Multiagent Systems Large Language Models (LLMs) hold promise in automating data analysis tasks, yet open-source models face significant limitations in these kinds of reasoning-intensive scenarios. In this work, we investigate strategies to enhance the data analysis capabilities of open-source LLMs. By curating a seed dataset of diverse, realistic scenarios, we evaluate model behavior across three core dimensions: data understanding, code generation, and strategic planning. Our analysis reveals three key findings: (1) Strategic planning quality serves as the primary determinant of model performance; (2) Interaction design and task complexity significantly influence reasoning capabilities; (3) Data quality demonstrates a greater impact than diversity in achieving optimal performance. We leverage these insights to develop a data synthesis methodology, demonstrating significant improvements in open-source LLMs' analytical reasoning capabilities. Code is available at https://github.com/zjunlp/DataMind. |
| title | Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study |
| topic | Computation and Language Artificial Intelligence Information Retrieval Machine Learning Multiagent Systems |
| url | https://arxiv.org/abs/2506.19794 |