Performance Evaluation of Large Language Models in Statistical Programming

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Xinyi, Xie, Kexin, Lee, Lina, Chen, Ruizhe, Clark, Jared M., He, Hao, He, Haoran, Min, Jie, Zhang, Xinlei, Zheng, Simin, Zhang, Zhiyang, Deng, Xinwei, Hong, Yili
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915159150166016
author Song, Xinyi
Xie, Kexin
Lee, Lina
Chen, Ruizhe
Clark, Jared M.
He, Hao
He, Haoran
Min, Jie
Zhang, Xinlei
Zheng, Simin
Zhang, Zhiyang
Deng, Xinwei
Hong, Yili
author_facet Song, Xinyi
Xie, Kexin
Lee, Lina
Chen, Ruizhe
Clark, Jared M.
He, Hao
He, Haoran
Min, Jie
Zhang, Xinlei
Zheng, Simin
Zhang, Zhiyang
Deng, Xinwei
Hong, Yili
contents The programming capabilities of large language models (LLMs) have revolutionized automatic code generation and opened new avenues for automatic statistical analysis. However, the validity and quality of these generated codes need to be systematically evaluated before they can be widely adopted. Despite their growing prominence, a comprehensive evaluation of statistical code generated by LLMs remains scarce in the literature. In this paper, we assess the performance of LLMs, including two versions of ChatGPT and one version of Llama, in the domain of SAS programming for statistical analysis. Our study utilizes a set of statistical analysis tasks encompassing diverse statistical topics and datasets. Each task includes a problem description, dataset information, and human-verified SAS code. We conduct a comprehensive assessment of the quality of SAS code generated by LLMs through human expert evaluation based on correctness, effectiveness, readability, executability, and the accuracy of output results. The analysis of rating scores reveals that while LLMs demonstrate usefulness in generating syntactically correct code, they struggle with tasks requiring deep domain understanding and may produce redundant or incorrect results. This study offers valuable insights into the capabilities and limitations of LLMs in statistical programming, providing guidance for future advancements in AI-assisted coding systems for statistical analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13117
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Performance Evaluation of Large Language Models in Statistical Programming
Song, Xinyi
Xie, Kexin
Lee, Lina
Chen, Ruizhe
Clark, Jared M.
He, Hao
He, Haoran
Min, Jie
Zhang, Xinlei
Zheng, Simin
Zhang, Zhiyang
Deng, Xinwei
Hong, Yili
Applications
Artificial Intelligence
The programming capabilities of large language models (LLMs) have revolutionized automatic code generation and opened new avenues for automatic statistical analysis. However, the validity and quality of these generated codes need to be systematically evaluated before they can be widely adopted. Despite their growing prominence, a comprehensive evaluation of statistical code generated by LLMs remains scarce in the literature. In this paper, we assess the performance of LLMs, including two versions of ChatGPT and one version of Llama, in the domain of SAS programming for statistical analysis. Our study utilizes a set of statistical analysis tasks encompassing diverse statistical topics and datasets. Each task includes a problem description, dataset information, and human-verified SAS code. We conduct a comprehensive assessment of the quality of SAS code generated by LLMs through human expert evaluation based on correctness, effectiveness, readability, executability, and the accuracy of output results. The analysis of rating scores reveals that while LLMs demonstrate usefulness in generating syntactically correct code, they struggle with tasks requiring deep domain understanding and may produce redundant or incorrect results. This study offers valuable insights into the capabilities and limitations of LLMs in statistical programming, providing guidance for future advancements in AI-assisted coding systems for statistical analysis.
title Performance Evaluation of Large Language Models in Statistical Programming
topic Applications
Artificial Intelligence
url https://arxiv.org/abs/2502.13117