StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Xinyi, Lee, Lina, Xie, Kexin, Liu, Xueying, Deng, Xinwei, Hong, Yili |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance Evaluation of Large Language Models in Statistical Programming
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
Applied Statistics in the Era of Artificial Intelligence: A Review and Vision
by: Min, Jie, et al.
Published: (2024)
by: Min, Jie, et al.
Published: (2024)
Adaptive Bi-Level Variable Selection of Conditional Main Effects for Generalized Linear Models
by: Xie, Kexin, et al.
Published: (2026)
by: Xie, Kexin, et al.
Published: (2026)
The Use of Variational Inference for Lifetime Data with Spatial Correlations
by: Wang, Yueyao, et al.
Published: (2025)
by: Wang, Yueyao, et al.
Published: (2025)
A Comprehensive Case Study on the Performance of Machine Learning Methods on the Classification of Solar Panel Electroluminescence Images
by: Song, Xinyi, et al.
Published: (2024)
by: Song, Xinyi, et al.
Published: (2024)
Bridging the Data Gap in AI Reliability Research and Establishing DR-AIR, a Comprehensive Data Repository for AI Reliability
by: Zheng, Simin, et al.
Published: (2025)
by: Zheng, Simin, et al.
Published: (2025)
A Detailed Historical and Statistical Analysis of the Influence of Hardware Artifacts on SPEC Integer Benchmark Performance
by: Wang, Yueyao, et al.
Published: (2024)
by: Wang, Yueyao, et al.
Published: (2024)
Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
by: Madden, Emma Rose
Published: (2025)
by: Madden, Emma Rose
Published: (2025)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
Modeling Multivariate Degradation Data with Dynamic Covariates Under a Bayesian Framework
by: Lin, Zhengzhi, et al.
Published: (2025)
by: Lin, Zhengzhi, et al.
Published: (2025)
Subnational Geocoding of Global Disasters Using Large Language Models
by: Ronco, Michele, et al.
Published: (2025)
by: Ronco, Michele, et al.
Published: (2025)
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
by: Xu, Zhiyu, et al.
Published: (2025)
by: Xu, Zhiyu, et al.
Published: (2025)
How Generalizable Is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation
by: Vincent, Joseph A., et al.
Published: (2024)
by: Vincent, Joseph A., et al.
Published: (2024)
From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
AI for Handball: predicting and explaining the 2024 Olympic Games tournament with Deep Learning and Large Language Models
by: Felice, Florian
Published: (2024)
by: Felice, Florian
Published: (2024)
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
by: Tschisgale, Paul, et al.
Published: (2026)
by: Tschisgale, Paul, et al.
Published: (2026)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
Statistical Analysis and End-to-End Performance Evaluation of Traffic Models for Automotive Data
by: Bullo, Marcello, et al.
Published: (2025)
by: Bullo, Marcello, et al.
Published: (2025)
Modeling Spatially Correlated Failure-time Data Under Two Distance Functions with an Application to Titan GPU Data
by: Clark, Jared M., et al.
Published: (2025)
by: Clark, Jared M., et al.
Published: (2025)
Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis
by: Jantre, Sanket, et al.
Published: (2025)
by: Jantre, Sanket, et al.
Published: (2025)
Efficient Prediction of Pass@k Scaling in Large Language Models
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
LLM4ED: Large Language Models for Automatic Equation Discovery
by: Du, Mengge, et al.
Published: (2024)
by: Du, Mengge, et al.
Published: (2024)
More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries
by: Song, Hongwen, et al.
Published: (2026)
by: Song, Hongwen, et al.
Published: (2026)
Automatic Generation of Cybersecurity Teaching Cases Using Large Language Models
by: Jiqiang Zhai, et al.
Published: (2025)
by: Jiqiang Zhai, et al.
Published: (2025)
Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects
by: Zhang, Kexin, et al.
Published: (2023)
by: Zhang, Kexin, et al.
Published: (2023)
Statistical Multicriteria Evaluation of LLM-Generated Text
by: Arias, Esteban Garces, et al.
Published: (2025)
by: Arias, Esteban Garces, et al.
Published: (2025)
Estimating Item Difficulty with Large Language Models as Experts
by: Kolesnikova, Diana, et al.
Published: (2026)
by: Kolesnikova, Diana, et al.
Published: (2026)
Classifying Metamorphic versus Single-Fold Proteins with Statistical Learning and AlphaFold2
by: Chen, Yongkai, et al.
Published: (2025)
by: Chen, Yongkai, et al.
Published: (2025)
Impact of Label Noise from Large Language Models Generated Annotations on Evaluation of Diagnostic Model Performance
by: Chavoshi, Mohammadreza, et al.
Published: (2025)
by: Chavoshi, Mohammadreza, et al.
Published: (2025)
Metacognitive Myopia in Large Language Models
by: Scholten, Florian, et al.
Published: (2024)
by: Scholten, Florian, et al.
Published: (2024)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
by: Xu, Yang, et al.
Published: (2026)
by: Xu, Yang, et al.
Published: (2026)
The EEPAS Model Revisited: Statistical Formalism and a High-Performance, Reproducible Open-Source Framework
by: Chung, Szu-Chi, et al.
Published: (2025)
by: Chung, Szu-Chi, et al.
Published: (2025)
Large Language Model Predicts Above Normal All India Summer Monsoon Rainfall in 2024
by: Sharma, Ujjawal, et al.
Published: (2024)
by: Sharma, Ujjawal, et al.
Published: (2024)
Are Large Language Models able to Predict Highly Cited Papers? Evidence from Statistical Publications
by: Ye, Zhanshuo, et al.
Published: (2026)
by: Ye, Zhanshuo, et al.
Published: (2026)
Wafer-Level Etch Spatial Profiling for Process Monitoring from Time-Series with Time-LLM
by: Kim, Hyunwoo, et al.
Published: (2026)
by: Kim, Hyunwoo, et al.
Published: (2026)
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
by: Devunuri, Saipraneeth, et al.
Published: (2024)
by: Devunuri, Saipraneeth, et al.
Published: (2024)
International Trade Network: Statistical Analysis and Modeling
by: Sosa, Juan, et al.
Published: (2024)
by: Sosa, Juan, et al.
Published: (2024)
Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys
by: Ye, Zikun, et al.
Published: (2026)
by: Ye, Zikun, et al.
Published: (2026)
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
by: Ye, Junze, et al.
Published: (2025)
by: Ye, Junze, et al.
Published: (2025)
Similar Items
-
Performance Evaluation of Large Language Models in Statistical Programming
by: Song, Xinyi, et al.
Published: (2025) -
Applied Statistics in the Era of Artificial Intelligence: A Review and Vision
by: Min, Jie, et al.
Published: (2024) -
Adaptive Bi-Level Variable Selection of Conditional Main Effects for Generalized Linear Models
by: Xie, Kexin, et al.
Published: (2026) -
The Use of Variational Inference for Lifetime Data with Spatial Correlations
by: Wang, Yueyao, et al.
Published: (2025) -
A Comprehensive Case Study on the Performance of Machine Learning Methods on the Classification of Solar Panel Electroluminescence Images
by: Song, Xinyi, et al.
Published: (2024)