Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Jin, Chen, Li, Xian, Xun, Luo, An, Tian, Fangqiao, Wang, Ganghua, Doss, Charles, Shen, Xiaotong, Ding, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Agentic AI Match the Performance of Human Data Scientists?
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
by: Luo, An, et al.
Published: (2026)
by: Luo, An, et al.
Published: (2026)
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
MRMS-Net and LMRMS-Net: Scalable Multi-Representation Multi-Scale Networks for Time Series Classification
by: Alagöz, Celal, et al.
Published: (2026)
by: Alagöz, Celal, et al.
Published: (2026)
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
by: Sharma, Akul, et al.
Published: (2025)
by: Sharma, Akul, et al.
Published: (2025)
Benchmarking Catastrophic Forgetting Mitigation Methods in Federated Time Series Forecasting
by: Hallak, Khaled, et al.
Published: (2025)
by: Hallak, Khaled, et al.
Published: (2025)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
by: Bell-Navas, Andrés, et al.
Published: (2025)
by: Bell-Navas, Andrés, et al.
Published: (2025)
Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models
by: Wibbeke, Jelke, et al.
Published: (2025)
by: Wibbeke, Jelke, et al.
Published: (2025)
Predicting Traffic Accident Severity with Deep Neural Networks
by: Bibb, Meghan, et al.
Published: (2025)
by: Bibb, Meghan, et al.
Published: (2025)
Location based Probabilistic Load Forecasting of EV Charging Sites: Deep Transfer Learning with Multi-Quantile Temporal Convolutional Network
by: Ali, Mohammad Wazed, et al.
Published: (2024)
by: Ali, Mohammad Wazed, et al.
Published: (2024)
A Standardized Benchmark for Multilabel Antimicrobial Peptide Classification
by: Ojeda, Sebastian, et al.
Published: (2025)
by: Ojeda, Sebastian, et al.
Published: (2025)
Less is More: Strategic Expert Selection Outperforms Ensemble Complexity in Traffic Forecasting
by: Guettala, Walid, et al.
Published: (2025)
by: Guettala, Walid, et al.
Published: (2025)
Improving Efficiency of Sampling-based Motion Planning via Message-Passing Monte Carlo
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
Randomized Spline Trees for Functional Data Classification: Theory and Application to Environmental Time Series
by: Riccio, Donato, et al.
Published: (2024)
by: Riccio, Donato, et al.
Published: (2024)
Benchmarking changepoint detection algorithms on cardiac time series
by: Cakmak, Ayse, et al.
Published: (2024)
by: Cakmak, Ayse, et al.
Published: (2024)
Machine learning technique for morphological classification of galaxies from SDSS. IV. Visual inspection vs CNN for merging, irregular, edge-on, barred, ringed, and with dust lanes galaxies at 0.02<z<0.1
by: V., Dobrycheva D., et al.
Published: (2026)
by: V., Dobrycheva D., et al.
Published: (2026)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
by: Mukherjee, Debdeep, et al.
Published: (2025)
by: Mukherjee, Debdeep, et al.
Published: (2025)
Single-Step Reconstruction-Free Anomaly Detection and Segmentation via Diffusion Models
by: Moradi, Mehrdad, et al.
Published: (2025)
by: Moradi, Mehrdad, et al.
Published: (2025)
MEG-to-MEG Transfer Learning and Cross-Task Speech/Silence Detection with Limited Data
by: de Zuazo, Xabier, et al.
Published: (2026)
by: de Zuazo, Xabier, et al.
Published: (2026)
A Class of Topological Pseudodistances for Fast Comparison of Persistence Diagrams
by: Nuñez, Rolando Kindelan, et al.
Published: (2024)
by: Nuñez, Rolando Kindelan, et al.
Published: (2024)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Development of ultra-high efficiency soft X-ray angle-resolved photoemission spectroscopy equipped with deep prior-based denoising method
by: Yamagami, Kohei, et al.
Published: (2025)
by: Yamagami, Kohei, et al.
Published: (2025)
Classifier Calibration at Scale: An Empirical Study of Model-Agnostic Post-Hoc Methods
by: Manokhin, Valery, et al.
Published: (2026)
by: Manokhin, Valery, et al.
Published: (2026)
GRALIS: A Unified Canonical Framework for Linear Attribution Methods via Riesz Representation
by: Fanale, Raimondo
Published: (2026)
by: Fanale, Raimondo
Published: (2026)
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
Robust Probabilistic Load Forecasting for a Single Household: A Comparative Study from SARIMA to Transformers on the REFIT Dataset
by: Manoj, Midhun
Published: (2025)
by: Manoj, Midhun
Published: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
by: Patel, Urjitkumar, et al.
Published: (2025)
by: Patel, Urjitkumar, et al.
Published: (2025)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025)
by: K, Prasanth K, et al.
Published: (2025)
Analyzing the Impact of Multimodal Perception on Sample Complexity and Optimization Landscapes in Imitation Learning
by: Abuelsamen, Luai, et al.
Published: (2025)
by: Abuelsamen, Luai, et al.
Published: (2025)
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
by: Lentsch, Ted, et al.
Published: (2026)
by: Lentsch, Ted, et al.
Published: (2026)
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
by: Lentsch, Ted, et al.
Published: (2024)
by: Lentsch, Ted, et al.
Published: (2024)
A Single Image Is All You Need: Zero-Shot Anomaly Localization Without Training Data
by: Moradi, Mehrdad, et al.
Published: (2025)
by: Moradi, Mehrdad, et al.
Published: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
Diverse capability and scaling of diffusion and auto-regressive models when learning abstract rules
by: Wang, Binxu, et al.
Published: (2024)
by: Wang, Binxu, et al.
Published: (2024)
Persistent Topological Structures and Cohomological Flows as a Mathematical Framework for Brain-Inspired Representation Learning
by: Girish, Preksha, et al.
Published: (2025)
by: Girish, Preksha, et al.
Published: (2025)
Tensor Neyman-Pearson Classification: Theory, Algorithms, and Error Control
by: Liu, Lingchong, et al.
Published: (2025)
by: Liu, Lingchong, et al.
Published: (2025)
Dense Video Understanding with Gated Residual Tokenization
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Similar Items
-
Can Agentic AI Match the Performance of Human Data Scientists?
by: Luo, An, et al.
Published: (2025) -
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
by: Luo, An, et al.
Published: (2026) -
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
by: Luo, An, et al.
Published: (2025) -
MRMS-Net and LMRMS-Net: Scalable Multi-Representation Multi-Scale Networks for Time Series Classification
by: Alagöz, Celal, et al.
Published: (2026) -
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
by: Sharma, Akul, et al.
Published: (2025)