HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
Fuente:
arXiv
Saved in:
| Main Authors: | Luettgau, Lennart, Coppock, Harry, Dubois, Magda, Summerfield, Christopher, Ududec, Cozmin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ask don't tell: Reducing sycophancy in large language models
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
Skewed Score: A statistical framework to assess autograders
by: Dubois, Magda, et al.
Published: (2025)
by: Dubois, Magda, et al.
Published: (2025)
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
by: Summerfield, Christopher, et al.
Published: (2025)
by: Summerfield, Christopher, et al.
Published: (2025)
Seven simple steps for log analysis in AI systems
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
Performance Evaluation of Large Language Models in Statistical Programming
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
Open-World Evaluations for Measuring Frontier AI Capabilities
by: Kapoor, Sayash, et al.
Published: (2026)
by: Kapoor, Sayash, et al.
Published: (2026)
What If They Took the Shot? A Hierarchical Bayesian Framework for Counterfactual Expected Goals
by: Mahmudlu, Mikayil, et al.
Published: (2025)
by: Mahmudlu, Mikayil, et al.
Published: (2025)
StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement
by: Sielinski, Ronald
Published: (2026)
by: Sielinski, Ronald
Published: (2026)
When Do LLM Preferences Predict Downstream Behavior?
by: Slama, Katarina, et al.
Published: (2026)
by: Slama, Katarina, et al.
Published: (2026)
Log analysis is necessary for credible evaluation of AI agents
by: Kirgis, Peter, et al.
Published: (2026)
by: Kirgis, Peter, et al.
Published: (2026)
Embracing Ambiguity: Bayesian Nonparametrics and Stakeholder Participation for Ambiguity-Aware Safety Evaluation
by: Long, Yanan
Published: (2025)
by: Long, Yanan
Published: (2025)
Decision Quality Evaluation Framework at Pinterest
by: Tian, Yuqi, et al.
Published: (2026)
by: Tian, Yuqi, et al.
Published: (2026)
Data-Driven Bayesian Network Models of Hurricane Evacuation Decision Making
by: Wang, Hui Sophie, et al.
Published: (2023)
by: Wang, Hui Sophie, et al.
Published: (2023)
One-shot emergency psychiatric triage across 15 frontier AI chatbots
by: Weilnhammer, Veith, et al.
Published: (2026)
by: Weilnhammer, Veith, et al.
Published: (2026)
Classifying Metamorphic versus Single-Fold Proteins with Statistical Learning and AlphaFold2
by: Chen, Yongkai, et al.
Published: (2025)
by: Chen, Yongkai, et al.
Published: (2025)
Unlocking the Potential of Past Research: Using Generative AI to Reconstruct Healthcare Simulation Models
by: Monks, Thomas, et al.
Published: (2025)
by: Monks, Thomas, et al.
Published: (2025)
Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
by: Madden, Emma Rose
Published: (2025)
by: Madden, Emma Rose
Published: (2025)
AI for Handball: predicting and explaining the 2024 Olympic Games tournament with Deep Learning and Large Language Models
by: Felice, Florian
Published: (2024)
by: Felice, Florian
Published: (2024)
Process-Aware Analysis of Treatment Paths in Heart Failure Patients: A Case Study
by: Beyel, Harry H., et al.
Published: (2024)
by: Beyel, Harry H., et al.
Published: (2024)
How Generalizable Is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation
by: Vincent, Joseph A., et al.
Published: (2024)
by: Vincent, Joseph A., et al.
Published: (2024)
Bayesian Networks for Causal Analysis in Socioecological Systems
by: Cabañas, Rafael, et al.
Published: (2024)
by: Cabañas, Rafael, et al.
Published: (2024)
A network analysis of decision strategies of human experts in steel manufacturing
by: Merten, Daniel Christopher, et al.
Published: (2021)
by: Merten, Daniel Christopher, et al.
Published: (2021)
Bridging the Data Gap in AI Reliability Research and Establishing DR-AIR, a Comprehensive Data Repository for AI Reliability
by: Zheng, Simin, et al.
Published: (2025)
by: Zheng, Simin, et al.
Published: (2025)
A Statistical Theory of Regularization-Based Continual Learning
by: Zhao, Xuyang, et al.
Published: (2024)
by: Zhao, Xuyang, et al.
Published: (2024)
The Advancement of Personalized Learning Potentially Accelerated by Generative AI
by: Wei, Yuang, et al.
Published: (2024)
by: Wei, Yuang, et al.
Published: (2024)
Analyzing the Impact of Climate Change With Major Emphasis on Pollution: A Comparative Study of ML and Statistical Models in Time Series Data
by: Mishra, Anurag, et al.
Published: (2024)
by: Mishra, Anurag, et al.
Published: (2024)
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
by: Hu, Zhengmian, et al.
Published: (2024)
by: Hu, Zhengmian, et al.
Published: (2024)
Explainability of Complex AI Models with Correlation Impact Ratio
by: Sengupta, Poushali, et al.
Published: (2026)
by: Sengupta, Poushali, et al.
Published: (2026)
Synergizing chemical and AI communities for advancing laboratories of the future
by: Oh, Saejin, et al.
Published: (2025)
by: Oh, Saejin, et al.
Published: (2025)
A More Realistic Evaluation of Cross-Frequency Transfer Learning and Foundation Forecasting Models
by: Olivares, Kin G., et al.
Published: (2025)
by: Olivares, Kin G., et al.
Published: (2025)
TCKAN:A Novel Integrated Network Model for Predicting Mortality Risk in Sepsis Patients
by: Dong, Fanglin
Published: (2024)
by: Dong, Fanglin
Published: (2024)
On the Practice of Deep Hierarchical Ensemble Network for Ad Conversion Rate Prediction
by: Zhuang, Jinfeng, et al.
Published: (2025)
by: Zhuang, Jinfeng, et al.
Published: (2025)
A Regression Mixture Model to understand the effect of the Covid-19 pandemic on Public Transport Ridership
by: Moreau, Hugues, et al.
Published: (2024)
by: Moreau, Hugues, et al.
Published: (2024)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
by: Hardy, Michael
Published: (2024)
by: Hardy, Michael
Published: (2024)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
by: Devunuri, Saipraneeth, et al.
Published: (2024)
by: Devunuri, Saipraneeth, et al.
Published: (2024)
A Distribution-Free Framework for Rewrite-Based Human-text Detection via Knockoff Filtering
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
Visual Error Patterns in Multi-Modal AI: A Statistical Approach
by: Wang, Ching-Yi
Published: (2024)
by: Wang, Ching-Yi
Published: (2024)
Vox Populi, Vox AI? Using Language Models to Estimate German Public Opinion
by: von der Heyde, Leah, et al.
Published: (2024)
by: von der Heyde, Leah, et al.
Published: (2024)
Decade-long Emission Forecasting with an Ensemble Model in Taiwan
by: Hung, Gordon, et al.
Published: (2025)
by: Hung, Gordon, et al.
Published: (2025)
Similar Items
-
Ask don't tell: Reducing sycophancy in large language models
by: Dubois, Magda, et al.
Published: (2026) -
Skewed Score: A statistical framework to assess autograders
by: Dubois, Magda, et al.
Published: (2025) -
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
by: Summerfield, Christopher, et al.
Published: (2025) -
Seven simple steps for log analysis in AI systems
by: Dubois, Magda, et al.
Published: (2026) -
Performance Evaluation of Large Language Models in Statistical Programming
by: Song, Xinyi, et al.
Published: (2025)