Bayesian Evaluation of Large Language Model Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | Longjohn, Rachel, Wu, Shang, Kher, Saatvik, Belém, Catarina, Smyth, Padhraic |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Assign Prediction Tasks to Agents with Capacity Constraints
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
Perceptions of Linguistic Uncertainty by Language Models and Humans
by: Belem, Catarina G, et al.
Published: (2024)
by: Belem, Catarina G, et al.
Published: (2024)
Improving Metacognition and Uncertainty Communication in Language Models
by: Steyvers, Mark, et al.
Published: (2025)
by: Steyvers, Mark, et al.
Published: (2025)
Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks
by: Longjohn, Rachel, et al.
Published: (2025)
by: Longjohn, Rachel, et al.
Published: (2025)
What Large Language Models Know and What People Think They Know
by: Steyvers, Mark, et al.
Published: (2024)
by: Steyvers, Mark, et al.
Published: (2024)
Improving and Evaluating Machine Learning Methods for Forensic Shoeprint Matching
by: Jain, Divij, et al.
Published: (2024)
by: Jain, Divij, et al.
Published: (2024)
Benchmark Data Repositories for Better Benchmarking
by: Longjohn, Rachel, et al.
Published: (2024)
by: Longjohn, Rachel, et al.
Published: (2024)
Semantic Probabilistic Control of Language Models
by: Ahmed, Kareem, et al.
Published: (2025)
by: Ahmed, Kareem, et al.
Published: (2025)
How to Choose a Threshold for an Evaluation Metric for Large Language Models
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
Likelihood ratios for changepoints in categorical event data with applications in digital forensics
by: Rachel Longjohn, et al.
Published: (2024)
by: Rachel Longjohn, et al.
Published: (2024)
Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
A Design-based Solution for Causal Inference with Text: Can a Language Model Be Too Large?
by: Tierney, Graham, et al.
Published: (2025)
by: Tierney, Graham, et al.
Published: (2025)
Domain-Shift-Aware Conformal Prediction for Large Language Models
by: Lin, Zhexiao, et al.
Published: (2025)
by: Lin, Zhexiao, et al.
Published: (2025)
Dynamic Topic Language Model on Heterogeneous Children's Mental Health Clinical Notes
by: Ye, Hanwen, et al.
Published: (2023)
by: Ye, Hanwen, et al.
Published: (2023)
Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis
by: Jantre, Sanket, et al.
Published: (2025)
by: Jantre, Sanket, et al.
Published: (2025)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
Understanding Gender Bias in AI-Generated Product Descriptions
by: Kelly, Markelle, et al.
Published: (2025)
by: Kelly, Markelle, et al.
Published: (2025)
How to Correctly Report LLM-as-a-Judge Evaluations
by: Lee, Chungpa, et al.
Published: (2025)
by: Lee, Chungpa, et al.
Published: (2025)
Augmented Risk Prediction for the Onset of Alzheimer's Disease from Electronic Health Records with Large Language Models
by: Wang, Jiankun, et al.
Published: (2024)
by: Wang, Jiankun, et al.
Published: (2024)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Explainable Automatic Grading with Neural Additive Models
by: Condor, Aubrey, et al.
Published: (2024)
by: Condor, Aubrey, et al.
Published: (2024)
Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
by: Hang, Haotian, et al.
Published: (2025)
by: Hang, Haotian, et al.
Published: (2025)
Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
by: Nagori, Aditya, et al.
Published: (2025)
by: Nagori, Aditya, et al.
Published: (2025)
The Use of a Large Language Model for Cyberbullying Detection
by: Ogunleye, Bayode, et al.
Published: (2024)
by: Ogunleye, Bayode, et al.
Published: (2024)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
by: Wasi, Azmine Toushik, et al.
Published: (2024)
by: Wasi, Azmine Toushik, et al.
Published: (2024)
Context-Alignment: Activating and Enhancing LLM Capabilities in Time Series
by: Hu, Yuxiao, et al.
Published: (2025)
by: Hu, Yuxiao, et al.
Published: (2025)
Classification errors distort findings in automated speech processing: examples and solutions from child-development research
by: Gautheron, Lucas, et al.
Published: (2025)
by: Gautheron, Lucas, et al.
Published: (2025)
Subjective Perspectives within Learned Representations Predict High-Impact Innovation
by: Cao, Likun, et al.
Published: (2025)
by: Cao, Likun, et al.
Published: (2025)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
A meta-analysis on the performance of machine-learning based language models for sentiment analysis
by: Rohde, Elena, et al.
Published: (2025)
by: Rohde, Elena, et al.
Published: (2025)
Ensemble Kalman filter for uncertainty in human language comprehension
by: Bhandari, Diksha, et al.
Published: (2025)
by: Bhandari, Diksha, et al.
Published: (2025)
Extracting Emotion Phrases from Tweets using BART
by: Rezapour, Mahdi
Published: (2024)
by: Rezapour, Mahdi
Published: (2024)
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
by: Li, Baishi, et al.
Published: (2026)
by: Li, Baishi, et al.
Published: (2026)
Specific language impairment (SLI) detection pipeline from transcriptions of spontaneous narratives
by: Arena, Santiago, et al.
Published: (2024)
by: Arena, Santiago, et al.
Published: (2024)
Causal Representation Learning with Generative Artificial Intelligence: Application to Texts as Treatments
by: Imai, Kosuke, et al.
Published: (2024)
by: Imai, Kosuke, et al.
Published: (2024)
Detecting LLM-Generated Text with Performance Guarantees
by: Zhou, Hongyi, et al.
Published: (2026)
by: Zhou, Hongyi, et al.
Published: (2026)
Deep literature reviews: an application of fine-tuned language models to migration research
by: Iacus, Stefano M., et al.
Published: (2025)
by: Iacus, Stefano M., et al.
Published: (2025)
Tailored Behavior-Change Messaging for Physical Activity: Integrating Contextual Bandits and Large Language Models
by: Song, Haochen, et al.
Published: (2025)
by: Song, Haochen, et al.
Published: (2025)
Deep R Programming
by: Gagolewski, Marek
Published: (2022)
by: Gagolewski, Marek
Published: (2022)
Similar Items
-
Learning to Assign Prediction Tasks to Agents with Capacity Constraints
by: Wu, Shang, et al.
Published: (2026) -
Perceptions of Linguistic Uncertainty by Language Models and Humans
by: Belem, Catarina G, et al.
Published: (2024) -
Improving Metacognition and Uncertainty Communication in Language Models
by: Steyvers, Mark, et al.
Published: (2025) -
Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks
by: Longjohn, Rachel, et al.
Published: (2025) -
What Large Language Models Know and What People Think They Know
by: Steyvers, Mark, et al.
Published: (2024)