Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Miller, Evan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Statistical Multicriteria Evaluation of LLM-Generated Text
par: Arias, Esteban Garces, et autres
Publié: (2025)
par: Arias, Esteban Garces, et autres
Publié: (2025)
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
par: Lum, Kristian, et autres
Publié: (2024)
par: Lum, Kristian, et autres
Publié: (2024)
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
par: Hobelsberger, Christian, et autres
Publié: (2025)
par: Hobelsberger, Christian, et autres
Publié: (2025)
Bayesian Evaluation of Large Language Model Behavior
par: Longjohn, Rachel, et autres
Publié: (2025)
par: Longjohn, Rachel, et autres
Publié: (2025)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
par: Hardy, Michael
Publié: (2024)
par: Hardy, Michael
Publié: (2024)
Does a Large Language Model Really Speak in Human-Like Language?
par: Park, Mose, et autres
Publié: (2025)
par: Park, Mose, et autres
Publié: (2025)
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios
par: Ackerman, Samuel, et autres
Publié: (2024)
par: Ackerman, Samuel, et autres
Publié: (2024)
Auditing the Use of Language Models to Guide Hiring Decisions
par: Gaebler, Johann D., et autres
Publié: (2024)
par: Gaebler, Johann D., et autres
Publié: (2024)
Large Language Models for Full-Text Methods Assessment: A Case Study on Mediation Analysis
par: Zhang, Wenqing, et autres
Publié: (2025)
par: Zhang, Wenqing, et autres
Publié: (2025)
The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control
par: Lommel, Arle, et autres
Publié: (2024)
par: Lommel, Arle, et autres
Publié: (2024)
LAVA: Language Model Assisted Verbal Autopsy for Cause-of-Death Determination
par: Chen, Yiqun T., et autres
Publié: (2025)
par: Chen, Yiqun T., et autres
Publié: (2025)
Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi
par: Kaptur, Dandan Chen, et autres
Publié: (2025)
par: Kaptur, Dandan Chen, et autres
Publié: (2025)
Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors
par: Abdelrahman, Ahmed S., et autres
Publié: (2025)
par: Abdelrahman, Ahmed S., et autres
Publié: (2025)
Judging It, Washing It: Scoring and Greenwashing Corporate Climate Disclosures using Large Language Models
par: Chuang, Marianne, et autres
Publié: (2025)
par: Chuang, Marianne, et autres
Publié: (2025)
Personalized Prediction of Perceived Message Effectiveness Using Large Language Model Based Digital Twins
par: Han, Jasmin, et autres
Publié: (2026)
par: Han, Jasmin, et autres
Publié: (2026)
Gender Inequality in English Textbooks Around the World: an NLP Approach
par: Liu, Tairan
Publié: (2025)
par: Liu, Tairan
Publié: (2025)
A Latent Dirichlet Allocation (LDA) Semantic Text Analytics Approach to Explore Topical Features in Charity Crowdfunding Campaigns
par: Muzumdar, Prathamesh, et autres
Publié: (2024)
par: Muzumdar, Prathamesh, et autres
Publié: (2024)
Sampling the Swadesh List to Identify Similar Languages with Tree Spaces
par: Ordway, Garett, et autres
Publié: (2024)
par: Ordway, Garett, et autres
Publié: (2024)
Language Markers of Emotion Flexibility Predict Depression and Anxiety Treatment Outcomes
par: Brindle, Benjamin, et autres
Publié: (2026)
par: Brindle, Benjamin, et autres
Publié: (2026)
Language Hierarchization Provides the Optimal Solution to Human Working Memory Limits
par: Chen, Luyao, et autres
Publié: (2026)
par: Chen, Luyao, et autres
Publié: (2026)
Classification errors distort findings in automated speech processing: examples and solutions from child-development research
par: Gautheron, Lucas, et autres
Publié: (2025)
par: Gautheron, Lucas, et autres
Publié: (2025)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
par: Ackerman, Samuel, et autres
Publié: (2025)
par: Ackerman, Samuel, et autres
Publié: (2025)
Repeated Sequences Reveal Gaps between Large Language Models and Natural Language
par: Tanaka-Ishii, Kumiko
Publié: (2026)
par: Tanaka-Ishii, Kumiko
Publié: (2026)
Documents Are People and Words Are Items: A Psychometric Approach to Textual Data with Contextual Embeddings
par: Chen, Jinsong
Publié: (2025)
par: Chen, Jinsong
Publié: (2025)
How to Choose a Threshold for an Evaluation Metric for Large Language Models
par: Sarmah, Bhaskarjit, et autres
Publié: (2024)
par: Sarmah, Bhaskarjit, et autres
Publié: (2024)
Statistics of punctuation in experimental literature -- the remarkable case of "Finnegans Wake" by James Joyce
par: Stanisz, Tomasz, et autres
Publié: (2024)
par: Stanisz, Tomasz, et autres
Publié: (2024)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
par: Devunuri, Saipraneeth, et autres
Publié: (2024)
par: Devunuri, Saipraneeth, et autres
Publié: (2024)
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
par: Hu, Zhengmian, et autres
Publié: (2024)
par: Hu, Zhengmian, et autres
Publié: (2024)
Dynamic Topic Language Model on Heterogeneous Children's Mental Health Clinical Notes
par: Ye, Hanwen, et autres
Publié: (2023)
par: Ye, Hanwen, et autres
Publié: (2023)
Metacognitive Myopia in Large Language Models
par: Scholten, Florian, et autres
Publié: (2024)
par: Scholten, Florian, et autres
Publié: (2024)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
par: Oleson, Jon
Publié: (2024)
par: Oleson, Jon
Publié: (2024)
Emotion Detection with Transformers: A Comparative Study
par: Rezapour, Mahdi
Publié: (2024)
par: Rezapour, Mahdi
Publié: (2024)
Still no evidence for an effect of the proportion of non-native speakers on language complexity -- A response to Kauhanen, Einhaus & Walkden (2023)
par: Koplenig, Alexander
Publié: (2023)
par: Koplenig, Alexander
Publié: (2023)
How to Correctly Report LLM-as-a-Judge Evaluations
par: Lee, Chungpa, et autres
Publié: (2025)
par: Lee, Chungpa, et autres
Publié: (2025)
A Design-based Solution for Causal Inference with Text: Can a Language Model Be Too Large?
par: Tierney, Graham, et autres
Publié: (2025)
par: Tierney, Graham, et autres
Publié: (2025)
Improving Probabilistic Models in Text Classification via Active Learning
par: Bosley, Mitchell, et autres
Publié: (2022)
par: Bosley, Mitchell, et autres
Publié: (2022)
From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages
par: Schöffel, Matthias, et autres
Publié: (2026)
par: Schöffel, Matthias, et autres
Publié: (2026)
Reliable and Efficient Amortized Model-based Evaluation
par: Truong, Sang, et autres
Publié: (2025)
par: Truong, Sang, et autres
Publié: (2025)
Limits of Large Language Models in Debating Humans
par: Flamino, James, et autres
Publié: (2024)
par: Flamino, James, et autres
Publié: (2024)
Exploring the Comprehension of ChatGPT in Traditional Chinese Medicine Knowledge
par: Yizhen, Li, et autres
Publié: (2024)
par: Yizhen, Li, et autres
Publié: (2024)
Documents similaires
-
Statistical Multicriteria Evaluation of LLM-Generated Text
par: Arias, Esteban Garces, et autres
Publié: (2025) -
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
par: Lum, Kristian, et autres
Publié: (2024) -
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
par: Hobelsberger, Christian, et autres
Publié: (2025) -
Bayesian Evaluation of Large Language Model Behavior
par: Longjohn, Rachel, et autres
Publié: (2025) -
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
par: Hardy, Michael
Publié: (2024)