Salvato in:
Dettagli Bibliografici
Autori principali: Hsu, Chi-Yang, Braylan, Alexander, Su, Yiheng, Lease, Matthew, Alonso, Omar
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2509.07309
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908760950177792
author Hsu, Chi-Yang
Braylan, Alexander
Su, Yiheng
Lease, Matthew
Alonso, Omar
author_facet Hsu, Chi-Yang
Braylan, Alexander
Su, Yiheng
Lease, Matthew
Alonso, Omar
contents Confidence estimation infers a probability for whether each model output is correct or not. While predicting such binary correctness is sensible for tasks with exact answers, free-form generation tasks are often more nuanced, with output quality being both fine-grained and multi-faceted. We thus propose Performance Interval Estimation (PIE) to predict both: 1) point estimates for any arbitrary set of continuous-valued evaluation metrics; and 2) calibrated uncertainty intervals around these point estimates. We then compare two approaches: LLM-as-judge vs. classic regression with confidence estimation features. Evaluation over 11 datasets spans summarization, translation, code generation, function-calling, and question answering. Regression is seen to achieve both: i) lower error point estimates of metric scores; and ii) well-calibrated uncertainty intervals. To support reproduction and follow-on work, we share our data and code.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07309
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PIE: Performance Interval Estimation for Free-Form Generation Tasks
Hsu, Chi-Yang
Braylan, Alexander
Su, Yiheng
Lease, Matthew
Alonso, Omar
Computation and Language
Machine Learning
Confidence estimation infers a probability for whether each model output is correct or not. While predicting such binary correctness is sensible for tasks with exact answers, free-form generation tasks are often more nuanced, with output quality being both fine-grained and multi-faceted. We thus propose Performance Interval Estimation (PIE) to predict both: 1) point estimates for any arbitrary set of continuous-valued evaluation metrics; and 2) calibrated uncertainty intervals around these point estimates. We then compare two approaches: LLM-as-judge vs. classic regression with confidence estimation features. Evaluation over 11 datasets spans summarization, translation, code generation, function-calling, and question answering. Regression is seen to achieve both: i) lower error point estimates of metric scores; and ii) well-calibrated uncertainty intervals. To support reproduction and follow-on work, we share our data and code.
title PIE: Performance Interval Estimation for Free-Form Generation Tasks
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.07309