QA-Calibration of Language Model Confidence Scores

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Manggala, Putra, Mastakouri, Atalanti, Kirschbaum, Elke, Kasiviswanathan, Shiva Prasad, Ramdas, Aaditya
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916637631840256
author Manggala, Putra
Mastakouri, Atalanti
Kirschbaum, Elke
Kasiviswanathan, Shiva Prasad
Ramdas, Aaditya
author_facet Manggala, Putra
Mastakouri, Atalanti
Kirschbaum, Elke
Kasiviswanathan, Shiva Prasad
Ramdas, Aaditya
contents To use generative question-and-answering (QA) systems for decision-making and in any critical application, these systems need to provide well-calibrated confidence scores that reflect the correctness of their answers. Existing calibration methods aim to ensure that the confidence score is, *on average*, indicative of the likelihood that the answer is correct. We argue, however, that this standard (average-case) notion of calibration is difficult to interpret for decision-making in generative QA. To address this, we generalize the standard notion of average calibration and introduce QA-calibration, which ensures calibration holds across different question-and-answer groups. We then propose discretized posthoc calibration schemes for achieving QA-calibration. We establish distribution-free guarantees on the performance of this method and validate our method on confidence scores returned by elicitation prompts across multiple QA benchmarks and large language models (LLMs).
format Preprint
id arxiv_https___arxiv_org_abs_2410_06615
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle QA-Calibration of Language Model Confidence Scores
Manggala, Putra
Mastakouri, Atalanti
Kirschbaum, Elke
Kasiviswanathan, Shiva Prasad
Ramdas, Aaditya
Computation and Language
Machine Learning
To use generative question-and-answering (QA) systems for decision-making and in any critical application, these systems need to provide well-calibrated confidence scores that reflect the correctness of their answers. Existing calibration methods aim to ensure that the confidence score is, *on average*, indicative of the likelihood that the answer is correct. We argue, however, that this standard (average-case) notion of calibration is difficult to interpret for decision-making in generative QA. To address this, we generalize the standard notion of average calibration and introduce QA-calibration, which ensures calibration holds across different question-and-answer groups. We then propose discretized posthoc calibration schemes for achieving QA-calibration. We establish distribution-free guarantees on the performance of this method and validate our method on confidence scores returned by elicitation prompts across multiple QA benchmarks and large language models (LLMs).
title QA-Calibration of Language Model Confidence Scores
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.06615