QA-Calibration of Language Model Confidence Scores
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916637631840256 |
|---|---|
| author | Manggala, Putra Mastakouri, Atalanti Kirschbaum, Elke Kasiviswanathan, Shiva Prasad Ramdas, Aaditya |
| author_facet | Manggala, Putra Mastakouri, Atalanti Kirschbaum, Elke Kasiviswanathan, Shiva Prasad Ramdas, Aaditya |
| contents | To use generative question-and-answering (QA) systems for decision-making and in any critical application, these systems need to provide well-calibrated confidence scores that reflect the correctness of their answers. Existing calibration methods aim to ensure that the confidence score is, *on average*, indicative of the likelihood that the answer is correct. We argue, however, that this standard (average-case) notion of calibration is difficult to interpret for decision-making in generative QA. To address this, we generalize the standard notion of average calibration and introduce QA-calibration, which ensures calibration holds across different question-and-answer groups. We then propose discretized posthoc calibration schemes for achieving QA-calibration. We establish distribution-free guarantees on the performance of this method and validate our method on confidence scores returned by elicitation prompts across multiple QA benchmarks and large language models (LLMs). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_06615 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | QA-Calibration of Language Model Confidence Scores Manggala, Putra Mastakouri, Atalanti Kirschbaum, Elke Kasiviswanathan, Shiva Prasad Ramdas, Aaditya Computation and Language Machine Learning To use generative question-and-answering (QA) systems for decision-making and in any critical application, these systems need to provide well-calibrated confidence scores that reflect the correctness of their answers. Existing calibration methods aim to ensure that the confidence score is, *on average*, indicative of the likelihood that the answer is correct. We argue, however, that this standard (average-case) notion of calibration is difficult to interpret for decision-making in generative QA. To address this, we generalize the standard notion of average calibration and introduce QA-calibration, which ensures calibration holds across different question-and-answer groups. We then propose discretized posthoc calibration schemes for achieving QA-calibration. We establish distribution-free guarantees on the performance of this method and validate our method on confidence scores returned by elicitation prompts across multiple QA benchmarks and large language models (LLMs). |
| title | QA-Calibration of Language Model Confidence Scores |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2410.06615 |