Schroedinger's Threshold: When the AUC doesn't predict Accuracy

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteur principal: Opitz, Juri
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911887823732736
author Opitz, Juri
author_facet Opitz, Juri
contents The Area Under Curve measure (AUC) seems apt to evaluate and compare diverse models, possibly without calibration. An important example of AUC application is the evaluation and benchmarking of models that predict faithfulness of generated text. But we show that the AUC yields an academic and optimistic notion of accuracy that can misalign with the actual accuracy observed in application, yielding significant changes in benchmark rankings. To paint a more realistic picture of downstream model performance (and prepare a model for actual application), we explore different calibration modes, testing calibration data and method.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03344
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Schroedinger's Threshold: When the AUC doesn't predict Accuracy
Opitz, Juri
Computation and Language
The Area Under Curve measure (AUC) seems apt to evaluate and compare diverse models, possibly without calibration. An important example of AUC application is the evaluation and benchmarking of models that predict faithfulness of generated text. But we show that the AUC yields an academic and optimistic notion of accuracy that can misalign with the actual accuracy observed in application, yielding significant changes in benchmark rankings. To paint a more realistic picture of downstream model performance (and prepare a model for actual application), we explore different calibration modes, testing calibration data and method.
title Schroedinger's Threshold: When the AUC doesn't predict Accuracy
topic Computation and Language
url https://arxiv.org/abs/2404.03344