Predicting generalization performance with correctness discriminators

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Yuekun, Koller, Alexander
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909618111774720
author Yao, Yuekun
Koller, Alexander
author_facet Yao, Yuekun
Koller, Alexander
contents The ability to predict an NLP model's accuracy on unseen, potentially out-of-distribution data is a prerequisite for trustworthiness. We present a novel model that establishes upper and lower bounds on the accuracy, without requiring gold labels for the unseen data. We achieve this by training a discriminator which predicts whether the output of a given sequence-to-sequence model is correct or not. We show across a variety of tagging, parsing, and semantic parsing tasks that the gold accuracy is reliably between the predicted upper and lower bounds, and that these bounds are remarkably close together.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09422
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Predicting generalization performance with correctness discriminators
Yao, Yuekun
Koller, Alexander
Computation and Language
The ability to predict an NLP model's accuracy on unseen, potentially out-of-distribution data is a prerequisite for trustworthiness. We present a novel model that establishes upper and lower bounds on the accuracy, without requiring gold labels for the unseen data. We achieve this by training a discriminator which predicts whether the output of a given sequence-to-sequence model is correct or not. We show across a variety of tagging, parsing, and semantic parsing tasks that the gold accuracy is reliably between the predicted upper and lower bounds, and that these bounds are remarkably close together.
title Predicting generalization performance with correctness discriminators
topic Computation and Language
url https://arxiv.org/abs/2311.09422