How to Select Datapoints for Efficient Human Evaluation of NLG Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zouhar, Vilém, Cui, Peng, Sachan, Mrinmaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915314941296640
author Zouhar, Vilém
Cui, Peng
Sachan, Mrinmaya
author_facet Zouhar, Vilém
Cui, Peng
Sachan, Mrinmaya
contents Human evaluation is the gold standard for evaluating text generation models. However, it is expensive. In order to fit budgetary constraints, a random subset of the test data is often chosen in practice for human evaluation. However, randomly selected data may not accurately represent test performance, making this approach economically inefficient for model comparison. Thus, in this work, we develop and analyze a suite of selectors to get the most informative datapoints for human evaluation, taking the evaluation costs into account. We show that selectors based on variance in automated metric scores, diversity in model outputs, or Item Response Theory outperform random selection. We further develop an approach to distill these selectors to the scenario where the model outputs are not yet available. In particular, we introduce source-based estimators, which predict item usefulness for human evaluation just based on the source texts. We demonstrate the efficacy of our selectors in two common NLG tasks, machine translation and summarization, and show that only $\sim$70\% of the test data is needed to produce the same evaluation result as the entire data.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18251
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How to Select Datapoints for Efficient Human Evaluation of NLG Models?
Zouhar, Vilém
Cui, Peng
Sachan, Mrinmaya
Computation and Language
Human evaluation is the gold standard for evaluating text generation models. However, it is expensive. In order to fit budgetary constraints, a random subset of the test data is often chosen in practice for human evaluation. However, randomly selected data may not accurately represent test performance, making this approach economically inefficient for model comparison. Thus, in this work, we develop and analyze a suite of selectors to get the most informative datapoints for human evaluation, taking the evaluation costs into account. We show that selectors based on variance in automated metric scores, diversity in model outputs, or Item Response Theory outperform random selection. We further develop an approach to distill these selectors to the scenario where the model outputs are not yet available. In particular, we introduce source-based estimators, which predict item usefulness for human evaluation just based on the source texts. We demonstrate the efficacy of our selectors in two common NLG tasks, machine translation and summarization, and show that only $\sim$70\% of the test data is needed to produce the same evaluation result as the entire data.
title How to Select Datapoints for Efficient Human Evaluation of NLG Models?
topic Computation and Language
url https://arxiv.org/abs/2501.18251