Reference-free Evaluation Metrics for Text Generation: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ito, Takumi, van Deemter, Kees, Suzuki, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913659159052288
author Ito, Takumi
van Deemter, Kees
Suzuki, Jun
author_facet Ito, Takumi
van Deemter, Kees
Suzuki, Jun
contents A number of automatic evaluation metrics have been proposed for natural language generation systems. The most common approach to automatic evaluation is the use of a reference-based metric that compares the model's output with gold-standard references written by humans. However, it is expensive to create such references, and for some tasks, such as response generation in dialogue, creating references is not a simple matter. Therefore, various reference-free metrics have been developed in recent years. In this survey, which intends to cover the full breadth of all NLG tasks, we investigate the most commonly used approaches, their application, and their other uses beyond evaluating models. The survey concludes by highlighting some promising directions for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12011
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reference-free Evaluation Metrics for Text Generation: A Survey
Ito, Takumi
van Deemter, Kees
Suzuki, Jun
Computation and Language
A number of automatic evaluation metrics have been proposed for natural language generation systems. The most common approach to automatic evaluation is the use of a reference-based metric that compares the model's output with gold-standard references written by humans. However, it is expensive to create such references, and for some tasks, such as response generation in dialogue, creating references is not a simple matter. Therefore, various reference-free metrics have been developed in recent years. In this survey, which intends to cover the full breadth of all NLG tasks, we investigate the most commonly used approaches, their application, and their other uses beyond evaluating models. The survey concludes by highlighting some promising directions for future research.
title Reference-free Evaluation Metrics for Text Generation: A Survey
topic Computation and Language
url https://arxiv.org/abs/2501.12011