Reference-based Metrics Disprove Themselves in Question Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nguyen, Bang, Yu, Mengxia, Huang, Yun, Jiang, Meng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909343795904512
author Nguyen, Bang
Yu, Mengxia
Huang, Yun
Jiang, Meng
author_facet Nguyen, Bang
Yu, Mengxia
Huang, Yun
Jiang, Meng
contents Reference-based metrics such as BLEU and BERTScore are widely used to evaluate question generation (QG). In this study, on QG benchmarks such as SQuAD and HotpotQA, we find that using human-written references cannot guarantee the effectiveness of the reference-based metrics. Most QG benchmarks have only one reference; we replicate the annotation process and collect another reference. A good metric is expected to grade a human-validated question no worse than generated questions. However, the results of reference-based metrics on our newly collected reference disproved the metrics themselves. We propose a reference-free metric consisted of multi-dimensional criteria such as naturalness, answerability, and complexity, utilizing large language models. These criteria are not constrained to the syntactic or semantic of a single reference question, and the metric does not require a diverse set of references. Experiments reveal that our metric accurately distinguishes between high-quality questions and flawed ones, and achieves state-of-the-art alignment with human judgment.
format Preprint
id arxiv_https___arxiv_org_abs_2403_12242
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reference-based Metrics Disprove Themselves in Question Generation
Nguyen, Bang
Yu, Mengxia
Huang, Yun
Jiang, Meng
Computation and Language
Artificial Intelligence
Machine Learning
Reference-based metrics such as BLEU and BERTScore are widely used to evaluate question generation (QG). In this study, on QG benchmarks such as SQuAD and HotpotQA, we find that using human-written references cannot guarantee the effectiveness of the reference-based metrics. Most QG benchmarks have only one reference; we replicate the annotation process and collect another reference. A good metric is expected to grade a human-validated question no worse than generated questions. However, the results of reference-based metrics on our newly collected reference disproved the metrics themselves. We propose a reference-free metric consisted of multi-dimensional criteria such as naturalness, answerability, and complexity, utilizing large language models. These criteria are not constrained to the syntactic or semantic of a single reference question, and the metric does not require a diverse set of references. Experiments reveal that our metric accurately distinguishes between high-quality questions and flawed ones, and achieves state-of-the-art alignment with human judgment.
title Reference-based Metrics Disprove Themselves in Question Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2403.12242