Saved in:
Bibliographic Details
Main Authors: Nguyen, Bang, Yu, Mengxia, Huang, Yun, Jiang, Meng
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.12242
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909343795904512
author Nguyen, Bang
Yu, Mengxia
Huang, Yun
Jiang, Meng
author_facet Nguyen, Bang
Yu, Mengxia
Huang, Yun
Jiang, Meng
contents Reference-based metrics such as BLEU and BERTScore are widely used to evaluate question generation (QG). In this study, on QG benchmarks such as SQuAD and HotpotQA, we find that using human-written references cannot guarantee the effectiveness of the reference-based metrics. Most QG benchmarks have only one reference; we replicate the annotation process and collect another reference. A good metric is expected to grade a human-validated question no worse than generated questions. However, the results of reference-based metrics on our newly collected reference disproved the metrics themselves. We propose a reference-free metric consisted of multi-dimensional criteria such as naturalness, answerability, and complexity, utilizing large language models. These criteria are not constrained to the syntactic or semantic of a single reference question, and the metric does not require a diverse set of references. Experiments reveal that our metric accurately distinguishes between high-quality questions and flawed ones, and achieves state-of-the-art alignment with human judgment.
format Preprint
id arxiv_https___arxiv_org_abs_2403_12242
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reference-based Metrics Disprove Themselves in Question Generation
Nguyen, Bang
Yu, Mengxia
Huang, Yun
Jiang, Meng
Computation and Language
Artificial Intelligence
Machine Learning
Reference-based metrics such as BLEU and BERTScore are widely used to evaluate question generation (QG). In this study, on QG benchmarks such as SQuAD and HotpotQA, we find that using human-written references cannot guarantee the effectiveness of the reference-based metrics. Most QG benchmarks have only one reference; we replicate the annotation process and collect another reference. A good metric is expected to grade a human-validated question no worse than generated questions. However, the results of reference-based metrics on our newly collected reference disproved the metrics themselves. We propose a reference-free metric consisted of multi-dimensional criteria such as naturalness, answerability, and complexity, utilizing large language models. These criteria are not constrained to the syntactic or semantic of a single reference question, and the metric does not require a diverse set of references. Experiments reveal that our metric accurately distinguishes between high-quality questions and flawed ones, and achieves state-of-the-art alignment with human judgment.
title Reference-based Metrics Disprove Themselves in Question Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2403.12242