QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Weiping, Wei, Bifan, Hu, Jianxiang, Cai, Zhongmin, Liu, Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912066034466816
author Fu, Weiping
Wei, Bifan
Hu, Jianxiang
Cai, Zhongmin
Liu, Jun
author_facet Fu, Weiping
Wei, Bifan
Hu, Jianxiang
Cai, Zhongmin
Liu, Jun
contents Automatically generated questions often suffer from problems such as unclear expression or factual inaccuracies, requiring a reliable and comprehensive evaluation of their quality. Human evaluation is widely used in the field of question generation (QG) and serves as the gold standard for automatic metrics. However, there is a lack of unified human evaluation criteria, which hampers consistent and reliable evaluations of both QG models and automatic metrics. To address this, we propose QGEval, a multi-dimensional Evaluation benchmark for Question Generation, which evaluates both generated questions and existing automatic metrics across 7 dimensions: fluency, clarity, conciseness, relevance, consistency, answerability, and answer consistency. We demonstrate the appropriateness of these dimensions by examining their correlations and distinctions. Through consistent evaluations of QG models and automatic metrics with QGEval, we find that 1) most QG models perform unsatisfactorily in terms of answerability and answer consistency, and 2) existing metrics fail to align well with human judgments when evaluating generated questions across the 7 dimensions. We expect this work to foster the development of both QG technologies and their evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05707
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
Fu, Weiping
Wei, Bifan
Hu, Jianxiang
Cai, Zhongmin
Liu, Jun
Computation and Language
Artificial Intelligence
Automatically generated questions often suffer from problems such as unclear expression or factual inaccuracies, requiring a reliable and comprehensive evaluation of their quality. Human evaluation is widely used in the field of question generation (QG) and serves as the gold standard for automatic metrics. However, there is a lack of unified human evaluation criteria, which hampers consistent and reliable evaluations of both QG models and automatic metrics. To address this, we propose QGEval, a multi-dimensional Evaluation benchmark for Question Generation, which evaluates both generated questions and existing automatic metrics across 7 dimensions: fluency, clarity, conciseness, relevance, consistency, answerability, and answer consistency. We demonstrate the appropriateness of these dimensions by examining their correlations and distinctions. Through consistent evaluations of QG models and automatic metrics with QGEval, we find that 1) most QG models perform unsatisfactorily in terms of answerability and answer consistency, and 2) existing metrics fail to align well with human judgments when evaluating generated questions across the 7 dimensions. We expect this work to foster the development of both QG technologies and their evaluation.
title QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.05707