A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Xinyu, Gao, Mingqi, Lin, Li, Yu, Zhenghan, Wan, Xiaojun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918125394460672
author Hu, Xinyu
Gao, Mingqi
Lin, Li
Yu, Zhenghan
Wan, Xiaojun
author_facet Hu, Xinyu
Gao, Mingqi
Lin, Li
Yu, Zhenghan
Wan, Xiaojun
contents In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling human ratings and ambiguous selections of correlation measures, which undermine the effectiveness of meta-evaluation. In this work, we propose a dual-perspective NLG meta-evaluation framework that focuses on different evaluation capabilities, thereby providing better interpretability. In addition, we introduce a method of automatically constructing the corresponding benchmarks without requiring new human annotations. Furthermore, we conduct experiments with 16 representative LLMs as the evaluators based on our proposed framework, comprehensively analyzing their evaluation performance from different perspectives.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12052
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
Hu, Xinyu
Gao, Mingqi
Lin, Li
Yu, Zhenghan
Wan, Xiaojun
Computation and Language
In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling human ratings and ambiguous selections of correlation measures, which undermine the effectiveness of meta-evaluation. In this work, we propose a dual-perspective NLG meta-evaluation framework that focuses on different evaluation capabilities, thereby providing better interpretability. In addition, we introduce a method of automatically constructing the corresponding benchmarks without requiring new human annotations. Furthermore, we conduct experiments with 16 representative LLMs as the evaluators based on our proposed framework, comprehensively analyzing their evaluation performance from different perspectives.
title A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
topic Computation and Language
url https://arxiv.org/abs/2502.12052