LLM-based NLG Evaluation: Current Status and Challenges

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gao, Mingqi, Hu, Xinyu, Ruan, Jie, Pu, Xiao, Wan, Xiaojun
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913835593498624
author Gao, Mingqi
Hu, Xinyu
Ruan, Jie
Pu, Xiao
Wan, Xiaojun
author_facet Gao, Mingqi
Hu, Xinyu
Ruan, Jie
Pu, Xiao
Wan, Xiaojun
contents Evaluating natural language generation (NLG) is a vital but challenging problem in natural language processing. Traditional evaluation metrics mainly capturing content (e.g. n-gram) overlap between system outputs and references are far from satisfactory, and large language models (LLMs) such as ChatGPT have demonstrated great potential in NLG evaluation in recent years. Various automatic evaluation methods based on LLMs have been proposed, including metrics derived from LLMs, prompting LLMs, fine-tuning LLMs, and human-LLM collaborative evaluation. In this survey, we first give a taxonomy of LLM-based NLG evaluation methods, and discuss their pros and cons, respectively. Lastly, we discuss several open problems in this area and point out future research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01383
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLM-based NLG Evaluation: Current Status and Challenges
Gao, Mingqi
Hu, Xinyu
Ruan, Jie
Pu, Xiao
Wan, Xiaojun
Computation and Language
Evaluating natural language generation (NLG) is a vital but challenging problem in natural language processing. Traditional evaluation metrics mainly capturing content (e.g. n-gram) overlap between system outputs and references are far from satisfactory, and large language models (LLMs) such as ChatGPT have demonstrated great potential in NLG evaluation in recent years. Various automatic evaluation methods based on LLMs have been proposed, including metrics derived from LLMs, prompting LLMs, fine-tuning LLMs, and human-LLM collaborative evaluation. In this survey, we first give a taxonomy of LLM-based NLG evaluation methods, and discuss their pros and cons, respectively. Lastly, we discuss several open problems in this area and point out future research directions.
title LLM-based NLG Evaluation: Current Status and Challenges
topic Computation and Language
url https://arxiv.org/abs/2402.01383