Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Xinyu, Lin, Li, Gao, Mingqi, Yin, Xunjian, Wan, Xiaojun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916425698902016
author Hu, Xinyu
Lin, Li
Gao, Mingqi
Yin, Xunjian
Wan, Xiaojun
author_facet Hu, Xinyu
Lin, Li
Gao, Mingqi
Yin, Xunjian
Wan, Xiaojun
contents The evaluation of natural language generation (NLG) tasks is a significant and longstanding research area. With the recent emergence of powerful large language models (LLMs), some studies have turned to LLM-based automatic evaluation methods, which demonstrate great potential to become a new evaluation paradigm following traditional string-based and model-based metrics. However, despite the improved performance of existing methods, they still possess some deficiencies, such as dependency on references and limited evaluation flexibility. Therefore, in this paper, we meticulously construct a large-scale NLG evaluation corpus NLG-Eval with annotations from both human and GPT-4 to alleviate the lack of relevant data in this field. Furthermore, we propose Themis, an LLM dedicated to NLG evaluation, which has been trained with our designed multi-perspective consistency verification and rating-oriented preference alignment methods. Themis can conduct flexible and interpretable evaluations without references, and it exhibits superior evaluation performance on various NLG tasks, simultaneously generalizing well to unseen tasks and surpassing other evaluation models, including GPT-4.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18365
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
Hu, Xinyu
Lin, Li
Gao, Mingqi
Yin, Xunjian
Wan, Xiaojun
Computation and Language
The evaluation of natural language generation (NLG) tasks is a significant and longstanding research area. With the recent emergence of powerful large language models (LLMs), some studies have turned to LLM-based automatic evaluation methods, which demonstrate great potential to become a new evaluation paradigm following traditional string-based and model-based metrics. However, despite the improved performance of existing methods, they still possess some deficiencies, such as dependency on references and limited evaluation flexibility. Therefore, in this paper, we meticulously construct a large-scale NLG evaluation corpus NLG-Eval with annotations from both human and GPT-4 to alleviate the lack of relevant data in this field. Furthermore, we propose Themis, an LLM dedicated to NLG evaluation, which has been trained with our designed multi-perspective consistency verification and rating-oriented preference alignment methods. Themis can conduct flexible and interpretable evaluations without references, and it exhibits superior evaluation performance on various NLG tasks, simultaneously generalizing well to unseen tasks and surpassing other evaluation models, including GPT-4.
title Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
topic Computation and Language
url https://arxiv.org/abs/2406.18365