Fusion-Eval: Integrating Assistant Evaluators with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shu, Lei, Wichers, Nevan, Luo, Liangchen, Zhu, Yun, Liu, Yinxiao, Chen, Jindong, Meng, Lei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917687212376064
author Shu, Lei
Wichers, Nevan
Luo, Liangchen
Zhu, Yun
Liu, Yinxiao
Chen, Jindong
Meng, Lei
author_facet Shu, Lei
Wichers, Nevan
Luo, Liangchen
Zhu, Yun
Liu, Yinxiao
Chen, Jindong
Meng, Lei
contents Evaluating natural language systems poses significant challenges, particularly in the realms of natural language understanding and high-level reasoning. In this paper, we introduce 'Fusion-Eval', an innovative approach that leverages Large Language Models (LLMs) to integrate insights from various assistant evaluators. The LLM is given the example to evaluate along with scores from the assistant evaluators. Each of these evaluators specializes in assessing distinct aspects of responses. Fusion-Eval achieves a 0.962 system-level Kendall-Tau correlation with humans on SummEval and a 0.744 turn-level Spearman correlation on TopicalChat, which is significantly higher than baseline methods. These results highlight Fusion-Eval's significant potential in the realm of natural language system evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09204
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Fusion-Eval: Integrating Assistant Evaluators with LLMs
Shu, Lei
Wichers, Nevan
Luo, Liangchen
Zhu, Yun
Liu, Yinxiao
Chen, Jindong
Meng, Lei
Computation and Language
Artificial Intelligence
Evaluating natural language systems poses significant challenges, particularly in the realms of natural language understanding and high-level reasoning. In this paper, we introduce 'Fusion-Eval', an innovative approach that leverages Large Language Models (LLMs) to integrate insights from various assistant evaluators. The LLM is given the example to evaluate along with scores from the assistant evaluators. Each of these evaluators specializes in assessing distinct aspects of responses. Fusion-Eval achieves a 0.962 system-level Kendall-Tau correlation with humans on SummEval and a 0.744 turn-level Spearman correlation on TopicalChat, which is significantly higher than baseline methods. These results highlight Fusion-Eval's significant potential in the realm of natural language system evaluation.
title Fusion-Eval: Integrating Assistant Evaluators with LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.09204