An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Hui, Bu, Xingyuan, Zhou, Hongli, Qu, Yingqi, Liu, Jing, Yang, Muyun, Xu, Bing, Zhao, Tiejun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908385645953024
author Huang, Hui
Bu, Xingyuan
Zhou, Hongli
Qu, Yingqi
Liu, Jing
Yang, Muyun
Xu, Bing
Zhao, Tiejun
author_facet Huang, Hui
Bu, Xingyuan
Zhou, Hongli
Qu, Yingqi
Liu, Jing
Yang, Muyun
Xu, Bing
Zhao, Tiejun
contents Recently, there has been a growing trend of utilizing Large Language Model (LLM) to evaluate the quality of other LLMs. Many studies have fine-tuned judge models based on open-source LLMs for evaluation. While the fine-tuned judge models are claimed to achieve comparable evaluation capability with GPT-4, in this work, we conduct an empirical study of LLM-as-a-Judge. Our findings indicate that although the fine-tuned judge models achieve high performance on in-domain test sets, even surpassing GPT-4, they underperform GPT-4 across several dimensions, including generalizability, fairness and adaptability. We also reveal that the fine-tuned judge model inherently operates as a task-specific classifier, consequently imposing the limitations.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Huang, Hui
Bu, Xingyuan
Zhou, Hongli
Qu, Yingqi
Liu, Jing
Yang, Muyun
Xu, Bing
Zhao, Tiejun
Computation and Language
Recently, there has been a growing trend of utilizing Large Language Model (LLM) to evaluate the quality of other LLMs. Many studies have fine-tuned judge models based on open-source LLMs for evaluation. While the fine-tuned judge models are claimed to achieve comparable evaluation capability with GPT-4, in this work, we conduct an empirical study of LLM-as-a-Judge. Our findings indicate that although the fine-tuned judge models achieve high performance on in-domain test sets, even surpassing GPT-4, they underperform GPT-4 across several dimensions, including generalizability, fairness and adaptability. We also reveal that the fine-tuned judge model inherently operates as a task-specific classifier, consequently imposing the limitations.
title An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
topic Computation and Language
url https://arxiv.org/abs/2403.02839