LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Son, Guijin, Ko, Hyunwoo, Lee, Hoyoung, Kim, Yewon, Hong, Seunghyeok
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912054910124032
author Son, Guijin
Ko, Hyunwoo
Lee, Hoyoung
Kim, Yewon
Hong, Seunghyeok
author_facet Son, Guijin
Ko, Hyunwoo
Lee, Hoyoung
Kim, Yewon
Hong, Seunghyeok
contents LLM-as-a-Judge and reward models are widely used alternatives of multiple-choice questions or human annotators for large language model (LLM) evaluation. Their efficacy shines in evaluating long-form responses, serving a critical role as evaluators of leaderboards and as proxies to align LLMs via reinforcement learning. However, despite their popularity, their effectiveness in diverse contexts, such as non-English prompts, factual verification, or challenging questions, remains unexplored. In this paper, we conduct a comprehensive analysis of automated evaluators, reporting several key findings on their behavior. First, we discover that English evaluation capabilities significantly influence language-specific evaluation capabilities, often more than the language proficiency itself, enabling evaluators trained in English to easily transfer their skills to other languages. Second, we identify critical shortcomings, where LLMs fail to detect and penalize errors, such as factual inaccuracies, cultural misrepresentations, and the presence of unwanted language. Finally, we find that state-of-the-art evaluators struggle with challenging prompts, in either English or Korean, underscoring their limitations in assessing or generating complex reasoning questions. We release the dataset and codes used.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11239
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
Son, Guijin
Ko, Hyunwoo
Lee, Hoyoung
Kim, Yewon
Hong, Seunghyeok
Computation and Language
LLM-as-a-Judge and reward models are widely used alternatives of multiple-choice questions or human annotators for large language model (LLM) evaluation. Their efficacy shines in evaluating long-form responses, serving a critical role as evaluators of leaderboards and as proxies to align LLMs via reinforcement learning. However, despite their popularity, their effectiveness in diverse contexts, such as non-English prompts, factual verification, or challenging questions, remains unexplored. In this paper, we conduct a comprehensive analysis of automated evaluators, reporting several key findings on their behavior. First, we discover that English evaluation capabilities significantly influence language-specific evaluation capabilities, often more than the language proficiency itself, enabling evaluators trained in English to easily transfer their skills to other languages. Second, we identify critical shortcomings, where LLMs fail to detect and penalize errors, such as factual inaccuracies, cultural misrepresentations, and the presence of unwanted language. Finally, we find that state-of-the-art evaluators struggle with challenging prompts, in either English or Korean, underscoring their limitations in assessing or generating complex reasoning questions. We release the dataset and codes used.
title LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
topic Computation and Language
url https://arxiv.org/abs/2409.11239