From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Dawei, Jiang, Bohan, Huang, Liangjie, Beigi, Alimohammad, Zhao, Chengshuai, Tan, Zhen, Bhattacharjee, Amrita, Jiang, Yuxuan, Chen, Canyu, Wu, Tianhao, Shu, Kai, Cheng, Lu, Liu, Huan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914063193210880
author Li, Dawei
Jiang, Bohan
Huang, Liangjie
Beigi, Alimohammad
Zhao, Chengshuai
Tan, Zhen
Bhattacharjee, Amrita
Jiang, Yuxuan
Chen, Canyu
Wu, Tianhao
Shu, Kai
Cheng, Lu
Liu, Huan
author_facet Li, Dawei
Jiang, Bohan
Huang, Liangjie
Beigi, Alimohammad
Zhao, Chengshuai
Tan, Zhen
Bhattacharjee, Amrita
Jiang, Yuxuan
Chen, Canyu
Wu, Tianhao
Shu, Kai
Cheng, Lu
Liu, Huan
contents Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or small model-based, often fall short in open-ended and dynamic scenarios. Recent advancements in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm, where LLMs are leveraged to perform scoring, ranking, or selection for various machine learning evaluation scenarios. This paper presents a comprehensive survey of LLM-based judgment and assessment, offering an in-depth overview to review this evolving field. We first provide the definition from both input and output perspectives. Then we introduce a systematic taxonomy to explore LLM-as-a-judge along three dimensions: what to judge, how to judge, and how to benchmark. Finally, we also highlight key challenges and promising future directions for this emerging area. More resources on LLM-as-a-judge are on the website: https://llm-as-a-judge.github.io and https://github.com/llm-as-a-judge/Awesome-LLM-as-a-judge.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16594
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
Li, Dawei
Jiang, Bohan
Huang, Liangjie
Beigi, Alimohammad
Zhao, Chengshuai
Tan, Zhen
Bhattacharjee, Amrita
Jiang, Yuxuan
Chen, Canyu
Wu, Tianhao
Shu, Kai
Cheng, Lu
Liu, Huan
Artificial Intelligence
Computation and Language
Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or small model-based, often fall short in open-ended and dynamic scenarios. Recent advancements in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm, where LLMs are leveraged to perform scoring, ranking, or selection for various machine learning evaluation scenarios. This paper presents a comprehensive survey of LLM-based judgment and assessment, offering an in-depth overview to review this evolving field. We first provide the definition from both input and output perspectives. Then we introduce a systematic taxonomy to explore LLM-as-a-judge along three dimensions: what to judge, how to judge, and how to benchmark. Finally, we also highlight key challenges and promising future directions for this emerging area. More resources on LLM-as-a-judge are on the website: https://llm-as-a-judge.github.io and https://github.com/llm-as-a-judge/Awesome-LLM-as-a-judge.
title From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.16594