Language Models can Evaluate Themselves via Probability Discrepancy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Tingyu, Yu, Bowen, Wu, Yuan, Chang, Yi, Zhou, Chang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913421663928320
author Xia, Tingyu
Yu, Bowen
Wu, Yuan
Chang, Yi
Zhou, Chang
author_facet Xia, Tingyu
Yu, Bowen
Wu, Yuan
Chang, Yi
Zhou, Chang
contents In this paper, we initiate our discussion by demonstrating how Large Language Models (LLMs), when tasked with responding to queries, display a more even probability distribution in their answers if they are more adept, as opposed to their less skilled counterparts. Expanding on this foundational insight, we propose a new self-evaluation method ProbDiff for assessing the efficacy of various LLMs. This approach obviates the necessity for an additional evaluation model or the dependence on external, proprietary models like GPT-4 for judgment. It uniquely utilizes the LLMs being tested to compute the probability discrepancy between the initial response and its revised versions. A higher discrepancy for a given query between two LLMs indicates a relatively weaker capability. Our findings reveal that ProbDiff achieves results on par with those obtained from evaluations based on GPT-4, spanning a range of scenarios that include natural language generation (NLG) tasks such as translation, summarization, and our proposed Xiaohongshu blog writing task, and benchmarks for LLM evaluation like AlignBench, MT-Bench, and AlpacaEval, across LLMs of varying magnitudes.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10516
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language Models can Evaluate Themselves via Probability Discrepancy
Xia, Tingyu
Yu, Bowen
Wu, Yuan
Chang, Yi
Zhou, Chang
Computation and Language
Artificial Intelligence
In this paper, we initiate our discussion by demonstrating how Large Language Models (LLMs), when tasked with responding to queries, display a more even probability distribution in their answers if they are more adept, as opposed to their less skilled counterparts. Expanding on this foundational insight, we propose a new self-evaluation method ProbDiff for assessing the efficacy of various LLMs. This approach obviates the necessity for an additional evaluation model or the dependence on external, proprietary models like GPT-4 for judgment. It uniquely utilizes the LLMs being tested to compute the probability discrepancy between the initial response and its revised versions. A higher discrepancy for a given query between two LLMs indicates a relatively weaker capability. Our findings reveal that ProbDiff achieves results on par with those obtained from evaluations based on GPT-4, spanning a range of scenarios that include natural language generation (NLG) tasks such as translation, summarization, and our proposed Xiaohongshu blog writing task, and benchmarks for LLM evaluation like AlignBench, MT-Bench, and AlpacaEval, across LLMs of varying magnitudes.
title Language Models can Evaluate Themselves via Probability Discrepancy
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.10516