PiCO: Peer Review in LLMs based on the Consistency Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ning, Kun-Peng, Yang, Shuo, Liu, Yu-Yang, Yao, Jia-Yu, Liu, Zhen-Hui, Tian, Yong-Hong, Song, Yibing, Yuan, Li
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910837695840256
author Ning, Kun-Peng
Yang, Shuo
Liu, Yu-Yang
Yao, Jia-Yu
Liu, Zhen-Hui
Tian, Yong-Hong
Song, Yibing
Yuan, Li
author_facet Ning, Kun-Peng
Yang, Shuo
Liu, Yu-Yang
Yao, Jia-Yu
Liu, Zhen-Hui
Tian, Yong-Hong
Song, Yibing
Yuan, Li
contents Existing large language models (LLMs) evaluation methods typically focus on testing the performance on some closed-environment and domain-specific benchmarks with human annotations. In this paper, we explore a novel unsupervised evaluation direction, utilizing peer-review mechanisms to measure LLMs automatically. In this setting, both open-source and closed-source LLMs lie in the same environment, capable of answering unlabeled questions and evaluating each other, where each LLM's response score is jointly determined by other anonymous ones. To obtain the ability hierarchy among these models, we assign each LLM a learnable capability parameter to adjust the final ranking. We formalize it as a constrained optimization problem, intending to maximize the consistency of each LLM's capabilities and scores. The key assumption behind is that high-level LLM can evaluate others' answers more accurately than low-level ones, while higher-level LLM can also achieve higher response scores. Moreover, we propose three metrics called PEN, CIN, and LIS to evaluate the gap in aligning human rankings. We perform experiments on multiple datasets with these metrics, validating the effectiveness of the proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01830
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PiCO: Peer Review in LLMs based on the Consistency Optimization
Ning, Kun-Peng
Yang, Shuo
Liu, Yu-Yang
Yao, Jia-Yu
Liu, Zhen-Hui
Tian, Yong-Hong
Song, Yibing
Yuan, Li
Computation and Language
Artificial Intelligence
Machine Learning
Existing large language models (LLMs) evaluation methods typically focus on testing the performance on some closed-environment and domain-specific benchmarks with human annotations. In this paper, we explore a novel unsupervised evaluation direction, utilizing peer-review mechanisms to measure LLMs automatically. In this setting, both open-source and closed-source LLMs lie in the same environment, capable of answering unlabeled questions and evaluating each other, where each LLM's response score is jointly determined by other anonymous ones. To obtain the ability hierarchy among these models, we assign each LLM a learnable capability parameter to adjust the final ranking. We formalize it as a constrained optimization problem, intending to maximize the consistency of each LLM's capabilities and scores. The key assumption behind is that high-level LLM can evaluate others' answers more accurately than low-level ones, while higher-level LLM can also achieve higher response scores. Moreover, we propose three metrics called PEN, CIN, and LIS to evaluate the gap in aligning human rankings. We perform experiments on multiple datasets with these metrics, validating the effectiveness of the proposed approach.
title PiCO: Peer Review in LLMs based on the Consistency Optimization
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2402.01830