League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Qianhong, Xie, Wei, Cai, Xiaofang, Wang, Enze, Ma, Shuoyoucheng, Sun, Xiaobing, Xia, Tian, Chen, Kai, Wang, Xiaofeng, Wang, Baosheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918444526469120
author Guo, Qianhong
Xie, Wei
Cai, Xiaofang
Wang, Enze
Ma, Shuoyoucheng
Sun, Xiaobing
Xia, Tian
Chen, Kai
Wang, Xiaofeng
Wang, Baosheng
author_facet Guo, Qianhong
Xie, Wei
Cai, Xiaofang
Wang, Enze
Ma, Shuoyoucheng
Sun, Xiaobing
Xia, Tian
Chen, Kai
Wang, Xiaofeng
Wang, Baosheng
contents Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination, opaque operation, and subjective preferences. To address these issues, we propose League of LLMs (LOL), a novel benchmark-free evaluation paradigm that organizes multiple LLMs into a self-governed league for multi-round mutual evaluation. LOL integrates four core criteria (dynamic, transparent, objective, and professional) to mitigate key limitations of existing paradigms. Experiments on eight mainstream LLMs in mathematics and programming demonstrate that LOL can effectively distinguish LLM capabilities while maintaining high internal ranking stability (Top-$k$ consistency $= 70.7\%$). Beyond ranking, LOL reveals empirical findings that are difficult for traditional paradigms to capture. For instance, ``memorization-based answering'' behaviors are observed in some models, and higher in-family scores are found in the OpenAI model family ($Δ= 9$, $p < 0.05$). Finally, we make our framework and code publicly available as a valuable complement to the current LLM evaluation ecosystem.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22359
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
Guo, Qianhong
Xie, Wei
Cai, Xiaofang
Wang, Enze
Ma, Shuoyoucheng
Sun, Xiaobing
Xia, Tian
Chen, Kai
Wang, Xiaofeng
Wang, Baosheng
Artificial Intelligence
Computation and Language
Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination, opaque operation, and subjective preferences. To address these issues, we propose League of LLMs (LOL), a novel benchmark-free evaluation paradigm that organizes multiple LLMs into a self-governed league for multi-round mutual evaluation. LOL integrates four core criteria (dynamic, transparent, objective, and professional) to mitigate key limitations of existing paradigms. Experiments on eight mainstream LLMs in mathematics and programming demonstrate that LOL can effectively distinguish LLM capabilities while maintaining high internal ranking stability (Top-$k$ consistency $= 70.7\%$). Beyond ranking, LOL reveals empirical findings that are difficult for traditional paradigms to capture. For instance, ``memorization-based answering'' behaviors are observed in some models, and higher in-family scores are found in the OpenAI model family ($Δ= 9$, $p < 0.05$). Finally, we make our framework and code publicly available as a valuable complement to the current LLM evaluation ecosystem.
title League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.22359