PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Margalit, Yanki, Avram, Erni, Taig, Ran, Margalit, Oded, Cohen-Inger, Nurit
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918320494608384
author Margalit, Yanki
Avram, Erni
Taig, Ran
Margalit, Oded
Cohen-Inger, Nurit
author_facet Margalit, Yanki
Avram, Erni
Taig, Ran
Margalit, Oded
Cohen-Inger, Nurit
contents Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly outdated, and mismatch open-world deployments that depend on web retrieval and synthesis. We introduce PeerRank, a fully autonomous end-to-end evaluation framework in which models generate evaluation tasks, answer them with category-scoped live web grounding, judge peer responses and aggregate dense peer assessments into relative performance estimates, without human supervision or gold references. PeerRank treats evaluation as a multi-agent process where each model participates symmetrically as task designer, respondent, and evaluator, while removing biased judgments. In a large-scale study over 12 commercially available models and 420 autonomously generated questions, PeerRank produces stable, discriminative rankings and reveals measurable identity and presentation biases. Rankings are robust, and mean peer scores agree with Elo. We further validate PeerRank on TruthfulQA and GSM8K, where peer scores correlate with objective accuracy. Together, these results suggest that bias-aware peer evaluation with selective web-grounded answering can scale open-world LLM assessment beyond static and human curated benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02589
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
Margalit, Yanki
Avram, Erni
Taig, Ran
Margalit, Oded
Cohen-Inger, Nurit
Artificial Intelligence
Machine Learning
Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly outdated, and mismatch open-world deployments that depend on web retrieval and synthesis. We introduce PeerRank, a fully autonomous end-to-end evaluation framework in which models generate evaluation tasks, answer them with category-scoped live web grounding, judge peer responses and aggregate dense peer assessments into relative performance estimates, without human supervision or gold references. PeerRank treats evaluation as a multi-agent process where each model participates symmetrically as task designer, respondent, and evaluator, while removing biased judgments. In a large-scale study over 12 commercially available models and 420 autonomously generated questions, PeerRank produces stable, discriminative rankings and reveals measurable identity and presentation biases. Rankings are robust, and mean peer scores agree with Elo. We further validate PeerRank on TruthfulQA and GSM8K, where peer scores correlate with objective accuracy. Together, these results suggest that bias-aware peer evaluation with selective web-grounded answering can scale open-world LLM assessment beyond static and human curated benchmarks.
title PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.02589