Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Wei-Hsiang, Wei, Sheng-Lun, Huang, Hen-Hsen, Chen, Hsin-Hsi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915511139303424
author Lin, Wei-Hsiang
Wei, Sheng-Lun
Huang, Hen-Hsen
Chen, Hsin-Hsi
author_facet Lin, Wei-Hsiang
Wei, Sheng-Lun
Huang, Hen-Hsen
Chen, Hsin-Hsi
contents LLM-as-Judge frameworks are increasingly popular for AI evaluation, yet research findings on the relationship between models' generation and judgment abilities remain inconsistent. We investigate this relationship through systematic dataset- and instance-level analyses across 11 models and 21 diverse tasks. Despite both capabilities relying on the same underlying knowledge, our analyses reveal they are only weakly correlated, primarily due to LLMs' sensitivity to the responses being judged. To address this, we propose a self-reference-guided evaluation strategy that leverages a model's own answers as references. This approach significantly strengthens the correlation between generation and judgment abilities, offering a practical path to align these skills and providing a reliable proxy for model selection in evaluation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19880
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
Lin, Wei-Hsiang
Wei, Sheng-Lun
Huang, Hen-Hsen
Chen, Hsin-Hsi
Computation and Language
Artificial Intelligence
LLM-as-Judge frameworks are increasingly popular for AI evaluation, yet research findings on the relationship between models' generation and judgment abilities remain inconsistent. We investigate this relationship through systematic dataset- and instance-level analyses across 11 models and 21 diverse tasks. Despite both capabilities relying on the same underlying knowledge, our analyses reveal they are only weakly correlated, primarily due to LLMs' sensitivity to the responses being judged. To address this, we propose a self-reference-guided evaluation strategy that leverages a model's own answers as references. This approach significantly strengthens the correlation between generation and judgment abilities, offering a practical path to align these skills and providing a reliable proxy for model selection in evaluation tasks.
title Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.19880