Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhuochun, Zhang, Yong, Li, Ming, Ji, Yuelyu, Zeng, Yiming, Cheng, Ning, Zhu, Yun, Wang, Yanmeng, Wang, Shaojun, Xiao, Jing, He, Daqing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918314775674880
author Li, Zhuochun
Zhang, Yong
Li, Ming
Ji, Yuelyu
Zeng, Yiming
Cheng, Ning
Zhu, Yun
Wang, Yanmeng
Wang, Shaojun
Xiao, Jing
He, Daqing
author_facet Li, Zhuochun
Zhang, Yong
Li, Ming
Ji, Yuelyu
Zeng, Yiming
Cheng, Ning
Zhu, Yun
Wang, Yanmeng
Wang, Shaojun
Xiao, Jing
He, Daqing
contents Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this "LLM-as-a-Judge" paradigm is costly, opaque, and sensitive to prompt design. In this work, we investigate whether smaller models can serve as efficient evaluators by leveraging internal representations instead of surface generation. We uncover a consistent empirical pattern: small LMs, despite with weak generative ability, encode rich evaluative signals in their hidden states. This motivates us to propose the Semantic Capacity Asymmetry Hypothesis: evaluation requires significantly less semantic capacity than generation and can be grounded in intermediate representations, suggesting that evaluation does not necessarily need to rely on large-scale generative models but can instead leverage latent features from smaller ones. Our findings motivate a paradigm shift from LLM-as-a-Judge to Representation-as-a-Judge, a decoding-free evaluation strategy that probes internal model structure rather than relying on prompted output. We instantiate this paradigm through INSPECTOR, a probing-based framework that predicts aspect-level evaluation scores from small model representations. Experiments on reasoning benchmarks (GSM8K, MATH, GPQA) show that INSPECTOR substantially outperforms prompting-based small LMs and closely approximates full LLM judges, while offering a more efficient, reliable, and interpretable alternative for scalable evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22588
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
Li, Zhuochun
Zhang, Yong
Li, Ming
Ji, Yuelyu
Zeng, Yiming
Cheng, Ning
Zhu, Yun
Wang, Yanmeng
Wang, Shaojun
Xiao, Jing
He, Daqing
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this "LLM-as-a-Judge" paradigm is costly, opaque, and sensitive to prompt design. In this work, we investigate whether smaller models can serve as efficient evaluators by leveraging internal representations instead of surface generation. We uncover a consistent empirical pattern: small LMs, despite with weak generative ability, encode rich evaluative signals in their hidden states. This motivates us to propose the Semantic Capacity Asymmetry Hypothesis: evaluation requires significantly less semantic capacity than generation and can be grounded in intermediate representations, suggesting that evaluation does not necessarily need to rely on large-scale generative models but can instead leverage latent features from smaller ones. Our findings motivate a paradigm shift from LLM-as-a-Judge to Representation-as-a-Judge, a decoding-free evaluation strategy that probes internal model structure rather than relying on prompted output. We instantiate this paradigm through INSPECTOR, a probing-based framework that predicts aspect-level evaluation scores from small model representations. Experiments on reasoning benchmarks (GSM8K, MATH, GPQA) show that INSPECTOR substantially outperforms prompting-based small LMs and closely approximates full LLM judges, while offering a more efficient, reliable, and interpretable alternative for scalable evaluation.
title Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.22588