Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhuochun, Zhang, Yong, Li, Ming, Ji, Yuelyu, Zeng, Yiming, Cheng, Ning, Zhu, Yun, Wang, Yanmeng, Wang, Shaojun, Xiao, Jing, He, Daqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering
by: Ji, Yuelyu, et al.
Published: (2026)
by: Ji, Yuelyu, et al.
Published: (2026)
Learning from Committee: Reasoning Distillation from a Mixture of Teachers with Peer-Review
by: Li, Zhuochun, et al.
Published: (2024)
by: Li, Zhuochun, et al.
Published: (2024)
Memory-Aware and Uncertainty-Guided Retrieval for Multi-Hop Question Answering
by: Ji, Yuelyu, et al.
Published: (2025)
by: Ji, Yuelyu, et al.
Published: (2025)
ReasoningRank: Teaching Student Models to Rank through Reasoning-Based Knowledge Distillation
by: Ji, Yuelyu, et al.
Published: (2024)
by: Ji, Yuelyu, et al.
Published: (2024)
Retrieval--Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends
by: Ji, Yuelyu, et al.
Published: (2026)
by: Ji, Yuelyu, et al.
Published: (2026)
Curriculum Guided Reinforcement Learning for Efficient Multi Hop Retrieval Augmented Generation
by: Ji, Yuelyu, et al.
Published: (2025)
by: Ji, Yuelyu, et al.
Published: (2025)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
by: Wang, Yutong, et al.
Published: (2025)
by: Wang, Yutong, et al.
Published: (2025)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
by: Zhang, Yong, et al.
Published: (2025)
by: Zhang, Yong, et al.
Published: (2025)
RAG-RLRC-LaySum at BioLaySumm: Integrating Retrieval-Augmented Generation and Readability Control for Layman Summarization of Biomedical Texts
by: Ji, Yuelyu, et al.
Published: (2024)
by: Ji, Yuelyu, et al.
Published: (2024)
Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models
by: Liu, Kainan, et al.
Published: (2026)
by: Liu, Kainan, et al.
Published: (2026)
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection
by: Hossain, Akram, et al.
Published: (2026)
by: Hossain, Akram, et al.
Published: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
SSPO: Subsentence-level Policy Optimization
by: Yang, Kun, et al.
Published: (2025)
by: Yang, Kun, et al.
Published: (2025)
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
by: Li, Chenyu, et al.
Published: (2026)
by: Li, Chenyu, et al.
Published: (2026)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
by: Lai, Peng, et al.
Published: (2026)
by: Lai, Peng, et al.
Published: (2026)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
by: Lai, Peng, et al.
Published: (2025)
by: Lai, Peng, et al.
Published: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
A Survey on LLM-as-a-Judge
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Self‐Knowledge and the Capacity to Judge
by: Matthew Parrott
Published: (2025)
by: Matthew Parrott
Published: (2025)
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
by: He, Peng, et al.
Published: (2026)
by: He, Peng, et al.
Published: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
by: Belmadani, Ikram, et al.
Published: (2026)
by: Belmadani, Ikram, et al.
Published: (2026)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
by: Soumik, Sadman Kabir
Published: (2026)
by: Soumik, Sadman Kabir
Published: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
by: Yuan, Tongxin, et al.
Published: (2024)
by: Yuan, Tongxin, et al.
Published: (2024)
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
by: Dey, Soumik, et al.
Published: (2025)
by: Dey, Soumik, et al.
Published: (2025)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
by: Saha, Swarnadeep, et al.
Published: (2025)
by: Saha, Swarnadeep, et al.
Published: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
by: Duo, Jiangshan, et al.
Published: (2026)
by: Duo, Jiangshan, et al.
Published: (2026)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
Efficient Inference for Noisy LLM-as-a-Judge Evaluation
by: Chen, Yiqun T, et al.
Published: (2026)
by: Chen, Yiqun T, et al.
Published: (2026)
Evaluating Scoring Bias in LLM-as-a-Judge
by: Li, Qingquan, et al.
Published: (2025)
by: Li, Qingquan, et al.
Published: (2025)
Learning to Adapt to Low-Resource Paraphrase Generation
by: Li, Zhigen, et al.
Published: (2024)
by: Li, Zhigen, et al.
Published: (2024)
LLM-as-Judge on a Budget
by: Saha, Aadirupa, et al.
Published: (2026)
by: Saha, Aadirupa, et al.
Published: (2026)
Similar Items
-
StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering
by: Ji, Yuelyu, et al.
Published: (2026) -
Learning from Committee: Reasoning Distillation from a Mixture of Teachers with Peer-Review
by: Li, Zhuochun, et al.
Published: (2024) -
Memory-Aware and Uncertainty-Guided Retrieval for Multi-Hop Question Answering
by: Ji, Yuelyu, et al.
Published: (2025) -
ReasoningRank: Teaching Student Models to Rank through Reasoning-Based Knowledge Distillation
by: Ji, Yuelyu, et al.
Published: (2024) -
Retrieval--Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends
by: Ji, Yuelyu, et al.
Published: (2026)