No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Krumdick, Michael, Lovering, Charles, Reddy, Varshini, Ebner, Seth, Tanner, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cost-Efficient Estimation of General Abilities Across Benchmarks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
On Finding Inconsistencies in Documents
by: Lovering, Charles J., et al.
Published: (2025)
by: Lovering, Charles J., et al.
Published: (2025)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)
by: Reddy, Varshini, et al.
Published: (2024)
Tokenization with Split Trees
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
The Effect of Scripts and Formats on LLM Numeracy
by: Reddy, Varshini, et al.
Published: (2026)
by: Reddy, Varshini, et al.
Published: (2026)
Language Model Probabilities are Not Calibrated in Numeric Contexts
by: Lovering, Charles, et al.
Published: (2024)
by: Lovering, Charles, et al.
Published: (2024)
An Analysis of Multilingual FActScore
by: Vu, Kim Trong, et al.
Published: (2024)
by: Vu, Kim Trong, et al.
Published: (2024)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
by: Reddy, Varshini, et al.
Published: (2025)
by: Reddy, Varshini, et al.
Published: (2025)
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
by: Schmidt, Craig W., et al.
Published: (2025)
by: Schmidt, Craig W., et al.
Published: (2025)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
by: Chang, Yapei, et al.
Published: (2025)
by: Chang, Yapei, et al.
Published: (2025)
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation
by: Sehanobish, Arijit, et al.
Published: (2026)
by: Sehanobish, Arijit, et al.
Published: (2026)
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
by: Sun, Xin, et al.
Published: (2026)
by: Sun, Xin, et al.
Published: (2026)
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024)
by: Schmidt, Craig W., et al.
Published: (2024)
AutoJudge: Judge Decoding Without Manual Annotation
by: Garipov, Roman, et al.
Published: (2025)
by: Garipov, Roman, et al.
Published: (2025)
MLLM-as-a-Judge for Image Safety without Human Labeling
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
by: Li, Yanran
Published: (2026)
by: Li, Yanran
Published: (2026)
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
by: Han, Steve, et al.
Published: (2025)
by: Han, Steve, et al.
Published: (2025)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
by: Wang, Victor, et al.
Published: (2025)
by: Wang, Victor, et al.
Published: (2025)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
by: Meisenbacher, Stephen, et al.
Published: (2025)
by: Meisenbacher, Stephen, et al.
Published: (2025)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
by: Zhou, Lingfeng, et al.
Published: (2025)
by: Zhou, Lingfeng, et al.
Published: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
by: Belmadani, Ikram, et al.
Published: (2026)
by: Belmadani, Ikram, et al.
Published: (2026)
Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
by: Lai, Peng, et al.
Published: (2025)
by: Lai, Peng, et al.
Published: (2025)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
A Closer Look at Claim Decomposition
by: Wanner, Miriam, et al.
Published: (2024)
by: Wanner, Miriam, et al.
Published: (2024)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
by: Zhang, Xinran
Published: (2026)
by: Zhang, Xinran
Published: (2026)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
by: Li, Weiyue, et al.
Published: (2026)
by: Li, Weiyue, et al.
Published: (2026)
Similar Items
-
Cost-Efficient Estimation of General Abilities Across Benchmarks
by: Krumdick, Michael, et al.
Published: (2026) -
On Finding Inconsistencies in Documents
by: Lovering, Charles J., et al.
Published: (2025) -
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023) -
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024) -
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)