Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Mingyang, Zheng, Mao, Xu, Chenning |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
von: Xu, Chenning, et al.
Veröffentlicht: (2026)
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
The Consensus Trap: Dissecting Subjectivity and the "Ground Truth" Illusion in Data Annotation
von: Munir, Sheza, et al.
Veröffentlicht: (2026)
von: Munir, Sheza, et al.
Veröffentlicht: (2026)
HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains
von: Li, Zheng, et al.
Veröffentlicht: (2026)
von: Li, Zheng, et al.
Veröffentlicht: (2026)
Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
von: Lai, Peng, et al.
Veröffentlicht: (2025)
von: Lai, Peng, et al.
Veröffentlicht: (2025)
Model Merging in the Era of Large Language Models: Methods, Applications, and Future Directions
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
A Survey of Query Optimization in Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
A Survey of On-Policy Distillation for Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
von: Lee, Dongryeol, et al.
Veröffentlicht: (2026)
von: Lee, Dongryeol, et al.
Veröffentlicht: (2026)
Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions
von: Wang, Dingzriui, et al.
Veröffentlicht: (2025)
von: Wang, Dingzriui, et al.
Veröffentlicht: (2025)
Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges
von: Li, Weiyuan, et al.
Veröffentlicht: (2025)
von: Li, Weiyuan, et al.
Veröffentlicht: (2025)
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
Permutation-Consensus Listwise Judging for Robust Factuality Evaluation
von: Huang, Tianyi, et al.
Veröffentlicht: (2026)
von: Huang, Tianyi, et al.
Veröffentlicht: (2026)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2024)
von: Li, Zheng, et al.
Veröffentlicht: (2024)
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
von: Sun, Xin, et al.
Veröffentlicht: (2026)
von: Sun, Xin, et al.
Veröffentlicht: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
von: Krumdick, Michael, et al.
Veröffentlicht: (2025)
von: Krumdick, Michael, et al.
Veröffentlicht: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
von: Huang, Hui, et al.
Veröffentlicht: (2024)
von: Huang, Hui, et al.
Veröffentlicht: (2024)
The Necessity of Setting Temperature in LLM-as-a-Judge
von: Li, Lujun, et al.
Veröffentlicht: (2026)
von: Li, Lujun, et al.
Veröffentlicht: (2026)
Evaluating Scoring Bias in LLM-as-a-Judge
von: Li, Qingquan, et al.
Veröffentlicht: (2025)
von: Li, Qingquan, et al.
Veröffentlicht: (2025)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
von: Zhang, Xinran
Veröffentlicht: (2026)
von: Zhang, Xinran
Veröffentlicht: (2026)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
von: Chen, Luyu, et al.
Veröffentlicht: (2025)
von: Chen, Luyu, et al.
Veröffentlicht: (2025)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
von: Yang, Bo, et al.
Veröffentlicht: (2026)
von: Yang, Bo, et al.
Veröffentlicht: (2026)
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
HY-MT1.5 Technical Report
von: Zheng, Mao, et al.
Veröffentlicht: (2025)
von: Zheng, Mao, et al.
Veröffentlicht: (2025)
Grounding LLM Reasoning with Knowledge Graphs
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2025)
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2025)
EventGround: Narrative Reasoning by Grounding to Eventuality-centric Knowledge Graphs
von: Jiayang, Cheng, et al.
Veröffentlicht: (2024)
von: Jiayang, Cheng, et al.
Veröffentlicht: (2024)
MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
von: Lavrinovics, Ernests, et al.
Veröffentlicht: (2025)
von: Lavrinovics, Ernests, et al.
Veröffentlicht: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
von: Xu, Chenning, et al.
Veröffentlicht: (2026) -
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
von: Xu, Chenning, et al.
Veröffentlicht: (2026) -
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
von: Song, Mingyang, et al.
Veröffentlicht: (2025) -
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
von: Shi, Zhichao, et al.
Veröffentlicht: (2025) -
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
von: Song, Mingyang, et al.
Veröffentlicht: (2025)