When LLM Judge Scores Look Good but Best-of-N Decisions Fail
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Landesberg, Eddie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Stability Fails: Hidden Failure Modes Of LLMS in Data-Constrained Scientific Decision-Making
von: Riasat, Nazia
Veröffentlicht: (2026)
von: Riasat, Nazia
Veröffentlicht: (2026)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
von: Salinas, David, et al.
Veröffentlicht: (2025)
von: Salinas, David, et al.
Veröffentlicht: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
von: Mehrafarin, Houman, et al.
Veröffentlicht: (2026)
von: Mehrafarin, Houman, et al.
Veröffentlicht: (2026)
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
von: Wang, Zehong, et al.
Veröffentlicht: (2026)
von: Wang, Zehong, et al.
Veröffentlicht: (2026)
Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
von: Landesberg, Eddie, et al.
Veröffentlicht: (2025)
von: Landesberg, Eddie, et al.
Veröffentlicht: (2025)
Best-of-N Jailbreaking
von: Hughes, John, et al.
Veröffentlicht: (2024)
von: Hughes, John, et al.
Veröffentlicht: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
When Bad Data Leads to Good Models
von: Li, Kenneth, et al.
Veröffentlicht: (2025)
von: Li, Kenneth, et al.
Veröffentlicht: (2025)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
von: Guo, Jizhou, et al.
Veröffentlicht: (2025)
von: Guo, Jizhou, et al.
Veröffentlicht: (2025)
Variational Best-of-N Alignment
von: Amini, Afra, et al.
Veröffentlicht: (2024)
von: Amini, Afra, et al.
Veröffentlicht: (2024)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
AdaBoN: Adaptive Best-of-N Alignment
von: Raman, Vinod, et al.
Veröffentlicht: (2025)
von: Raman, Vinod, et al.
Veröffentlicht: (2025)
Learning Generative Selection for Best-of-N
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2026)
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2026)
Majority of the Bests: Improving Best-of-N via Bootstrapping
von: Rakhsha, Amin, et al.
Veröffentlicht: (2025)
von: Rakhsha, Amin, et al.
Veröffentlicht: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Investigating Non-Transitivity in LLM-as-a-Judge
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
BOND: Aligning LLMs with Best-of-N Distillation
von: Sessa, Pier Giuseppe, et al.
Veröffentlicht: (2024)
von: Sessa, Pier Giuseppe, et al.
Veröffentlicht: (2024)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
von: Mazumder, Aritra, et al.
Veröffentlicht: (2026)
von: Mazumder, Aritra, et al.
Veröffentlicht: (2026)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
von: Gera, Ariel, et al.
Veröffentlicht: (2024)
von: Gera, Ariel, et al.
Veröffentlicht: (2024)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
von: Hallaç, İbrahim Rıza, et al.
Veröffentlicht: (2026)
von: Hallaç, İbrahim Rıza, et al.
Veröffentlicht: (2026)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
von: Gupta, Manan, et al.
Veröffentlicht: (2026)
von: Gupta, Manan, et al.
Veröffentlicht: (2026)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
von: Hu, Renjun, et al.
Veröffentlicht: (2025)
von: Hu, Renjun, et al.
Veröffentlicht: (2025)
Compliance-Scored Best-of-N Guardrail Orchestration for Multimodal Document Generation in Payments Dispute Defense
von: Sundar, Nataraj Agaram, et al.
Veröffentlicht: (2026)
von: Sundar, Nataraj Agaram, et al.
Veröffentlicht: (2026)
PredictaBoard: Benchmarking LLM Score Predictability
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2025)
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2025)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
von: Zhang, Wenbo, et al.
Veröffentlicht: (2026)
von: Zhang, Wenbo, et al.
Veröffentlicht: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024)
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
Towards Best Practices for Open Datasets for LLM Training
von: Baack, Stefan, et al.
Veröffentlicht: (2025)
von: Baack, Stefan, et al.
Veröffentlicht: (2025)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
When Stability Fails: Hidden Failure Modes Of LLMS in Data-Constrained Scientific Decision-Making
von: Riasat, Nazia
Veröffentlicht: (2026) -
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025) -
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
von: Salinas, David, et al.
Veröffentlicht: (2025) -
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026) -
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
von: Mehrafarin, Houman, et al.
Veröffentlicht: (2026)