Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
Fuente:
arXiv
Guardado en:
| Autores principales: | Gupta, Manan, Kumar, Dhruv |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Context Over Content: Exposing Evaluation Faking in Automated Judges
por: Gupta, Manan, et al.
Publicado: (2026)
por: Gupta, Manan, et al.
Publicado: (2026)
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
por: Gupta, Manan, et al.
Publicado: (2026)
por: Gupta, Manan, et al.
Publicado: (2026)
Investigating Non-Transitivity in LLM-as-a-Judge
por: Xu, Yi, et al.
Publicado: (2025)
por: Xu, Yi, et al.
Publicado: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
por: Xu, Austin, et al.
Publicado: (2025)
por: Xu, Austin, et al.
Publicado: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
por: Hong, Yihan, et al.
Publicado: (2026)
por: Hong, Yihan, et al.
Publicado: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction
por: Ye, Yuanchang, et al.
Publicado: (2025)
por: Ye, Yuanchang, et al.
Publicado: (2025)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
por: Liu, Zijun, et al.
Publicado: (2024)
por: Liu, Zijun, et al.
Publicado: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
por: Zhong, Tianle, et al.
Publicado: (2026)
por: Zhong, Tianle, et al.
Publicado: (2026)
MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media
por: Pramanik, Souvik, et al.
Publicado: (2026)
por: Pramanik, Souvik, et al.
Publicado: (2026)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
por: Yang, Jinming, et al.
Publicado: (2026)
por: Yang, Jinming, et al.
Publicado: (2026)
Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories
por: Hamilton, Sil, et al.
Publicado: (2026)
por: Hamilton, Sil, et al.
Publicado: (2026)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
por: Li, Zhuochun, et al.
Publicado: (2026)
por: Li, Zhuochun, et al.
Publicado: (2026)
Diagnosing Structural Failures in LLM-Based Evidence Extraction for Meta-Analysis
por: Tan, Zhiyin, et al.
Publicado: (2026)
por: Tan, Zhiyin, et al.
Publicado: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
por: Gera, Ariel, et al.
Publicado: (2024)
por: Gera, Ariel, et al.
Publicado: (2024)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
por: Salinas, David, et al.
Publicado: (2025)
por: Salinas, David, et al.
Publicado: (2025)
Mitigating LLM Hallucinations via Conformal Abstention
por: Yadkori, Yasin Abbasi, et al.
Publicado: (2024)
por: Yadkori, Yasin Abbasi, et al.
Publicado: (2024)
Set-LLM: A Permutation-Invariant LLM
por: Egressy, Beni, et al.
Publicado: (2025)
por: Egressy, Beni, et al.
Publicado: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
por: Liu, Yixin, et al.
Publicado: (2026)
por: Liu, Yixin, et al.
Publicado: (2026)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
por: Yehudai, Asaf, et al.
Publicado: (2025)
por: Yehudai, Asaf, et al.
Publicado: (2025)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
por: Hu, Renjun, et al.
Publicado: (2025)
por: Hu, Renjun, et al.
Publicado: (2025)
ALAS: Autonomous Learning Agent for Self-Updating Language Models
por: Atreja, Dhruv
Publicado: (2025)
por: Atreja, Dhruv
Publicado: (2025)
Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibration
por: Mouzouni, Charafeddine
Publicado: (2026)
por: Mouzouni, Charafeddine
Publicado: (2026)
Policy-Invisible Violations in LLM-Based Agents
por: Wu, Jie, et al.
Publicado: (2026)
por: Wu, Jie, et al.
Publicado: (2026)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
por: Landesberg, Eddie
Publicado: (2026)
por: Landesberg, Eddie
Publicado: (2026)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
por: Zhang, Wenbo, et al.
Publicado: (2026)
por: Zhang, Wenbo, et al.
Publicado: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
por: Whitehouse, Chenxi, et al.
Publicado: (2025)
por: Whitehouse, Chenxi, et al.
Publicado: (2025)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
por: Xu, Ran, et al.
Publicado: (2025)
por: Xu, Ran, et al.
Publicado: (2025)
LLM-as-a-Judge for Time Series Explanations
por: Sivalingam, Preetham, et al.
Publicado: (2026)
por: Sivalingam, Preetham, et al.
Publicado: (2026)
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
por: Zhao, Zirui, et al.
Publicado: (2024)
por: Zhao, Zirui, et al.
Publicado: (2024)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
por: Muhamed, Aashiq
Publicado: (2025)
por: Muhamed, Aashiq
Publicado: (2025)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
por: Guo, Kevin H., et al.
Publicado: (2026)
por: Guo, Kevin H., et al.
Publicado: (2026)
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
por: Jain, Ojas, et al.
Publicado: (2026)
por: Jain, Ojas, et al.
Publicado: (2026)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
por: Duo, Jiangshan, et al.
Publicado: (2026)
por: Duo, Jiangshan, et al.
Publicado: (2026)
RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
por: Sharma, Raghav, et al.
Publicado: (2025)
por: Sharma, Raghav, et al.
Publicado: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
por: Ramesh, Shyam Sundhar, et al.
Publicado: (2026)
por: Ramesh, Shyam Sundhar, et al.
Publicado: (2026)
Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
por: Zheng, Hang, et al.
Publicado: (2025)
por: Zheng, Hang, et al.
Publicado: (2025)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
por: Swain, Sankalp Tattwadarshi, et al.
Publicado: (2025)
por: Swain, Sankalp Tattwadarshi, et al.
Publicado: (2025)
Ejemplares similares
-
Context Over Content: Exposing Evaluation Faking in Automated Judges
por: Gupta, Manan, et al.
Publicado: (2026) -
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
por: Gupta, Manan, et al.
Publicado: (2026) -
Investigating Non-Transitivity in LLM-as-a-Judge
por: Xu, Yi, et al.
Publicado: (2025) -
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
por: Xu, Austin, et al.
Publicado: (2025) -
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
por: Hong, Yihan, et al.
Publicado: (2026)