Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Dadkhahi, Hamid, Trabelsi, Firas, Riley, Parker, Juraska, Juraj, Mirzazadeh, Mehdi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
by: Zhang, Lily H., et al.
Published: (2024)
by: Zhang, Lily H., et al.
Published: (2024)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
by: Halder, Indranil, et al.
Published: (2025)
by: Halder, Indranil, et al.
Published: (2025)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
by: Whitehouse, Chenxi, et al.
Published: (2025)
by: Whitehouse, Chenxi, et al.
Published: (2025)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
by: Wang, Yutong, et al.
Published: (2025)
by: Wang, Yutong, et al.
Published: (2025)
Auto-Prompt Ensemble for LLM Judge
by: Li, Jiajie, et al.
Published: (2025)
by: Li, Jiajie, et al.
Published: (2025)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
by: Wagner, Nico, et al.
Published: (2024)
by: Wagner, Nico, et al.
Published: (2024)
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
by: Zhang, Yasi, et al.
Published: (2026)
by: Zhang, Yasi, et al.
Published: (2026)
Towards Trustworthy Machine Learning in Production: An Overview of the Robustness in MLOps Approach
by: Bayram, Firas, et al.
Published: (2024)
by: Bayram, Firas, et al.
Published: (2024)
Calibrated Test-Time Guidance for Bayesian Inference
by: Geyfman, Daniel, et al.
Published: (2026)
by: Geyfman, Daniel, et al.
Published: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
by: Adamska, Marta, et al.
Published: (2025)
by: Adamska, Marta, et al.
Published: (2025)
Investigating Non-Transitivity in LLM-as-a-Judge
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Automatic Calibration for Membership Inference Attack on Large Language Models
by: Zade, Saleh Zare, et al.
Published: (2025)
by: Zade, Saleh Zare, et al.
Published: (2025)
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
by: Feuer, Benjamin, et al.
Published: (2024)
by: Feuer, Benjamin, et al.
Published: (2024)
Toward Conditional Distribution Calibration in Survival Prediction
by: Qi, Shi-ang, et al.
Published: (2024)
by: Qi, Shi-ang, et al.
Published: (2024)
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
by: Dey, Soumik, et al.
Published: (2025)
by: Dey, Soumik, et al.
Published: (2025)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
by: Li, Zhuochun, et al.
Published: (2026)
by: Li, Zhuochun, et al.
Published: (2026)
Think like a Scientist: Physics-guided LLM Agent for Equation Discovery
by: Yang, Jianke, et al.
Published: (2026)
by: Yang, Jianke, et al.
Published: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
by: Agnihotri, Rudransh, et al.
Published: (2025)
by: Agnihotri, Rudransh, et al.
Published: (2025)
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
by: Shen, William F., et al.
Published: (2026)
by: Shen, William F., et al.
Published: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
by: Fabbri, Francesco, et al.
Published: (2025)
by: Fabbri, Francesco, et al.
Published: (2025)
Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
by: Wu, Jiayun, et al.
Published: (2025)
by: Wu, Jiayun, et al.
Published: (2025)
Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
by: Luo, Beier, et al.
Published: (2025)
by: Luo, Beier, et al.
Published: (2025)
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering
by: Wu, Fangzhou, et al.
Published: (2026)
by: Wu, Fangzhou, et al.
Published: (2026)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
by: Zhan, Weixiao, et al.
Published: (2026)
by: Zhan, Weixiao, et al.
Published: (2026)
Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits
by: Pershin, Maksim, et al.
Published: (2026)
by: Pershin, Maksim, et al.
Published: (2026)
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
by: Sharma, Aman, et al.
Published: (2025)
by: Sharma, Aman, et al.
Published: (2025)
Robust Calibration For Improved Weather Prediction Under Distributional Shift
by: Gilda, Sankalp, et al.
Published: (2024)
by: Gilda, Sankalp, et al.
Published: (2024)
Semantic-based Distributed Learning for Diverse and Discriminative Representations
by: Tian, Zhuojun, et al.
Published: (2026)
by: Tian, Zhuojun, et al.
Published: (2026)
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
by: Noh, Kangjun, et al.
Published: (2026)
by: Noh, Kangjun, et al.
Published: (2026)
Distributional Multi-objective Black-box Optimization for Diffusion-model Inference-time Multi-Target Generation
by: Tan, Kim Yong, et al.
Published: (2025)
by: Tan, Kim Yong, et al.
Published: (2025)
Distilling Calibration via Conformalized Credal Inference
by: Huang, Jiayi, et al.
Published: (2025)
by: Huang, Jiayi, et al.
Published: (2025)
Similar Items
-
Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
by: Zhang, Lily H., et al.
Published: (2024) -
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
by: Radharapu, Bhaktipriya, et al.
Published: (2025) -
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
by: Halder, Indranil, et al.
Published: (2025) -
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025) -
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
by: Whitehouse, Chenxi, et al.
Published: (2025)