Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Radharapu, Bhaktipriya, Saxena, Eshika, Li, Kenneth, Whitehouse, Chenxi, Williams, Adina, Cancedda, Nicola |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
by: Sturman, Olivia, et al.
Published: (2024)
by: Sturman, Olivia, et al.
Published: (2024)
RealSeal: Revolutionizing Media Authentication with Real-Time Realism Scoring
by: Radharapu, Bhaktipriya, et al.
Published: (2024)
by: Radharapu, Bhaktipriya, et al.
Published: (2024)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
by: Whitehouse, Chenxi, et al.
Published: (2025)
by: Whitehouse, Chenxi, et al.
Published: (2025)
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
by: Shen, William F., et al.
Published: (2026)
by: Shen, William F., et al.
Published: (2026)
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
by: Koishekenov, Yeskendir, et al.
Published: (2025)
by: Koishekenov, Yeskendir, et al.
Published: (2025)
Why Uncertainty Calibration Matters for Reliable Perturbation-based Explanations
by: Decker, Thomas, et al.
Published: (2025)
by: Decker, Thomas, et al.
Published: (2025)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
by: Dadkhahi, Hamid, et al.
Published: (2025)
by: Dadkhahi, Hamid, et al.
Published: (2025)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
by: Wagner, Nico, et al.
Published: (2024)
by: Wagner, Nico, et al.
Published: (2024)
LLM Unlearning via Neural Activation Redirection
by: Shen, William F., et al.
Published: (2025)
by: Shen, William F., et al.
Published: (2025)
Enhancing In-context Learning via Linear Probe Calibration
by: Abbas, Momin, et al.
Published: (2024)
by: Abbas, Momin, et al.
Published: (2024)
Verifying Chain-of-Thought Reasoning via Its Computational Graph
by: Zhao, Zheng, et al.
Published: (2025)
by: Zhao, Zheng, et al.
Published: (2025)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
A Monte Carlo Framework for Calibrated Uncertainty Estimation in Sequence Prediction
by: Yang, Qidong, et al.
Published: (2024)
by: Yang, Qidong, et al.
Published: (2024)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Diffusion Tensor Estimation with Uncertainty Calibration
by: Karimi, Davood, et al.
Published: (2021)
by: Karimi, Davood, et al.
Published: (2021)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
by: Agarwal, Chirag, et al.
Published: (2022)
by: Agarwal, Chirag, et al.
Published: (2022)
Reconsidering LLM Uncertainty Estimation Methods in the Wild
by: Bakman, Yavuz, et al.
Published: (2025)
by: Bakman, Yavuz, et al.
Published: (2025)
Auto-Prompt Ensemble for LLM Judge
by: Li, Jiajie, et al.
Published: (2025)
by: Li, Jiajie, et al.
Published: (2025)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
by: Wang, Yutong, et al.
Published: (2025)
by: Wang, Yutong, et al.
Published: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
by: Hong, Yihan, et al.
Published: (2026)
by: Hong, Yihan, et al.
Published: (2026)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
by: Tao, Linwei, et al.
Published: (2025)
by: Tao, Linwei, et al.
Published: (2025)
Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity
by: Li, Bojie
Published: (2026)
by: Li, Bojie
Published: (2026)
Calibrated Explanations: with Uncertainty Information and Counterfactuals
by: Lofstrom, Helena, et al.
Published: (2023)
by: Lofstrom, Helena, et al.
Published: (2023)
Fast Calibrated Explanations: Efficient and Uncertainty-Aware Explanations for Machine Learning Models
by: Löfström, Tuwe, et al.
Published: (2024)
by: Löfström, Tuwe, et al.
Published: (2024)
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Spectral Filters, Dark Signals, and Attention Sinks
by: Cancedda, Nicola
Published: (2024)
by: Cancedda, Nicola
Published: (2024)
Enhancing Deep Neural Network Reliability with Refinement and Calibration
by: Hebbalaguppe, Ramya, et al.
Published: (2026)
by: Hebbalaguppe, Ramya, et al.
Published: (2026)
Adjusting Regression Models for Conditional Uncertainty Calibration
by: Gao, Ruijiang, et al.
Published: (2024)
by: Gao, Ruijiang, et al.
Published: (2024)
From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty
by: Jenane, Azza, et al.
Published: (2026)
by: Jenane, Azza, et al.
Published: (2026)
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
by: Dey, Soumik, et al.
Published: (2025)
by: Dey, Soumik, et al.
Published: (2025)
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
by: Fathullah, Yassir, et al.
Published: (2025)
by: Fathullah, Yassir, et al.
Published: (2025)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
by: Agnimo, Yedidia, et al.
Published: (2026)
by: Agnimo, Yedidia, et al.
Published: (2026)
Rhetorical Questions in LLM Representations: A Linear Probing Study
by: Yao, Louie Hong, et al.
Published: (2026)
by: Yao, Louie Hong, et al.
Published: (2026)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
by: Dai, Hui, et al.
Published: (2026)
by: Dai, Hui, et al.
Published: (2026)
ConformaDecompose: Explaining Uncertainty via Calibration Localization
by: Yapicioglu, Fatima Rabia, et al.
Published: (2026)
by: Yapicioglu, Fatima Rabia, et al.
Published: (2026)
Uncertainty-Calibrated Spatiotemporal Field Diffusion with Sparse Supervision
by: Valencia, Kevin, et al.
Published: (2026)
by: Valencia, Kevin, et al.
Published: (2026)
Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
by: Liu, Zijun, et al.
Published: (2024)
by: Liu, Zijun, et al.
Published: (2024)
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
by: Whitehouse, Chenxi, et al.
Published: (2023)
by: Whitehouse, Chenxi, et al.
Published: (2023)
Similar Items
-
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
by: Sturman, Olivia, et al.
Published: (2024) -
RealSeal: Revolutionizing Media Authentication with Real-Time Realism Scoring
by: Radharapu, Bhaktipriya, et al.
Published: (2024) -
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
by: Whitehouse, Chenxi, et al.
Published: (2025) -
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
by: Shen, William F., et al.
Published: (2026) -
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
by: Koishekenov, Yeskendir, et al.
Published: (2025)