Black-box Uncertainty Quantification Method for LLM-as-a-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Wagner, Nico, Desmond, Michael, Nair, Rahul, Ashktorab, Zahra, Daly, Elizabeth M., Pan, Qian, Cooper, Martín Santillán, Johnson, James M., Geyer, Werner |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human-Centered Design Recommendations for LLM-as-a-Judge
by: Pan, Qian, et al.
Published: (2024)
by: Pan, Qian, et al.
Published: (2024)
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
by: Ashktorab, Zahra, et al.
Published: (2025)
by: Ashktorab, Zahra, et al.
Published: (2025)
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
by: Ashktorab, Zahra, et al.
Published: (2024)
by: Ashktorab, Zahra, et al.
Published: (2024)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
by: Do, Hyo Jin, et al.
Published: (2025)
by: Do, Hyo Jin, et al.
Published: (2025)
Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations
by: Ashktorab, Zahra, et al.
Published: (2024)
by: Ashktorab, Zahra, et al.
Published: (2024)
Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
by: Song, Jaeyoon, et al.
Published: (2025)
by: Song, Jaeyoon, et al.
Published: (2025)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
by: Chiang, Charles, et al.
Published: (2026)
by: Chiang, Charles, et al.
Published: (2026)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
by: Gajcin, Jasmina, et al.
Published: (2025)
by: Gajcin, Jasmina, et al.
Published: (2025)
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
by: Lin, Zhen, et al.
Published: (2023)
by: Lin, Zhen, et al.
Published: (2023)
Granite Guardian
by: Padhi, Inkit, et al.
Published: (2024)
by: Padhi, Inkit, et al.
Published: (2024)
Helping the Helper: Supporting Peer Counselors via AI-Empowered Practice and Feedback
by: Hsu, Shang-Ling, et al.
Published: (2023)
by: Hsu, Shang-Ling, et al.
Published: (2023)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
A Representation-Level Assessment of Bias Mitigation in Foundation Models
by: Nizhnichenkov, Svetoslav, et al.
Published: (2026)
by: Nizhnichenkov, Svetoslav, et al.
Published: (2026)
Bias and Uncertainty in LLM-as-a-Judge Estimation
by: Fiedler, James
Published: (2026)
by: Fiedler, James
Published: (2026)
A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development
by: Geyer, Werner, et al.
Published: (2025)
by: Geyer, Werner, et al.
Published: (2025)
Humble AI in the real-world: the case of algorithmic hiring
by: Nair, Rahul, et al.
Published: (2025)
by: Nair, Rahul, et al.
Published: (2025)
Paying Alignment Tax with Contrastive Learning
by: Korkmaz, Buse Sibel, et al.
Published: (2025)
by: Korkmaz, Buse Sibel, et al.
Published: (2025)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
by: Gebreegziabher, Simret Araya, et al.
Published: (2026)
by: Gebreegziabher, Simret Araya, et al.
Published: (2026)
Estimating the Black-box LLM Uncertainty with Distribution-Aligned Adversarial Distillation
by: Cui, Huizi, et al.
Published: (2026)
by: Cui, Huizi, et al.
Published: (2026)
Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability
by: Song, Jaeyoon, et al.
Published: (2025)
by: Song, Jaeyoon, et al.
Published: (2025)
Ranking Large Language Models without Ground Truth
by: Dhurandhar, Amit, et al.
Published: (2024)
by: Dhurandhar, Amit, et al.
Published: (2024)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
by: Ye, Jiayi, et al.
Published: (2024)
by: Ye, Jiayi, et al.
Published: (2024)
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
Fast HARDI Uncertainty Quantification and Visualization with Spherical Sampling
by: Tark Patel, et al.
Published: (2025)
by: Tark Patel, et al.
Published: (2025)
From Black-box to Causal-box: Towards Building More Interpretable Models
by: Hwang, Inwoo, et al.
Published: (2025)
by: Hwang, Inwoo, et al.
Published: (2025)
Am I Confused or Is This Confusing?: Deep Ensembles for ENSO Uncertainty Quantification
by: McAfee, Devin M., et al.
Published: (2025)
by: McAfee, Devin M., et al.
Published: (2025)
Foundation Models at Work: Fine-Tuning for Fairness in Algorithmic Hiring
by: Korkmaz, Buse Sibel, et al.
Published: (2025)
by: Korkmaz, Buse Sibel, et al.
Published: (2025)
Uncertainty Quantification in Probabilistic Machine Learning Models: Theory, Methods, and Insights
by: Ajirak, Marzieh, et al.
Published: (2025)
by: Ajirak, Marzieh, et al.
Published: (2025)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
Uncertainty Quantification and Decomposition for LLM-based Recommendation
by: Kweon, Wonbin, et al.
Published: (2025)
by: Kweon, Wonbin, et al.
Published: (2025)
Holographic Equidistribution
by: Cooper, Nico
Published: (2026)
by: Cooper, Nico
Published: (2026)
Convolutional Neural Networks For Turbulent Model Uncertainty Quantification
by: Chu, Minghan, et al.
Published: (2024)
by: Chu, Minghan, et al.
Published: (2024)
Black-box Optimization of LLM Outputs by Asking for Directions
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Statistically Optimal Uncertainty Quantification for Expensive Black-Box Models
by: He, Shengyi, et al.
Published: (2024)
by: He, Shengyi, et al.
Published: (2024)
Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
by: McCabe, Lucas H., et al.
Published: (2025)
by: McCabe, Lucas H., et al.
Published: (2025)
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
A Survey on Uncertainty Quantification Methods for Deep Learning
by: He, Wenchong, et al.
Published: (2023)
by: He, Wenchong, et al.
Published: (2023)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
by: Seleznyov, Mikhail, et al.
Published: (2026)
by: Seleznyov, Mikhail, et al.
Published: (2026)
Schrodinger Neural Network and Uncertainty Quantification: Quantum Machine
by: Hammad, M. M.
Published: (2025)
by: Hammad, M. M.
Published: (2025)
Similar Items
-
Human-Centered Design Recommendations for LLM-as-a-Judge
by: Pan, Qian, et al.
Published: (2024) -
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
by: Ashktorab, Zahra, et al.
Published: (2025) -
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
by: Ashktorab, Zahra, et al.
Published: (2024) -
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
by: Do, Hyo Jin, et al.
Published: (2025) -
Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations
by: Ashktorab, Zahra, et al.
Published: (2024)