Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
Fuente:
arXiv
Salvato in:
| Autori principali: | Fathullah, Yassir, Gales, Mark J. F. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Sample-Specific Encoder Perturbations
di: Fathullah, Yassir, et al.
Pubblicazione: (2024)
di: Fathullah, Yassir, et al.
Pubblicazione: (2024)
Who can we trust? LLM-as-a-jury for Comparative Assessment
di: Qian, Mengjie, et al.
Pubblicazione: (2026)
di: Qian, Mengjie, et al.
Pubblicazione: (2026)
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
di: Liusie, Adian, et al.
Pubblicazione: (2024)
di: Liusie, Adian, et al.
Pubblicazione: (2024)
Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
di: Raju, Ravi, et al.
Pubblicazione: (2024)
di: Raju, Ravi, et al.
Pubblicazione: (2024)
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
di: Liusie, Adian, et al.
Pubblicazione: (2024)
di: Liusie, Adian, et al.
Pubblicazione: (2024)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
di: Li, Dawei, et al.
Pubblicazione: (2025)
di: Li, Dawei, et al.
Pubblicazione: (2025)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
di: Sun, Guangzhi, et al.
Pubblicazione: (2025)
di: Sun, Guangzhi, et al.
Pubblicazione: (2025)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
Probabilistic Neural Networks (PNNs) for Modeling Aleatoric Uncertainty in Scientific Machine Learning
di: Pourkamali-Anaraki, Farhad, et al.
Pubblicazione: (2024)
di: Pourkamali-Anaraki, Farhad, et al.
Pubblicazione: (2024)
Reconsidering LLM Uncertainty Estimation Methods in the Wild
di: Bakman, Yavuz, et al.
Pubblicazione: (2025)
di: Bakman, Yavuz, et al.
Pubblicazione: (2025)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
di: Weltevrede, Max, et al.
Pubblicazione: (2025)
di: Weltevrede, Max, et al.
Pubblicazione: (2025)
Rational Tuning of LLM Cascades via Probabilistic Modeling
di: Zellinger, Michael J., et al.
Pubblicazione: (2025)
di: Zellinger, Michael J., et al.
Pubblicazione: (2025)
Generalisation of Total Uncertainty in AI: A Theoretical Study
di: Shariatmadar, Keivan
Pubblicazione: (2024)
di: Shariatmadar, Keivan
Pubblicazione: (2024)
Beyond Static Uncertainty: Modeling Temporal Uncertainty Dynamics for Probabilistic Time Series Forecasting
di: Wang, Yijun, et al.
Pubblicazione: (2026)
di: Wang, Yijun, et al.
Pubblicazione: (2026)
REPEAT: Improving Uncertainty Estimation in Representation Learning Explainability
di: Wickstrøm, Kristoffer K., et al.
Pubblicazione: (2024)
di: Wickstrøm, Kristoffer K., et al.
Pubblicazione: (2024)
A Research Agenda for Usability and Generalisation in Reinforcement Learning
di: Soemers, Dennis J. N. J., et al.
Pubblicazione: (2024)
di: Soemers, Dennis J. N. J., et al.
Pubblicazione: (2024)
Improving Uncertainty Estimation through Semantically Diverse Language Generation
di: Aichberger, Lukas, et al.
Pubblicazione: (2024)
di: Aichberger, Lukas, et al.
Pubblicazione: (2024)
Uncertainty Quantification in Probabilistic Machine Learning Models: Theory, Methods, and Insights
di: Ajirak, Marzieh, et al.
Pubblicazione: (2025)
di: Ajirak, Marzieh, et al.
Pubblicazione: (2025)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
When Astronomy Meets AI: Manazel For Crescent Visibility Prediction in Morocco
di: Lairgi, Yassir
Pubblicazione: (2025)
di: Lairgi, Yassir
Pubblicazione: (2025)
An Uncertainty-Aware ED-LSTM for Probabilistic Suffix Prediction
di: Mustroph, Henryk, et al.
Pubblicazione: (2025)
di: Mustroph, Henryk, et al.
Pubblicazione: (2025)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
di: Kirk, Robert, et al.
Pubblicazione: (2023)
di: Kirk, Robert, et al.
Pubblicazione: (2023)
Wasserstein Gradient Flows for Scalable and Regularized Barycenter Computation
di: Montesuma, Eduardo Fernandes, et al.
Pubblicazione: (2025)
di: Montesuma, Eduardo Fernandes, et al.
Pubblicazione: (2025)
How Uncertainty Estimation Scales with Sampling in Reasoning Models
di: Del, Maksym, et al.
Pubblicazione: (2026)
di: Del, Maksym, et al.
Pubblicazione: (2026)
The Probabilistic Tsetlin Machine: A Novel Approach to Uncertainty Quantification
di: Abeyrathna, K. Darshana, et al.
Pubblicazione: (2024)
di: Abeyrathna, K. Darshana, et al.
Pubblicazione: (2024)
Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation
di: Lee, JoonHo, et al.
Pubblicazione: (2024)
di: Lee, JoonHo, et al.
Pubblicazione: (2024)
Context-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting
di: Yao, Yueyang, et al.
Pubblicazione: (2025)
di: Yao, Yueyang, et al.
Pubblicazione: (2025)
Towards Generalisable Imitation Learning Through Conditioned Transition Estimation and Online Behaviour Alignment
di: Gavenski, Nathan, et al.
Pubblicazione: (2026)
di: Gavenski, Nathan, et al.
Pubblicazione: (2026)
Uncertainty Estimation by Human Perception versus Neural Models
di: Mendes, Pedro, et al.
Pubblicazione: (2025)
di: Mendes, Pedro, et al.
Pubblicazione: (2025)
Credal Wrapper of Model Averaging for Uncertainty Estimation in Classification
di: Wang, Kaizheng, et al.
Pubblicazione: (2024)
di: Wang, Kaizheng, et al.
Pubblicazione: (2024)
On the Relationship between Bayesian Networks and Probabilistic Structural Causal Models
di: Lucas, Peter J. F., et al.
Pubblicazione: (2026)
di: Lucas, Peter J. F., et al.
Pubblicazione: (2026)
Efficient Epistemic Uncertainty Estimation in Regression Ensemble Models Using Pairwise-Distance Estimators
di: Berry, Lucas, et al.
Pubblicazione: (2023)
di: Berry, Lucas, et al.
Pubblicazione: (2023)
Large Language Models in Fire Engineering: An Examination of Technical Questions Against Domain Knowledge
di: Hostetter, Haley, et al.
Pubblicazione: (2024)
di: Hostetter, Haley, et al.
Pubblicazione: (2024)
Quantifying Generalisation in Imitation Learning
di: Gavenski, Nathan, et al.
Pubblicazione: (2025)
di: Gavenski, Nathan, et al.
Pubblicazione: (2025)
Understanding and Improving Adversarial Robustness of Neural Probabilistic Circuits
di: Chen, Weixin, et al.
Pubblicazione: (2025)
di: Chen, Weixin, et al.
Pubblicazione: (2025)
A Cross Modal Knowledge Distillation & Data Augmentation Recipe for Improving Transcriptomics Representations through Morphological Features
di: Bendidi, Ihab, et al.
Pubblicazione: (2025)
di: Bendidi, Ihab, et al.
Pubblicazione: (2025)
Bayesian Modelling in Practice: Using Uncertainty to Improve Trustworthiness in Medical Applications
di: Ruhe, David, et al.
Pubblicazione: (2019)
di: Ruhe, David, et al.
Pubblicazione: (2019)
Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning
di: Manchingal, Shireen Kudukkil, et al.
Pubblicazione: (2026)
di: Manchingal, Shireen Kudukkil, et al.
Pubblicazione: (2026)
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2025)
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Sample-Specific Encoder Perturbations
di: Fathullah, Yassir, et al.
Pubblicazione: (2024) -
Who can we trust? LLM-as-a-jury for Comparative Assessment
di: Qian, Mengjie, et al.
Pubblicazione: (2026) -
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
di: Liusie, Adian, et al.
Pubblicazione: (2024) -
Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
di: Raju, Ravi, et al.
Pubblicazione: (2024) -
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
di: Liusie, Adian, et al.
Pubblicazione: (2024)