Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport
Fuente:
arXiv
Salvato in:
| Autore principale: | Morandi, Andrea |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Scoring Bias in LLM-as-a-Judge
di: Li, Qingquan, et al.
Pubblicazione: (2025)
di: Li, Qingquan, et al.
Pubblicazione: (2025)
Correcting Selection Bias in Sparse User Feedback for Large Language Model Quality Estimation: A Multi-Agent Hierarchical Bayesian Approach
di: Morandi, Andrea, et al.
Pubblicazione: (2026)
di: Morandi, Andrea, et al.
Pubblicazione: (2026)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
di: Fujinuma, Yoshinari
Pubblicazione: (2025)
di: Fujinuma, Yoshinari
Pubblicazione: (2025)
RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning
di: Morandi, Andrea
Pubblicazione: (2026)
di: Morandi, Andrea
Pubblicazione: (2026)
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
di: Zhou, Karen, et al.
Pubblicazione: (2026)
di: Zhou, Karen, et al.
Pubblicazione: (2026)
Same Input, Different Scores: A Multi Model Study on the Inconsistency of LLM Judge
di: Lau, Fiona
Pubblicazione: (2026)
di: Lau, Fiona
Pubblicazione: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024)
di: Shi, Lin, et al.
Pubblicazione: (2024)
Self-Preference Bias in LLM-as-a-Judge
di: Wataoka, Koki, et al.
Pubblicazione: (2024)
di: Wataoka, Koki, et al.
Pubblicazione: (2024)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
di: Barnes, Jeremy, et al.
Pubblicazione: (2025)
di: Barnes, Jeremy, et al.
Pubblicazione: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
di: Landesberg, Eddie
Pubblicazione: (2026)
di: Landesberg, Eddie
Pubblicazione: (2026)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
di: Kucia, Filip J., et al.
Pubblicazione: (2026)
di: Kucia, Filip J., et al.
Pubblicazione: (2026)
QA-Calibration of Language Model Confidence Scores
di: Manggala, Putra, et al.
Pubblicazione: (2024)
di: Manggala, Putra, et al.
Pubblicazione: (2024)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
di: He, Zhonghao, et al.
Pubblicazione: (2025)
di: He, Zhonghao, et al.
Pubblicazione: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
di: Feng, Yuanning, et al.
Pubblicazione: (2025)
di: Feng, Yuanning, et al.
Pubblicazione: (2025)
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
di: Wang, Yinjie, et al.
Pubblicazione: (2025)
di: Wang, Yinjie, et al.
Pubblicazione: (2025)
CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
di: Li, Haitao, et al.
Pubblicazione: (2024)
di: Li, Haitao, et al.
Pubblicazione: (2024)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
di: Balkır, Esma, et al.
Pubblicazione: (2026)
di: Balkır, Esma, et al.
Pubblicazione: (2026)
A Finite-Calibration Regime Map for LLM Judge Panels
di: Zhu, Bin, et al.
Pubblicazione: (2026)
di: Zhu, Bin, et al.
Pubblicazione: (2026)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
di: Aksoy, Sinan G., et al.
Pubblicazione: (2026)
di: Aksoy, Sinan G., et al.
Pubblicazione: (2026)
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
di: Liu, Zhuo, et al.
Pubblicazione: (2025)
di: Liu, Zhuo, et al.
Pubblicazione: (2025)
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges
di: Tapwal, Riya, et al.
Pubblicazione: (2026)
di: Tapwal, Riya, et al.
Pubblicazione: (2026)
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
di: Guo, Shuchen, et al.
Pubblicazione: (2025)
di: Guo, Shuchen, et al.
Pubblicazione: (2025)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
di: Karim, Ahmed, et al.
Pubblicazione: (2025)
di: Karim, Ahmed, et al.
Pubblicazione: (2025)
From Isolated Scoring to Collaborative Ranking: A Comparison-Native Framework for LLM-Based Paper Evaluation
di: Zheng, Pujun, et al.
Pubblicazione: (2026)
di: Zheng, Pujun, et al.
Pubblicazione: (2026)
Judging It, Washing It: Scoring and Greenwashing Corporate Climate Disclosures using Large Language Models
di: Chuang, Marianne, et al.
Pubblicazione: (2025)
di: Chuang, Marianne, et al.
Pubblicazione: (2025)
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
di: Blackwell, Robert E., et al.
Pubblicazione: (2024)
di: Blackwell, Robert E., et al.
Pubblicazione: (2024)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
di: Lai, Peng, et al.
Pubblicazione: (2026)
di: Lai, Peng, et al.
Pubblicazione: (2026)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
di: Hallaç, İbrahim Rıza, et al.
Pubblicazione: (2026)
di: Hallaç, İbrahim Rıza, et al.
Pubblicazione: (2026)
Phrase-Level Adversarial Training for Mitigating Bias in Neural Network-based Automatic Essay Scoring
di: Philip, Haddad, et al.
Pubblicazione: (2024)
di: Philip, Haddad, et al.
Pubblicazione: (2024)
Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring
di: Li, Jiazheng, et al.
Pubblicazione: (2024)
di: Li, Jiazheng, et al.
Pubblicazione: (2024)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
di: Zhou, Xiaolin, et al.
Pubblicazione: (2026)
di: Zhou, Xiaolin, et al.
Pubblicazione: (2026)
Beyond Bias Scores: Unmasking Vacuous Neutrality in Small Language Models
di: Manduru, Sumanth, et al.
Pubblicazione: (2025)
di: Manduru, Sumanth, et al.
Pubblicazione: (2025)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
di: Chu, SeongYeub, et al.
Pubblicazione: (2024)
di: Chu, SeongYeub, et al.
Pubblicazione: (2024)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
di: Zheng, Danna, et al.
Pubblicazione: (2024)
di: Zheng, Danna, et al.
Pubblicazione: (2024)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
di: Yang, Jinming, et al.
Pubblicazione: (2026)
di: Yang, Jinming, et al.
Pubblicazione: (2026)
Beyond Raw Detection Scores: Markov-Informed Calibration for Boosting Machine-Generated Text Detection
di: Wu, Chenwang, et al.
Pubblicazione: (2026)
di: Wu, Chenwang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Evaluating Scoring Bias in LLM-as-a-Judge
di: Li, Qingquan, et al.
Pubblicazione: (2025) -
Correcting Selection Bias in Sparse User Feedback for Large Language Model Quality Estimation: A Multi-Agent Hierarchical Bayesian Approach
di: Morandi, Andrea, et al.
Pubblicazione: (2026) -
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
di: Fujinuma, Yoshinari
Pubblicazione: (2025) -
RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning
di: Morandi, Andrea
Pubblicazione: (2026) -
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
di: Zhou, Karen, et al.
Pubblicazione: (2026)