Same Input, Different Scores: A Multi Model Study on the Inconsistency of LLM Judge
Fuente:
arXiv
Guardado en:
| Autor principal: | Lau, Fiona |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
por: Kostić, Bogdan, et al.
Publicado: (2026)
por: Kostić, Bogdan, et al.
Publicado: (2026)
Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
por: Haldar, Rajarshi, et al.
Publicado: (2025)
por: Haldar, Rajarshi, et al.
Publicado: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
por: Wang, Yidong, et al.
Publicado: (2025)
por: Wang, Yidong, et al.
Publicado: (2025)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2026)
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2026)
Evaluating Scoring Bias in LLM-as-a-Judge
por: Li, Qingquan, et al.
Publicado: (2025)
por: Li, Qingquan, et al.
Publicado: (2025)
Same Content, Different Representations: A Controlled Study for Table QA
por: Zhang, Yue, et al.
Publicado: (2025)
por: Zhang, Yue, et al.
Publicado: (2025)
Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias
por: Tonneau, Manuel, et al.
Publicado: (2026)
por: Tonneau, Manuel, et al.
Publicado: (2026)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
Personalized LLM for Generating Customized Responses to the Same Query from Different Users
por: Zeng, Hang, et al.
Publicado: (2024)
por: Zeng, Hang, et al.
Publicado: (2024)
The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling
por: Zhang, Ruochen, et al.
Publicado: (2024)
por: Zhang, Ruochen, et al.
Publicado: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
por: Shi, Lin, et al.
Publicado: (2024)
por: Shi, Lin, et al.
Publicado: (2024)
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
por: Zhou, Karen, et al.
Publicado: (2026)
por: Zhou, Karen, et al.
Publicado: (2026)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
por: Tang, Zhenwei, et al.
Publicado: (2026)
por: Tang, Zhenwei, et al.
Publicado: (2026)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
por: Huang, Hui, et al.
Publicado: (2024)
por: Huang, Hui, et al.
Publicado: (2024)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
por: Fujinuma, Yoshinari
Publicado: (2025)
por: Fujinuma, Yoshinari
Publicado: (2025)
Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing
por: Ahn, Jihyun Janice, et al.
Publicado: (2025)
por: Ahn, Jihyun Janice, et al.
Publicado: (2025)
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
por: Lin, Zizhuo, et al.
Publicado: (2026)
por: Lin, Zizhuo, et al.
Publicado: (2026)
Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport
por: Morandi, Andrea
Publicado: (2026)
por: Morandi, Andrea
Publicado: (2026)
Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
por: Levy, Mosh, et al.
Publicado: (2024)
por: Levy, Mosh, et al.
Publicado: (2024)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
por: Barnes, Jeremy, et al.
Publicado: (2025)
por: Barnes, Jeremy, et al.
Publicado: (2025)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
por: Kamoi, Ryo, et al.
Publicado: (2026)
por: Kamoi, Ryo, et al.
Publicado: (2026)
Misleading through Inconsistency: A Benchmark for Political Inconsistencies Detection
por: Sagimbayeva, Nursulu, et al.
Publicado: (2025)
por: Sagimbayeva, Nursulu, et al.
Publicado: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
por: Hong, Yihan, et al.
Publicado: (2026)
por: Hong, Yihan, et al.
Publicado: (2026)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
por: Landesberg, Eddie
Publicado: (2026)
por: Landesberg, Eddie
Publicado: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
por: Yang, Bo, et al.
Publicado: (2026)
por: Yang, Bo, et al.
Publicado: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
por: Marioriyad, Arash, et al.
Publicado: (2025)
por: Marioriyad, Arash, et al.
Publicado: (2025)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
por: Bellibatlu, Rohith Reddy, et al.
Publicado: (2026)
por: Bellibatlu, Rohith Reddy, et al.
Publicado: (2026)
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
por: Wu, Junjie, et al.
Publicado: (2026)
por: Wu, Junjie, et al.
Publicado: (2026)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
por: Su, Jiamin, et al.
Publicado: (2025)
por: Su, Jiamin, et al.
Publicado: (2025)
Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG
por: Li, Yubo, et al.
Publicado: (2026)
por: Li, Yubo, et al.
Publicado: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
por: Belmadani, Ikram, et al.
Publicado: (2026)
por: Belmadani, Ikram, et al.
Publicado: (2026)
Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs
por: Ford, Casey, et al.
Publicado: (2026)
por: Ford, Casey, et al.
Publicado: (2026)
AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
por: Miao, Yongliang, et al.
Publicado: (2026)
por: Miao, Yongliang, et al.
Publicado: (2026)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
por: Zhu, Ziyi, et al.
Publicado: (2026)
por: Zhu, Ziyi, et al.
Publicado: (2026)
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
por: Lee, Yoonsang, et al.
Publicado: (2024)
por: Lee, Yoonsang, et al.
Publicado: (2024)
Input Order Shapes LLM Semantic Alignment in Multi-Document Summarization
por: Ma, Jing
Publicado: (2025)
por: Ma, Jing
Publicado: (2025)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
por: Le, Hieu Xuan, et al.
Publicado: (2026)
por: Le, Hieu Xuan, et al.
Publicado: (2026)
Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs
por: Alkaeed, Mahdi, et al.
Publicado: (2026)
por: Alkaeed, Mahdi, et al.
Publicado: (2026)
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
por: Tang, Yuqi, et al.
Publicado: (2025)
por: Tang, Yuqi, et al.
Publicado: (2025)
Debatrix: Multi-dimensional Debate Judge with Iterative Chronological Analysis Based on LLM
por: Liang, Jingcong, et al.
Publicado: (2024)
por: Liang, Jingcong, et al.
Publicado: (2024)
Ejemplares similares
-
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
por: Kostić, Bogdan, et al.
Publicado: (2026) -
Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
por: Haldar, Rajarshi, et al.
Publicado: (2025) -
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
por: Wang, Yidong, et al.
Publicado: (2025) -
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2026) -
Evaluating Scoring Bias in LLM-as-a-Judge
por: Li, Qingquan, et al.
Publicado: (2025)