The Necessity of Setting Temperature in LLM-as-a-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Lujun, Sleem, Lama, Xu, Yangjie, Song, Yewei, Jia, Aolin, Francois, Jerome, State, Radu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Impact of Temperature on Large Language Models:Hot or Cold?
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
Small Language Models in the Real World: Insights from Industrial Text Classification
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
Do Large Language Models Grasp The Grammar? Evidence from Grammar-Book-Guided Probing in Luxembourgish
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
Is Small Language Model the Silver Bullet to Low-Resource Languages Machine Translation?
by: Song, Yewei, et al.
Published: (2025)
by: Song, Yewei, et al.
Published: (2025)
Uncovering Zero-Shot Generalization Gaps in Time-Series Foundation Models Using Real-World Videos
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
Agent Skill Framework: Perspectives on the Potential of Small Language Models in Industrial Environments
by: Xu, Yangjie, et al.
Published: (2026)
by: Xu, Yangjie, et al.
Published: (2026)
NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
by: Sleem, Lama, et al.
Published: (2025)
by: Sleem, Lama, et al.
Published: (2025)
HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
by: Bergeron, Loris, et al.
Published: (2025)
by: Bergeron, Loris, et al.
Published: (2025)
Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
by: Petrova, Tatiana, et al.
Published: (2026)
by: Petrova, Tatiana, et al.
Published: (2026)
Temporal-Spatial Tubelet Embedding for Cloud-Robust MSI Reconstruction using MSI-SAR Fusion: A Multi-Head Self-Attention Video Vision Transformer Approach
by: Wang, Yiqun, et al.
Published: (2025)
by: Wang, Yiqun, et al.
Published: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
by: Gajcin, Jasmina, et al.
Published: (2025)
by: Gajcin, Jasmina, et al.
Published: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?
by: Rajput, Prateek, et al.
Published: (2026)
by: Rajput, Prateek, et al.
Published: (2026)
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
Limitations of Normalization in Attention Mechanism
by: Mudarisov, Timur, et al.
Published: (2025)
by: Mudarisov, Timur, et al.
Published: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
by: Huang, Hui, et al.
Published: (2024)
by: Huang, Hui, et al.
Published: (2024)
Exploring the Necessity of Reasoning in LLM-based Agent Scenarios
by: Zhou, Xueyang, et al.
Published: (2025)
by: Zhou, Xueyang, et al.
Published: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
by: Li, Qingquan, et al.
Published: (2025)
by: Li, Qingquan, et al.
Published: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
by: Belmadani, Ikram, et al.
Published: (2026)
by: Belmadani, Ikram, et al.
Published: (2026)
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
by: Kumar, Shachi H, et al.
Published: (2024)
by: Kumar, Shachi H, et al.
Published: (2024)
Measuring LLM Code Generation Stability via Structural Entropy
by: Song, Yewei, et al.
Published: (2025)
by: Song, Yewei, et al.
Published: (2025)
A Survey on LLM-as-a-Judge
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Improve LLM-as-a-Judge Ability as a General Ability
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
Counterargument for Critical Thinking as Judged by AI and Humans
by: Adewumi, Tosin, et al.
Published: (2026)
by: Adewumi, Tosin, et al.
Published: (2026)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
by: Zhu, Ziyi, et al.
Published: (2026)
by: Zhu, Ziyi, et al.
Published: (2026)
Self-Preference Bias in LLM-as-a-Judge
by: Wataoka, Koki, et al.
Published: (2024)
by: Wataoka, Koki, et al.
Published: (2024)
How Reliable is Multilingual LLM-as-a-Judge?
by: Fu, Xiyan, et al.
Published: (2025)
by: Fu, Xiyan, et al.
Published: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
by: Chen, Junjie, et al.
Published: (2026)
by: Chen, Junjie, et al.
Published: (2026)
Tokenization Falling Short: On Subword Robustness in Large Language Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
Quantitative LLM Judges
by: Sahoo, Aishwarya, et al.
Published: (2025)
by: Sahoo, Aishwarya, et al.
Published: (2025)
Similar Items
-
Exploring the Impact of Temperature on Large Language Models:Hot or Cold?
by: Li, Lujun, et al.
Published: (2025) -
Small Language Models in the Real World: Insights from Industrial Text Classification
by: Li, Lujun, et al.
Published: (2025) -
Do Large Language Models Grasp The Grammar? Evidence from Grammar-Book-Guided Probing in Luxembourgish
by: Li, Lujun, et al.
Published: (2025) -
Is Small Language Model the Silver Bullet to Low-Resource Languages Machine Translation?
by: Song, Yewei, et al.
Published: (2025) -
Uncovering Zero-Shot Generalization Gaps in Time-Series Foundation Models Using Real-World Videos
by: Li, Lujun, et al.
Published: (2025)