Prompting a Weighting Mechanism into LLM-as-a-Judge in Two-Step: A Case Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Wenwen, Gwizdz, Gray, Feng, Dongji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
von: Bellibatlu, Rohith Reddy, et al.
Veröffentlicht: (2026)
von: Bellibatlu, Rohith Reddy, et al.
Veröffentlicht: (2026)
LLM for Complex Reasoning Task: An Exploratory Study in Fermi Problems
von: Liu, Zishuo, et al.
Veröffentlicht: (2025)
von: Liu, Zishuo, et al.
Veröffentlicht: (2025)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
von: Le, Hieu Xuan, et al.
Veröffentlicht: (2026)
von: Le, Hieu Xuan, et al.
Veröffentlicht: (2026)
LLM for Comparative Narrative Analysis
von: Kampen, Leo, et al.
Veröffentlicht: (2025)
von: Kampen, Leo, et al.
Veröffentlicht: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
von: Yang, Bo, et al.
Veröffentlicht: (2026)
von: Yang, Bo, et al.
Veröffentlicht: (2026)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
von: Zhang, Xinran
Veröffentlicht: (2026)
von: Zhang, Xinran
Veröffentlicht: (2026)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
von: Wei, Hui, et al.
Veröffentlicht: (2024)
von: Wei, Hui, et al.
Veröffentlicht: (2024)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
von: Huang, Hui, et al.
Veröffentlicht: (2024)
von: Huang, Hui, et al.
Veröffentlicht: (2024)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
von: Liu, Zhuo, et al.
Veröffentlicht: (2025)
von: Liu, Zhuo, et al.
Veröffentlicht: (2025)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
Can LLM be a Personalized Judge?
von: Dong, Yijiang River, et al.
Veröffentlicht: (2024)
von: Dong, Yijiang River, et al.
Veröffentlicht: (2024)
Are We on the Right Way to Assessing LLM-as-a-Judge?
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
von: Chhabra, Mukul, et al.
Veröffentlicht: (2026)
von: Chhabra, Mukul, et al.
Veröffentlicht: (2026)
Evaluating Scoring Bias in LLM-as-a-Judge
von: Li, Qingquan, et al.
Veröffentlicht: (2025)
von: Li, Qingquan, et al.
Veröffentlicht: (2025)
How Reliable is Multilingual LLM-as-a-Judge?
von: Fu, Xiyan, et al.
Veröffentlicht: (2025)
von: Fu, Xiyan, et al.
Veröffentlicht: (2025)
The Necessity of Setting Temperature in LLM-as-a-Judge
von: Li, Lujun, et al.
Veröffentlicht: (2026)
von: Li, Lujun, et al.
Veröffentlicht: (2026)
Self-Preference Bias in LLM-as-a-Judge
von: Wataoka, Koki, et al.
Veröffentlicht: (2024)
von: Wataoka, Koki, et al.
Veröffentlicht: (2024)
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
von: Akarajaradwong, Pawitsapak, et al.
Veröffentlicht: (2026)
von: Akarajaradwong, Pawitsapak, et al.
Veröffentlicht: (2026)
LLM as a Complementary Optimizer to Gradient Descent: A Case Study in Prompt Tuning
von: Guo, Zixian, et al.
Veröffentlicht: (2024)
von: Guo, Zixian, et al.
Veröffentlicht: (2024)
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
von: Yamauchi, Yusuke, et al.
Veröffentlicht: (2025)
von: Yamauchi, Yusuke, et al.
Veröffentlicht: (2025)
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
von: Schroeder, Kayla, et al.
Veröffentlicht: (2024)
von: Schroeder, Kayla, et al.
Veröffentlicht: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
Humans or LLMs as the Judge? A Study on Judgement Biases
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
von: Wang, Victor, et al.
Veröffentlicht: (2025)
von: Wang, Victor, et al.
Veröffentlicht: (2025)
On Cost-Effective LLM-as-a-Judge Improvement Techniques
von: Lail, Ryan, et al.
Veröffentlicht: (2026)
von: Lail, Ryan, et al.
Veröffentlicht: (2026)
Mediocrity is the key for LLM as a Judge Anchor Selection
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2026)
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2026)
Improve LLM-as-a-Judge Ability as a General Ability
von: Yu, Jiachen, et al.
Veröffentlicht: (2025)
von: Yu, Jiachen, et al.
Veröffentlicht: (2025)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
Same Input, Different Scores: A Multi Model Study on the Inconsistency of LLM Judge
von: Lau, Fiona
Veröffentlicht: (2026)
von: Lau, Fiona
Veröffentlicht: (2026)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
von: Bellibatlu, Rohith Reddy, et al.
Veröffentlicht: (2026) -
LLM for Complex Reasoning Task: An Exploratory Study in Fermi Problems
von: Liu, Zishuo, et al.
Veröffentlicht: (2025) -
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
von: Le, Hieu Xuan, et al.
Veröffentlicht: (2026) -
LLM for Comparative Narrative Analysis
von: Kampen, Leo, et al.
Veröffentlicht: (2025) -
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)