An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Fuente:
arXiv
Guardado en:
| Autores principales: | Yamauchi, Yusuke, Yano, Taro, Oyamada, Masafumi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
por: Ishibashi, Yoichi, et al.
Publicado: (2025)
por: Ishibashi, Yoichi, et al.
Publicado: (2025)
LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
por: Yano, Taro, et al.
Publicado: (2025)
por: Yano, Taro, et al.
Publicado: (2025)
Can Large Language Models Invent Algorithms to Improve Themselves?: Algorithm Discovery for Recursive Self-Improvement through Reinforcement Learning
por: Ishibashi, Yoichi, et al.
Publicado: (2024)
por: Ishibashi, Yoichi, et al.
Publicado: (2024)
Can a Crow Hatch a Falcon? Lineage Matters in Predicting Large Language Model Performance
por: Tamura, Takuya, et al.
Publicado: (2025)
por: Tamura, Takuya, et al.
Publicado: (2025)
Effective Harness Engineering for Algorithm Discovery with Coding Agents
por: Ishibashi, Yoichi, et al.
Publicado: (2026)
por: Ishibashi, Yoichi, et al.
Publicado: (2026)
Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework
por: Enomoto, Masafumi, et al.
Publicado: (2026)
por: Enomoto, Masafumi, et al.
Publicado: (2026)
How Reliable is Multilingual LLM-as-a-Judge?
por: Fu, Xiyan, et al.
Publicado: (2025)
por: Fu, Xiyan, et al.
Publicado: (2025)
Read More, Think More: Revisiting Observation Reduction for Web Agents
por: Enomoto, Masafumi, et al.
Publicado: (2026)
por: Enomoto, Masafumi, et al.
Publicado: (2026)
$M^3$ Scaling Law: Optimizing Multi-Epoch, Multi-Lingual, and Multi-Stage Training for Low-Resource Language Models
por: Akimoto, Kosuke, et al.
Publicado: (2024)
por: Akimoto, Kosuke, et al.
Publicado: (2024)
Context Quality Matters in Training Fusion-in-Decoder for Extractive Open-Domain Question Answering
por: Akimoto, Kosuke, et al.
Publicado: (2024)
por: Akimoto, Kosuke, et al.
Publicado: (2024)
LightPAL: Lightweight Passage Retrieval for Open Domain Multi-Document Summarization
por: Enomoto, Masafumi, et al.
Publicado: (2024)
por: Enomoto, Masafumi, et al.
Publicado: (2024)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
por: Huang, Hui, et al.
Publicado: (2024)
por: Huang, Hui, et al.
Publicado: (2024)
Understanding the Impact of Confidence in Retrieval Augmented Generation: A Case Study in the Medical Domain
por: Ozaki, Shintaro, et al.
Publicado: (2024)
por: Ozaki, Shintaro, et al.
Publicado: (2024)
Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality
por: Harada, Yuto, et al.
Publicado: (2025)
por: Harada, Yuto, et al.
Publicado: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
por: Bologna, Federica, et al.
Publicado: (2026)
por: Bologna, Federica, et al.
Publicado: (2026)
How Sensitive Are Safety Benchmarks to Judge Configuration Choices?
por: Zhang, Xinran
Publicado: (2026)
por: Zhang, Xinran
Publicado: (2026)
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
por: Schroeder, Kayla, et al.
Publicado: (2024)
por: Schroeder, Kayla, et al.
Publicado: (2024)
How to Correctly Report LLM-as-a-Judge Evaluations
por: Lee, Chungpa, et al.
Publicado: (2025)
por: Lee, Chungpa, et al.
Publicado: (2025)
Are Emotions Arranged in a Circle? Geometric Analysis of Emotion Representations via Hyperspherical Contrastive Learning
por: Yamauchi, Yusuke, et al.
Publicado: (2026)
por: Yamauchi, Yusuke, et al.
Publicado: (2026)
Reliability Crisis of Reference-free Metrics for Grammatical Error Correction
por: Goto, Takumi, et al.
Publicado: (2025)
por: Goto, Takumi, et al.
Publicado: (2025)
Jellyfish: A Large Language Model for Data Preprocessing
por: Zhang, Haochen, et al.
Publicado: (2023)
por: Zhang, Haochen, et al.
Publicado: (2023)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
por: Zhou, Yuhang, et al.
Publicado: (2025)
por: Zhou, Yuhang, et al.
Publicado: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
por: Belmadani, Ikram, et al.
Publicado: (2026)
por: Belmadani, Ikram, et al.
Publicado: (2026)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
por: Wang, Yidong, et al.
Publicado: (2025)
por: Wang, Yidong, et al.
Publicado: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
por: Li, Qingquan, et al.
Publicado: (2025)
por: Li, Qingquan, et al.
Publicado: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
por: Hong, Yihan, et al.
Publicado: (2026)
por: Hong, Yihan, et al.
Publicado: (2026)
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human?
por: Goto, Takumi, et al.
Publicado: (2025)
por: Goto, Takumi, et al.
Publicado: (2025)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
por: Zhu, Ziyi, et al.
Publicado: (2026)
por: Zhu, Ziyi, et al.
Publicado: (2026)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
por: Shi, Lin, et al.
Publicado: (2024)
por: Shi, Lin, et al.
Publicado: (2024)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
por: Wang, Qian, et al.
Publicado: (2025)
por: Wang, Qian, et al.
Publicado: (2025)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
por: Cavalin, Paulo, et al.
Publicado: (2025)
por: Cavalin, Paulo, et al.
Publicado: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
por: Zhou, Yilun, et al.
Publicado: (2025)
por: Zhou, Yilun, et al.
Publicado: (2025)
Edit-level Majority Voting Mitigates Over-Correction in LLM-based Grammatical Error Correction
por: Goto, Takumi, et al.
Publicado: (2026)
por: Goto, Takumi, et al.
Publicado: (2026)
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
por: Bavaresco, Anna, et al.
Publicado: (2024)
por: Bavaresco, Anna, et al.
Publicado: (2024)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
por: Lee, Dongryeol, et al.
Publicado: (2026)
por: Lee, Dongryeol, et al.
Publicado: (2026)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
por: Chen, Junjie, et al.
Publicado: (2026)
por: Chen, Junjie, et al.
Publicado: (2026)
Grammatical Error Correction Evaluation by Optimally Transporting Edit Representation
por: Goto, Takumi, et al.
Publicado: (2026)
por: Goto, Takumi, et al.
Publicado: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation
por: Goto, Takumi, et al.
Publicado: (2025)
por: Goto, Takumi, et al.
Publicado: (2025)
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
por: Zhang, Ziyin, et al.
Publicado: (2024)
por: Zhang, Ziyin, et al.
Publicado: (2024)
Ejemplares similares
-
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
por: Ishibashi, Yoichi, et al.
Publicado: (2025) -
LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
por: Yano, Taro, et al.
Publicado: (2025) -
Can Large Language Models Invent Algorithms to Improve Themselves?: Algorithm Discovery for Recursive Self-Improvement through Reinforcement Learning
por: Ishibashi, Yoichi, et al.
Publicado: (2024) -
Can a Crow Hatch a Falcon? Lineage Matters in Predicting Large Language Model Performance
por: Tamura, Takuya, et al.
Publicado: (2025) -
Effective Harness Engineering for Algorithm Discovery with Coding Agents
por: Ishibashi, Yoichi, et al.
Publicado: (2026)