On Evaluating LLM Alignment by Evaluating LLMs as Judges
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yixin, Liu, Pengfei, Cohan, Arman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
di: Liu, Yixin, et al.
Pubblicazione: (2026)
di: Liu, Yixin, et al.
Pubblicazione: (2026)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
References Improve LLM Alignment in Non-Verifiable Domains
di: Shi, Kejian, et al.
Pubblicazione: (2026)
di: Shi, Kejian, et al.
Pubblicazione: (2026)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
di: Liu, Yixin, et al.
Pubblicazione: (2024)
di: Liu, Yixin, et al.
Pubblicazione: (2024)
Calibrating Long-form Generations from Large Language Models
di: Huang, Yukun, et al.
Pubblicazione: (2024)
di: Huang, Yukun, et al.
Pubblicazione: (2024)
ReIFE: Re-evaluating Instruction-Following Evaluation
di: Liu, Yixin, et al.
Pubblicazione: (2024)
di: Liu, Yixin, et al.
Pubblicazione: (2024)
Understanding Reference Policies in Direct Preference Optimization
di: Liu, Yixin, et al.
Pubblicazione: (2024)
di: Liu, Yixin, et al.
Pubblicazione: (2024)
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
di: Alam, Firoj, et al.
Pubblicazione: (2026)
di: Alam, Firoj, et al.
Pubblicazione: (2026)
Bayesian Calibration of Win Rate Estimation with LLM Evaluators
di: Gao, Yicheng, et al.
Pubblicazione: (2024)
di: Gao, Yicheng, et al.
Pubblicazione: (2024)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
di: Muhamed, Aashiq
Pubblicazione: (2025)
di: Muhamed, Aashiq
Pubblicazione: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025)
di: Collot, Stephane, et al.
Pubblicazione: (2025)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
di: Li, Chuhan, et al.
Pubblicazione: (2024)
di: Li, Chuhan, et al.
Pubblicazione: (2024)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
di: Liu, Zijun, et al.
Pubblicazione: (2024)
di: Liu, Zijun, et al.
Pubblicazione: (2024)
Sample-Efficient Alignment for LLMs
di: Liu, Zichen, et al.
Pubblicazione: (2024)
di: Liu, Zichen, et al.
Pubblicazione: (2024)
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
di: Zhao, Yilun, et al.
Pubblicazione: (2025)
di: Zhao, Yilun, et al.
Pubblicazione: (2025)
Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
di: Jwa, Seungyeon, et al.
Pubblicazione: (2025)
di: Jwa, Seungyeon, et al.
Pubblicazione: (2025)
Context Over Content: Exposing Evaluation Faking in Automated Judges
di: Gupta, Manan, et al.
Pubblicazione: (2026)
di: Gupta, Manan, et al.
Pubblicazione: (2026)
Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes
di: Nath, Abhijnan, et al.
Pubblicazione: (2025)
di: Nath, Abhijnan, et al.
Pubblicazione: (2025)
One-shot Optimized Steering Vectors Mediate Safety-relevant Behaviors in LLMs
di: Dunefsky, Jacob, et al.
Pubblicazione: (2025)
di: Dunefsky, Jacob, et al.
Pubblicazione: (2025)
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
di: Liu, Minqian, et al.
Pubblicazione: (2023)
di: Liu, Minqian, et al.
Pubblicazione: (2023)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
di: Liu, Mingyi
Pubblicazione: (2026)
di: Liu, Mingyi
Pubblicazione: (2026)
AgentBench: Evaluating LLMs as Agents
di: Liu, Xiao, et al.
Pubblicazione: (2023)
di: Liu, Xiao, et al.
Pubblicazione: (2023)
Reformatted Alignment
di: Fan, Run-Ze, et al.
Pubblicazione: (2024)
di: Fan, Run-Ze, et al.
Pubblicazione: (2024)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
di: Wang, Xingyao, et al.
Pubblicazione: (2023)
di: Wang, Xingyao, et al.
Pubblicazione: (2023)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
di: Kim, Jane Paik
Pubblicazione: (2026)
di: Kim, Jane Paik
Pubblicazione: (2026)
Investigating Non-Transitivity in LLM-as-a-Judge
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
FormalAlign: Automated Alignment Evaluation for Autoformalization
di: Lu, Jianqiao, et al.
Pubblicazione: (2024)
di: Lu, Jianqiao, et al.
Pubblicazione: (2024)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
di: Wang, Duo, et al.
Pubblicazione: (2024)
di: Wang, Duo, et al.
Pubblicazione: (2024)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
di: Mahdavi, Sadegh, et al.
Pubblicazione: (2025)
di: Mahdavi, Sadegh, et al.
Pubblicazione: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
di: Yang, Jinming, et al.
Pubblicazione: (2026)
di: Yang, Jinming, et al.
Pubblicazione: (2026)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
di: Liu, Yixin, et al.
Pubblicazione: (2026) -
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
di: Gao, Mingqi, et al.
Pubblicazione: (2024) -
References Improve LLM Alignment in Non-Verifiable Domains
di: Shi, Kejian, et al.
Pubblicazione: (2026) -
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
di: Liu, Yixin, et al.
Pubblicazione: (2024) -
Calibrating Long-form Generations from Large Language Models
di: Huang, Yukun, et al.
Pubblicazione: (2024)