Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Jiayi, Wang, Yanbo, Huang, Yue, Chen, Dongping, Zhang, Qihui, Moniz, Nuno, Gao, Tian, Geyer, Werner, Huang, Chao, Chen, Pin-Yu, Chawla, Nitesh V, Zhang, Xiangliang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Capability-Oriented Training Induced Alignment Risk
by: Zhou, Yujun, et al.
Published: (2026)
by: Zhou, Yujun, et al.
Published: (2026)
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024)
by: Pan, Deng, et al.
Published: (2024)
Conformalized Selective Regression
by: Sokol, Anna, et al.
Published: (2024)
by: Sokol, Anna, et al.
Published: (2024)
Are we making much progress? Revisiting chemical reaction yield prediction from an imbalanced regression perspective
by: Ma, Yihong, et al.
Published: (2024)
by: Ma, Yihong, et al.
Published: (2024)
AnyLoss: Transforming Classification Metrics into Loss Functions
by: Han, Doheon, et al.
Published: (2024)
by: Han, Doheon, et al.
Published: (2024)
Intersectional Divergence: Measuring Fairness in Regression
by: Germino, Joe, et al.
Published: (2025)
by: Germino, Joe, et al.
Published: (2025)
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
by: Sokol, Anna, et al.
Published: (2024)
by: Sokol, Anna, et al.
Published: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Differentially-Private Data Synthetisation for Efficient Re-Identification Risk Control
by: Carvalho, Tânia, et al.
Published: (2022)
by: Carvalho, Tânia, et al.
Published: (2022)
Class-Aware Contrastive Optimization for Imbalanced Text Classification
by: Khvatskii, Grigorii, et al.
Published: (2024)
by: Khvatskii, Grigorii, et al.
Published: (2024)
HonestLLM: Toward an Honest and Helpful Large Language Model
by: Gao, Chujie, et al.
Published: (2024)
by: Gao, Chujie, et al.
Published: (2024)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
by: Guo, Taicheng, et al.
Published: (2026)
by: Guo, Taicheng, et al.
Published: (2026)
Context Attribution with Multi-Armed Bandit Optimization
by: Pan, Deng, et al.
Published: (2025)
by: Pan, Deng, et al.
Published: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Spectral Manifold Harmonization for Graph Imbalanced Regression
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
Genotype-Conditioned Molecular Generation via Evidence-Grounded Multi-Objective Latent Perturbation in Diffusion Models
by: Nogueira, Brenda, et al.
Published: (2026)
by: Nogueira, Brenda, et al.
Published: (2026)
Emergent Social Intelligence Risks in Generative Multi-Agent Systems
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
My Favorite Streamer is an LLM: Discovering, Bonding, and Co-Creating in AI VTuber Fandom
by: Ye, Jiayi, et al.
Published: (2025)
by: Ye, Jiayi, et al.
Published: (2025)
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
UGMAE: A Unified Framework for Graph Masked Autoencoders
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
SPECTRA: Spectral Domain-Aware Graph Generation for Imbalanced Molecular Property Regression
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
by: Bao, Han, et al.
Published: (2024)
by: Bao, Han, et al.
Published: (2024)
DataGen: Unified Synthetic Dataset Generation via Large Language Models
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
by: Zhang, Ziqian, et al.
Published: (2026)
by: Zhang, Ziqian, et al.
Published: (2026)
Quantifying LLM Biases Across Instruction Boundary in Mixed Question Forms
by: Ling, Zipeng, et al.
Published: (2025)
by: Ling, Zipeng, et al.
Published: (2025)
LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected?
by: Zhang, Qihui, et al.
Published: (2024)
by: Zhang, Qihui, et al.
Published: (2024)
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation
by: Zhang, Qihui, et al.
Published: (2025)
by: Zhang, Qihui, et al.
Published: (2025)
Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models
by: Xu, Zixiang, et al.
Published: (2025)
by: Xu, Zixiang, et al.
Published: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
by: Feng, Yuanning, et al.
Published: (2025)
by: Feng, Yuanning, et al.
Published: (2025)
ChefFusion: Multimodal Foundation Model Integrating Recipe and Food Image Generation
by: Li, Peiyu, et al.
Published: (2024)
by: Li, Peiyu, et al.
Published: (2024)
Jailbreaking Large Language Models Through Alignment Vulnerabilities in Out-of-Distribution Settings
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
by: Huang, Hui, et al.
Published: (2026)
by: Huang, Hui, et al.
Published: (2026)
A Survey of Large Language Models for Graphs
by: Ren, Xubin, et al.
Published: (2024)
by: Ren, Xubin, et al.
Published: (2024)
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
by: Ma, Chiyu, et al.
Published: (2025)
by: Ma, Chiyu, et al.
Published: (2025)
Large Language Model based Multi-Agents: A Survey of Progress and Challenges
by: Guo, Taicheng, et al.
Published: (2024)
by: Guo, Taicheng, et al.
Published: (2024)
Similar Items
-
Capability-Oriented Training Induced Alignment Risk
by: Zhou, Yujun, et al.
Published: (2026) -
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024) -
Conformalized Selective Regression
by: Sokol, Anna, et al.
Published: (2024) -
Are we making much progress? Revisiting chemical reaction yield prediction from an imbalanced regression perspective
by: Ma, Yihong, et al.
Published: (2024) -
AnyLoss: Transforming Classification Metrics into Loss Functions
by: Han, Doheon, et al.
Published: (2024)