Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Huang, Hui, Wu, Xuanxin, Yang, Muyun, Arase, Yuki |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
par: Wu, Xuanxin, et autres
Publié: (2025)
par: Wu, Xuanxin, et autres
Publié: (2025)
An In-depth Evaluation of Large Language Models in Sentence Simplification with Error-based Human Assessment
par: Wu, Xuanxin, et autres
Publié: (2024)
par: Wu, Xuanxin, et autres
Publié: (2024)
DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier
par: Huang, Hui, et autres
Publié: (2026)
par: Huang, Hui, et autres
Publié: (2026)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
par: Huang, Hui, et autres
Publié: (2024)
par: Huang, Hui, et autres
Publié: (2024)
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation
par: Miyano, Ryota, et autres
Publié: (2025)
par: Miyano, Ryota, et autres
Publié: (2025)
Distilling Monolingual and Crosslingual Word-in-Context Representations
par: Arase, Yuki, et autres
Publié: (2024)
par: Arase, Yuki, et autres
Publié: (2024)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
par: Zhou, Hongli, et autres
Publié: (2026)
par: Zhou, Hongli, et autres
Publié: (2026)
Hallucinated Span Detection with Multi-View Attention Features
par: Ogasa, Yuya, et autres
Publié: (2025)
par: Ogasa, Yuya, et autres
Publié: (2025)
Aligning Sentence Simplification with ESL Learner's Proficiency for Language Acquisition
par: Li, Guanlin, et autres
Publié: (2025)
par: Li, Guanlin, et autres
Publié: (2025)
Edit-Constrained Decoding for Sentence Simplification
par: Zetsu, Tatsuya, et autres
Publié: (2024)
par: Zetsu, Tatsuya, et autres
Publié: (2024)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
par: Ye, Jiayi, et autres
Publié: (2024)
par: Ye, Jiayi, et autres
Publié: (2024)
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
par: Xu, Mufan, et autres
Publié: (2024)
par: Xu, Mufan, et autres
Publié: (2024)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
par: Moon, Jiwon, et autres
Publié: (2025)
par: Moon, Jiwon, et autres
Publié: (2025)
Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges
par: Li, Weiyuan, et autres
Publié: (2025)
par: Li, Weiyuan, et autres
Publié: (2025)
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
par: Liu, Yihong, et autres
Publié: (2026)
par: Liu, Yihong, et autres
Publié: (2026)
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
par: Kumar, Shachi H, et autres
Publié: (2024)
par: Kumar, Shachi H, et autres
Publié: (2024)
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
par: Liu, Yang, et autres
Publié: (2025)
par: Liu, Yang, et autres
Publié: (2025)
Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoning
par: Xu, Mufan, et autres
Publié: (2025)
par: Xu, Mufan, et autres
Publié: (2025)
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
par: Zhou, Hongli, et autres
Publié: (2026)
par: Zhou, Hongli, et autres
Publié: (2026)
JudgeLRM: Large Reasoning Models as a Judge
par: Chen, Nuo, et autres
Publié: (2025)
par: Chen, Nuo, et autres
Publié: (2025)
GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
par: Mewes, Fabian, et autres
Publié: (2026)
par: Mewes, Fabian, et autres
Publié: (2026)
Imagination Helps Visual Reasoning, But Not Yet in Latent Space
par: Li, You, et autres
Publié: (2026)
par: Li, You, et autres
Publié: (2026)
Self-Evaluation of Large Language Model based on Glass-box Features
par: Huang, Hui, et autres
Publié: (2024)
par: Huang, Hui, et autres
Publié: (2024)
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
par: Ling Team, et autres
Publié: (2025)
par: Ling Team, et autres
Publié: (2025)
Large Language Models Cannot Self-Correct Reasoning Yet
par: Huang, Jie, et autres
Publié: (2023)
par: Huang, Jie, et autres
Publié: (2023)
Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models
par: Xu, Mufan, et autres
Publié: (2026)
par: Xu, Mufan, et autres
Publié: (2026)
Humans or LLMs as the Judge? A Study on Judgement Biases
par: Chen, Guiming Hardy, et autres
Publié: (2024)
par: Chen, Guiming Hardy, et autres
Publié: (2024)
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
par: Zhang, Qiyuan, et autres
Publié: (2025)
par: Zhang, Qiyuan, et autres
Publié: (2025)
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
par: Feng, Zhaoxin, et autres
Publié: (2026)
par: Feng, Zhaoxin, et autres
Publié: (2026)
Mitigating the Bias of Large Language Model Evaluation
par: Zhou, Hongli, et autres
Publié: (2024)
par: Zhou, Hongli, et autres
Publié: (2024)
LLM-based Human Simulations Have Not Yet Been Reliable
par: Wang, Qian, et autres
Publié: (2025)
par: Wang, Qian, et autres
Publié: (2025)
Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
par: Naous, Tarek, et autres
Publié: (2025)
par: Naous, Tarek, et autres
Publié: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
par: Yang, Bo, et autres
Publié: (2026)
par: Yang, Bo, et autres
Publié: (2026)
MR. Judge: Multimodal Reasoner as a Judge
par: Pi, Renjie, et autres
Publié: (2025)
par: Pi, Renjie, et autres
Publié: (2025)
Yet Another Watermark for Large Language Models
par: Bao, Siyuan, et autres
Publié: (2025)
par: Bao, Siyuan, et autres
Publié: (2025)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
par: Bansal, Hritik, et autres
Publié: (2024)
par: Bansal, Hritik, et autres
Publié: (2024)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
par: Xu, Ran, et autres
Publié: (2025)
par: Xu, Ran, et autres
Publié: (2025)
Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
par: Tang, Yihong, et autres
Publié: (2025)
par: Tang, Yihong, et autres
Publié: (2025)
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
par: Zhang, Lan, et autres
Publié: (2025)
par: Zhang, Lan, et autres
Publié: (2025)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
par: Hwang, Yerin, et autres
Publié: (2025)
par: Hwang, Yerin, et autres
Publié: (2025)
Documents similaires
-
Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
par: Wu, Xuanxin, et autres
Publié: (2025) -
An In-depth Evaluation of Large Language Models in Sentence Simplification with Error-based Human Assessment
par: Wu, Xuanxin, et autres
Publié: (2024) -
DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier
par: Huang, Hui, et autres
Publié: (2026) -
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
par: Huang, Hui, et autres
Publié: (2024) -
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation
par: Miyano, Ryota, et autres
Publié: (2025)