Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Hongli, Huang, Hui, Zhang, Rui, Chen, Kehai, Xu, Bing, Zhu, Conghui, Zhao, Tiejun, Yang, Muyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating the Bias of Large Language Model Evaluation
di: Zhou, Hongli, et al.
Pubblicazione: (2024)
di: Zhou, Hongli, et al.
Pubblicazione: (2024)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
di: Huang, Hui, et al.
Pubblicazione: (2026)
di: Huang, Hui, et al.
Pubblicazione: (2026)
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
di: Zhou, Hongli, et al.
Pubblicazione: (2025)
di: Zhou, Hongli, et al.
Pubblicazione: (2025)
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
di: Xu, Mufan, et al.
Pubblicazione: (2024)
di: Xu, Mufan, et al.
Pubblicazione: (2024)
LoRA-drop: Efficient LoRA Parameter Pruning based on Output Evaluation
di: Zhou, Hongyun, et al.
Pubblicazione: (2024)
di: Zhou, Hongyun, et al.
Pubblicazione: (2024)
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
di: Huang, Hui, et al.
Pubblicazione: (2025)
di: Huang, Hui, et al.
Pubblicazione: (2025)
Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis
di: Chen, Andong, et al.
Pubblicazione: (2025)
di: Chen, Andong, et al.
Pubblicazione: (2025)
Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoning
di: Xu, Mufan, et al.
Pubblicazione: (2025)
di: Xu, Mufan, et al.
Pubblicazione: (2025)
Self-Evaluation of Large Language Model based on Glass-box Features
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
di: Huang, Hui, et al.
Pubblicazione: (2026)
di: Huang, Hui, et al.
Pubblicazione: (2026)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
User-Aware Active Knowledge Acquisition for Emotional Support Dialogue
di: Xu, Mufan, et al.
Pubblicazione: (2026)
di: Xu, Mufan, et al.
Pubblicazione: (2026)
LLM-based Translation Inference with Iterative Bilingual Understanding
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models
di: Xu, Mufan, et al.
Pubblicazione: (2026)
di: Xu, Mufan, et al.
Pubblicazione: (2026)
Large Language Models for Classical Chinese Poetry Translation: Benchmarking, Evaluating, and Improving
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
di: Liu, Dongxu, et al.
Pubblicazione: (2024)
di: Liu, Dongxu, et al.
Pubblicazione: (2024)
Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
di: Tang, Yihong, et al.
Pubblicazione: (2025)
di: Tang, Yihong, et al.
Pubblicazione: (2025)
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
di: Yang, Bo, et al.
Pubblicazione: (2026)
di: Yang, Bo, et al.
Pubblicazione: (2026)
DUAL-REFLECT: Enhancing Large Language Models for Reflective Translation through Dual Learning Feedback Mechanisms
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
di: Lai, Peng, et al.
Pubblicazione: (2026)
di: Lai, Peng, et al.
Pubblicazione: (2026)
Think-J: Learning to Think for Generative LLM-as-a-Judge
di: Huang, Hui, et al.
Pubblicazione: (2025)
di: Huang, Hui, et al.
Pubblicazione: (2025)
A Survey on Human Preference Learning for Large Language Models
di: Jiang, Ruili, et al.
Pubblicazione: (2024)
di: Jiang, Ruili, et al.
Pubblicazione: (2024)
Evaluating Scoring Bias in LLM-as-a-Judge
di: Li, Qingquan, et al.
Pubblicazione: (2025)
di: Li, Qingquan, et al.
Pubblicazione: (2025)
DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier
di: Huang, Hui, et al.
Pubblicazione: (2026)
di: Huang, Hui, et al.
Pubblicazione: (2026)
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
di: Zhang, Yudi, et al.
Pubblicazione: (2025)
di: Zhang, Yudi, et al.
Pubblicazione: (2025)
Bias Beyond English: Evaluating Social Bias and Debiasing Methods in a Low-Resource Setting
di: Zhou, Ej, et al.
Pubblicazione: (2025)
di: Zhou, Ej, et al.
Pubblicazione: (2025)
Dynamic Planning for LLM-based Graphical User Interface Automation
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
di: Song, Yuchen, et al.
Pubblicazione: (2025)
di: Song, Yuchen, et al.
Pubblicazione: (2025)
Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
di: Yang, Haoyan, et al.
Pubblicazione: (2025)
di: Yang, Haoyan, et al.
Pubblicazione: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
di: Yang, Dingkang, et al.
Pubblicazione: (2024)
di: Yang, Dingkang, et al.
Pubblicazione: (2024)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
di: Yang, Jinming, et al.
Pubblicazione: (2026)
di: Yang, Jinming, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Mitigating the Bias of Large Language Model Evaluation
di: Zhou, Hongli, et al.
Pubblicazione: (2024) -
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024) -
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
di: Huang, Hui, et al.
Pubblicazione: (2026) -
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
di: Zhou, Hongli, et al.
Pubblicazione: (2025) -
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
di: Xu, Mufan, et al.
Pubblicazione: (2024)