Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Haoyan, Bao, Runxue, Xiao, Cao, Ma, Jun, Bhatia, Parminder, Gao, Shangqian, Kass-Hout, Taha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Uncertainty Ranking: Enhancing Retrieval-Augmented In-Context Learning for Long-Tail Knowledge in LLMs
by: Yu, Shuyang, et al.
Published: (2024)
by: Yu, Shuyang, et al.
Published: (2024)
MedHEval: Benchmarking Hallucinations and Mitigation Strategies in Medical Large Vision-Language Models
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
by: Zhou, Sicheng, et al.
Published: (2025)
by: Zhou, Sicheng, et al.
Published: (2025)
Hyper Hawkes Processes: Interpretable Models of Marked Temporal Point Processes
by: Boyd, Alex, et al.
Published: (2025)
by: Boyd, Alex, et al.
Published: (2025)
Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval
by: Jiang, Pengcheng, et al.
Published: (2024)
by: Jiang, Pengcheng, et al.
Published: (2024)
Segment as You Wish -- Free-Form Language-Based Segmentation for Medical Images
by: Da, Longchao, et al.
Published: (2024)
by: Da, Longchao, et al.
Published: (2024)
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
Deep Continuous-Time State-Space Models for Marked Event Sequences
by: Chang, Yuxin, et al.
Published: (2024)
by: Chang, Yuxin, et al.
Published: (2024)
Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image Segmentation
by: Konwer, Aishik, et al.
Published: (2025)
by: Konwer, Aishik, et al.
Published: (2025)
Bi-level Contrastive Learning for Knowledge-Enhanced Molecule Representations
by: Jiang, Pengcheng, et al.
Published: (2023)
by: Jiang, Pengcheng, et al.
Published: (2023)
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
by: Li, Chenliang, et al.
Published: (2025)
by: Li, Chenliang, et al.
Published: (2025)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
by: Wang, Zhepeng, et al.
Published: (2024)
by: Wang, Zhepeng, et al.
Published: (2024)
TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale
by: Jiang, Pengcheng, et al.
Published: (2024)
by: Jiang, Pengcheng, et al.
Published: (2024)
BIPEFT: Budget-Guided Iterative Search for Parameter Efficient Fine-Tuning of Large Pretrained Language Models
by: Chang, Aofei, et al.
Published: (2024)
by: Chang, Aofei, et al.
Published: (2024)
Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
by: Yang, Zhijian, et al.
Published: (2025)
by: Yang, Zhijian, et al.
Published: (2025)
Auto-Train-Once: Controller Network Guided Automatic Network Pruning from Scratch
by: Wu, Xidong, et al.
Published: (2024)
by: Wu, Xidong, et al.
Published: (2024)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
KG-FIT: Knowledge Graph Fine-Tuning Upon Open-World Knowledge
by: Jiang, Pengcheng, et al.
Published: (2024)
by: Jiang, Pengcheng, et al.
Published: (2024)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
by: Zhou, Hongli, et al.
Published: (2026)
by: Zhou, Hongli, et al.
Published: (2026)
UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Capability Self-Assessment: Teaching LLMs to Know Their Limits
by: Yang, Haoyan, et al.
Published: (2026)
by: Yang, Haoyan, et al.
Published: (2026)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
Transfer Learning with Clinical Concept Embeddings from Large Language Models
by: Gao, Yuhe, et al.
Published: (2024)
by: Gao, Yuhe, et al.
Published: (2024)
Debate on Graph: a Flexible and Reliable Reasoning Framework for Large Language Models
by: Ma, Jie, et al.
Published: (2024)
by: Ma, Jie, et al.
Published: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
by: Zhao, Yuwei, et al.
Published: (2024)
by: Zhao, Yuwei, et al.
Published: (2024)
BiasFilter: An Inference-Time Debiasing Framework for Large Language Models
by: Cheng, Xiaoqing, et al.
Published: (2025)
by: Cheng, Xiaoqing, et al.
Published: (2025)
Debiasing Classifiers by Amplifying Bias with Latent Diffusion and Large Language Models
by: Ko, Donggeun, et al.
Published: (2024)
by: Ko, Donggeun, et al.
Published: (2024)
Safe Screening Rules for Group SLOPE
by: Bao, Runxue, et al.
Published: (2025)
by: Bao, Runxue, et al.
Published: (2025)
Safe Screening Rules for Group OWL Models
by: Bao, Runxue, et al.
Published: (2025)
by: Bao, Runxue, et al.
Published: (2025)
All-in-One Tuning and Structural Pruning for Domain-Specific LLMs
by: Lu, Lei, et al.
Published: (2024)
by: Lu, Lei, et al.
Published: (2024)
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
by: Le, Khiem, et al.
Published: (2026)
by: Le, Khiem, et al.
Published: (2026)
Causal Debiasing for Visual Commonsense Reasoning
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
Judge Anything: MLLM as a Judge Across Any Modality
by: Pu, Shu, et al.
Published: (2025)
by: Pu, Shu, et al.
Published: (2025)
Dynamic Noise Preference Optimization: Self-Improvement of Large Language Models with Self-Synthetic Data
by: Yang, Haoyan, et al.
Published: (2025)
by: Yang, Haoyan, et al.
Published: (2025)
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
Debias Can be Unreliable: Mitigating Bias Issue in Evaluating Debiasing Recommendation
by: Wang, Chengbing, et al.
Published: (2024)
by: Wang, Chengbing, et al.
Published: (2024)
FortisAVQA and MAVEN: a Benchmark Dataset and Debiasing Framework for Robust Multimodal Reasoning
by: Ma, Jie, et al.
Published: (2025)
by: Ma, Jie, et al.
Published: (2025)
Similar Items
-
Dynamic Uncertainty Ranking: Enhancing Retrieval-Augmented In-Context Learning for Long-Tail Knowledge in LLMs
by: Yu, Shuyang, et al.
Published: (2024) -
MedHEval: Benchmarking Hallucinations and Mitigation Strategies in Medical Large Vision-Language Models
by: Chang, Aofei, et al.
Published: (2025) -
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
by: Zhou, Sicheng, et al.
Published: (2025) -
Hyper Hawkes Processes: Interpretable Models of Marked Temporal Point Processes
by: Boyd, Alex, et al.
Published: (2025) -
Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval
by: Jiang, Pengcheng, et al.
Published: (2024)