Saved in:
| Main Authors: | Tapwal, Riya, Kumar, Abhishek, Maple, Carsten |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.23970 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DriveSafe: A Hierarchical Risk Taxonomy for Safety-Critical LLM-Based Driving Assistants
by: Kumar, Abhishek, et al.
Published: (2026)
by: Kumar, Abhishek, et al.
Published: (2026)
PRISM: Generation-Time Detection and Mitigation of Secret Leakage in Multi-Agent LLM Pipelines
by: Tapwal, Riya, et al.
Published: (2026)
by: Tapwal, Riya, et al.
Published: (2026)
Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
by: Maple, Carsten, et al.
Published: (2026)
by: Maple, Carsten, et al.
Published: (2026)
Field-Localized Forgery Detection for Digital Identity Documents
by: Kumar, Abhishek, et al.
Published: (2026)
by: Kumar, Abhishek, et al.
Published: (2026)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
by: Zhu, Ziyi, et al.
Published: (2026)
by: Zhu, Ziyi, et al.
Published: (2026)
Evaluating Scoring Bias in LLM-as-a-Judge
by: Li, Qingquan, et al.
Published: (2025)
by: Li, Qingquan, et al.
Published: (2025)
Self-Preference Bias in LLM-as-a-Judge
by: Wataoka, Koki, et al.
Published: (2024)
by: Wataoka, Koki, et al.
Published: (2024)
Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
by: Zhang, Zhenliang, et al.
Published: (2025)
by: Zhang, Zhenliang, et al.
Published: (2025)
Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality
by: Hu, Zhimin, et al.
Published: (2026)
by: Hu, Zhimin, et al.
Published: (2026)
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
by: Liu, Zhuo, et al.
Published: (2025)
by: Liu, Zhuo, et al.
Published: (2025)
On Positional Bias of Faithfulness for Long-form Summarization
by: Wan, David, et al.
Published: (2024)
by: Wan, David, et al.
Published: (2024)
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
by: Kumar, Shachi H, et al.
Published: (2024)
by: Kumar, Shachi H, et al.
Published: (2024)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)
by: Yang, Jinming, et al.
Published: (2026)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
by: Lai, Peng, et al.
Published: (2026)
by: Lai, Peng, et al.
Published: (2026)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
by: Zhou, Hongli, et al.
Published: (2026)
by: Zhou, Hongli, et al.
Published: (2026)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
by: Zhou, Xiaolin, et al.
Published: (2026)
by: Zhou, Xiaolin, et al.
Published: (2026)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
by: Fujinuma, Yoshinari
Published: (2025)
by: Fujinuma, Yoshinari
Published: (2025)
ClickGuard: A Trustworthy Adaptive Fusion Framework for Clickbait Detection
by: Dhiman, Chhavi, et al.
Published: (2026)
by: Dhiman, Chhavi, et al.
Published: (2026)
Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
by: Haldar, Rajarshi, et al.
Published: (2025)
by: Haldar, Rajarshi, et al.
Published: (2025)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
eDIF: A European Deep Inference Fabric for Remote Interpretability of LLM
by: Guggenberger, Irma Heithoff. Marc, et al.
Published: (2025)
by: Guggenberger, Irma Heithoff. Marc, et al.
Published: (2025)
PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
by: Divekar, Abhishek, et al.
Published: (2026)
by: Divekar, Abhishek, et al.
Published: (2026)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
by: Zhang, Hongbin, et al.
Published: (2026)
by: Zhang, Hongbin, et al.
Published: (2026)
A Causal Lens for Evaluating Faithfulness Metrics
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
LLM-as-a-Judge for Time Series Explanations
by: Sivalingam, Preetham, et al.
Published: (2026)
by: Sivalingam, Preetham, et al.
Published: (2026)
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
by: Xu, Yuzheng, et al.
Published: (2026)
by: Xu, Yuzheng, et al.
Published: (2026)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
by: Bhan, Milan, et al.
Published: (2025)
by: Bhan, Milan, et al.
Published: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
by: Lee, Sua, et al.
Published: (2026)
by: Lee, Sua, et al.
Published: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
by: Belmadani, Ikram, et al.
Published: (2026)
by: Belmadani, Ikram, et al.
Published: (2026)
Are LLM Decisions Faithful to Verbal Confidence?
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
by: Li, Chenyu, et al.
Published: (2026)
by: Li, Chenyu, et al.
Published: (2026)
Quantitative LLM Judges
by: Sahoo, Aishwarya, et al.
Published: (2025)
by: Sahoo, Aishwarya, et al.
Published: (2025)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
Similar Items
-
DriveSafe: A Hierarchical Risk Taxonomy for Safety-Critical LLM-Based Driving Assistants
by: Kumar, Abhishek, et al.
Published: (2026) -
PRISM: Generation-Time Detection and Mitigation of Secret Leakage in Multi-Agent LLM Pipelines
by: Tapwal, Riya, et al.
Published: (2026) -
Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
by: Maple, Carsten, et al.
Published: (2026) -
Field-Localized Forgery Detection for Digital Identity Documents
by: Kumar, Abhishek, et al.
Published: (2026) -
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
by: Mittal, Avni, et al.
Published: (2026)