BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lai, Peng, Ou, Zhihao, Wang, Yong, Wang, Longyue, Yang, Jian, Chen, Yun, Chen, Guanhua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
von: Zhao, Zixiao, et al.
Veröffentlicht: (2026)
von: Zhao, Zixiao, et al.
Veröffentlicht: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
von: Ghosh, Himel, et al.
Veröffentlicht: (2026)
von: Ghosh, Himel, et al.
Veröffentlicht: (2026)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
von: Chen, Zaoyu, et al.
Veröffentlicht: (2025)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2025)
Mitigating Gender Bias in Code Large Language Models via Model Editing
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024)
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024)
Efficient Fairness Testing in Large Language Models: Prioritizing Metamorphic Relations for Bias Detection
von: Giramata, Suavis, et al.
Veröffentlicht: (2025)
von: Giramata, Suavis, et al.
Veröffentlicht: (2025)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
Bias Testing and Mitigation in LLM-based Code Generation
von: Huang, Dong, et al.
Veröffentlicht: (2023)
von: Huang, Dong, et al.
Veröffentlicht: (2023)
Executing as You Generate: Hiding Execution Latency in LLM Code Generation
von: Sun, Zhensu, et al.
Veröffentlicht: (2026)
von: Sun, Zhensu, et al.
Veröffentlicht: (2026)
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
von: Ling, Lin, et al.
Veröffentlicht: (2024)
von: Ling, Lin, et al.
Veröffentlicht: (2024)
Towards Fair Machine Learning Software: Understanding and Addressing Model Bias Through Counterfactual Thinking
von: Wang, Zichong, et al.
Veröffentlicht: (2023)
von: Wang, Zichong, et al.
Veröffentlicht: (2023)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
von: Moon, Jiwon, et al.
Veröffentlicht: (2025)
von: Moon, Jiwon, et al.
Veröffentlicht: (2025)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
von: Huang, Yuheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuheng, et al.
Veröffentlicht: (2025)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
ACECODER: Acing Coder RL via Automated Test-Case Synthesis
von: Zeng, Huaye, et al.
Veröffentlicht: (2025)
von: Zeng, Huaye, et al.
Veröffentlicht: (2025)
Towards Automated Data Sciences with Natural Language and SageCopilot: Practices and Lessons Learned
von: Liao, Yuan, et al.
Veröffentlicht: (2024)
von: Liao, Yuan, et al.
Veröffentlicht: (2024)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
von: Chen, Kexin, et al.
Veröffentlicht: (2024)
von: Chen, Kexin, et al.
Veröffentlicht: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications
von: Shen, Leming, et al.
Veröffentlicht: (2025)
von: Shen, Leming, et al.
Veröffentlicht: (2025)
Social Bias in LLM-Generated Code: Benchmark and Mitigation
von: Rabbi, Fazle, et al.
Veröffentlicht: (2026)
von: Rabbi, Fazle, et al.
Veröffentlicht: (2026)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
von: Zhang, Ziyao, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyao, et al.
Veröffentlicht: (2024)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
von: Chen, Haorui, et al.
Veröffentlicht: (2025)
von: Chen, Haorui, et al.
Veröffentlicht: (2025)
RovoDev Code Reviewer: A Large-Scale Online Evaluation of LLM-based Code Review Automation at Atlassian
von: Tantithamthavorn, Kla, et al.
Veröffentlicht: (2026)
von: Tantithamthavorn, Kla, et al.
Veröffentlicht: (2026)
LLM-as-a-Judge for Reference-less Automatic Code Validation and Refinement for Natural Language to Bash in IT Automation
von: Vo, Ngoc Phuoc An, et al.
Veröffentlicht: (2025)
von: Vo, Ngoc Phuoc An, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
von: Huang, Yue, et al.
Veröffentlicht: (2026)
von: Huang, Yue, et al.
Veröffentlicht: (2026)
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
von: De Michele, William, et al.
Veröffentlicht: (2025)
von: De Michele, William, et al.
Veröffentlicht: (2025)
Evaluating and Achieving Controllable Code Completion in Code LLM
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
von: Lee, Seongmin, et al.
Veröffentlicht: (2025)
von: Lee, Seongmin, et al.
Veröffentlicht: (2025)
Towards an Understanding of Context Utilization in Code Intelligence
von: Wang, Yanlin, et al.
Veröffentlicht: (2025)
von: Wang, Yanlin, et al.
Veröffentlicht: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
von: Peng, Yun, et al.
Veröffentlicht: (2024)
von: Peng, Yun, et al.
Veröffentlicht: (2024)
A Taxonomy of Prompt Defects in LLM Systems
von: Tian, Haoye, et al.
Veröffentlicht: (2025)
von: Tian, Haoye, et al.
Veröffentlicht: (2025)
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
von: Chen, Simin, et al.
Veröffentlicht: (2024)
von: Chen, Simin, et al.
Veröffentlicht: (2024)
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
von: Dong, Honghua, et al.
Veröffentlicht: (2025)
von: Dong, Honghua, et al.
Veröffentlicht: (2025)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
Rethinking Code Refinement: Learning to Judge Code Efficiency
von: Seo, Minju, et al.
Veröffentlicht: (2024)
von: Seo, Minju, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
von: Zhao, Zixiao, et al.
Veröffentlicht: (2026) -
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025) -
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
von: Li, Yuanyang, et al.
Veröffentlicht: (2026) -
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
von: Ghosh, Himel, et al.
Veröffentlicht: (2026) -
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)