Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Moon, Jiwon, Hwang, Yerin, Lee, Dongryeol, Kang, Taegwan, Kim, Yongil, Jung, Kyomin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
by: Lee, Dongryeol, et al.
Published: (2026)
by: Lee, Dongryeol, et al.
Published: (2026)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
by: Lee, Dongryeol, et al.
Published: (2024)
by: Lee, Dongryeol, et al.
Published: (2024)
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
Rethinking Code Refinement: Learning to Judge Code Efficiency
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
by: Zhao, Yuwei, et al.
Published: (2024)
by: Zhao, Yuwei, et al.
Published: (2024)
CodeJudge: Evaluating Code Generation with Large Language Models
by: Tong, Weixi, et al.
Published: (2024)
by: Tong, Weixi, et al.
Published: (2024)
Comparing Developer and LLM Biases in Code Evaluation
by: Mittal, Aditya, et al.
Published: (2026)
by: Mittal, Aditya, et al.
Published: (2026)
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
by: Koo, Jahyun, et al.
Published: (2024)
by: Koo, Jahyun, et al.
Published: (2024)
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
by: Mu, Wenhan, et al.
Published: (2025)
by: Mu, Wenhan, et al.
Published: (2025)
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
by: Weyssow, Martin, et al.
Published: (2024)
by: Weyssow, Martin, et al.
Published: (2024)
LLM-as-a-Judge for Reference-less Automatic Code Validation and Refinement for Natural Language to Bash in IT Automation
by: Vo, Ngoc Phuoc An, et al.
Published: (2025)
by: Vo, Ngoc Phuoc An, et al.
Published: (2025)
Generating Diverse Hypotheses for Inductive Reasoning
by: Lee, Kang-il, et al.
Published: (2024)
by: Lee, Kang-il, et al.
Published: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Coding Agents Don't Know When to Act
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
Evaluating and Achieving Controllable Code Completion in Code LLM
by: Zhang, Jiajun, et al.
Published: (2026)
by: Zhang, Jiajun, et al.
Published: (2026)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025)
by: Yang, Chenyang, et al.
Published: (2025)
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
by: Chon, Heejae, et al.
Published: (2024)
by: Chon, Heejae, et al.
Published: (2024)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
by: Moon, Seungjun, et al.
Published: (2023)
by: Moon, Seungjun, et al.
Published: (2023)
LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
by: Han, Hojae, et al.
Published: (2026)
by: Han, Hojae, et al.
Published: (2026)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
by: Lai, Peng, et al.
Published: (2026)
by: Lai, Peng, et al.
Published: (2026)
From Code to Courtroom: LLMs as the New Software Judges
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
by: Han, Hojae, et al.
Published: (2024)
by: Han, Hojae, et al.
Published: (2024)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
by: Guo, Jiaxun, et al.
Published: (2026)
by: Guo, Jiaxun, et al.
Published: (2026)
Evaluation of Code LLMs on Geospatial Code Generation
by: Gramacki, Piotr, et al.
Published: (2024)
by: Gramacki, Piotr, et al.
Published: (2024)
From Prediction to Application: Language Model-based Code Knowledge Tracing with Domain Adaptive Pre-Training and Automatic Feedback System with Pedagogical Prompting for Comprehensive Programming Education
by: Lee, Unggi, et al.
Published: (2024)
by: Lee, Unggi, et al.
Published: (2024)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
by: Zheng, Zihan, et al.
Published: (2025)
by: Zheng, Zihan, et al.
Published: (2025)
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
by: Yoo, Jaeseok, et al.
Published: (2024)
by: Yoo, Jaeseok, et al.
Published: (2024)
LLM4VV: Exploring LLM-as-a-Judge for Validation and Verification Testsuites
by: Sollenberger, Zachariah, et al.
Published: (2024)
by: Sollenberger, Zachariah, et al.
Published: (2024)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
by: Zhang, Chenchen, et al.
Published: (2025)
by: Zhang, Chenchen, et al.
Published: (2025)
Don't Complete It! Preventing Unhelpful Code Completion for Productive and Sustainable Neural Code Completion Systems
by: Sun, Zhensu, et al.
Published: (2022)
by: Sun, Zhensu, et al.
Published: (2022)
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
VersiCode: Towards Version-controllable Code Generation
by: Wu, Tongtong, et al.
Published: (2024)
by: Wu, Tongtong, et al.
Published: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning
by: Dihan, Mahir Labib, et al.
Published: (2026)
by: Dihan, Mahir Labib, et al.
Published: (2026)
Similar Items
-
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025) -
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
by: Lee, Dongryeol, et al.
Published: (2026) -
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025) -
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
by: Lee, Dongryeol, et al.
Published: (2024) -
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)