Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moon, Jiwon, Hwang, Yerin, Lee, Dongryeol, Kang, Taegwan, Kim, Yongil, Jung, Kyomin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911354353352704
author Moon, Jiwon
Hwang, Yerin
Lee, Dongryeol
Kang, Taegwan
Kim, Yongil
Jung, Kyomin
author_facet Moon, Jiwon
Hwang, Yerin
Lee, Dongryeol
Kang, Taegwan
Kim, Yongil
Jung, Kyomin
contents With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While this offers scalability and flexibility, it also raises a critical, unresolved question: Can LLM judges fairly and robustly evaluate semantically equivalent code with superficial variations? Functionally correct code often exhibits variations-such as differences in variable names, comments, or formatting-that should not influence its correctness. Yet, whether LLM judges can reliably handle these variations remains unclear. We present the first comprehensive study of this issue, defining six types of potential bias in code evaluation and revealing their systematic impact on LLM judges. Across five programming languages and multiple LLMs, we empirically demonstrate that all tested LLM judges are susceptible to both positive and negative biases, resulting in inflated or unfairly low scores. Moreover, we observe that LLM judges remain vulnerable to these biases even when prompted to generate test cases before scoring, highlighting the need for more robust code evaluation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
Moon, Jiwon
Hwang, Yerin
Lee, Dongryeol
Kang, Taegwan
Kim, Yongil
Jung, Kyomin
Computation and Language
Software Engineering
With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While this offers scalability and flexibility, it also raises a critical, unresolved question: Can LLM judges fairly and robustly evaluate semantically equivalent code with superficial variations? Functionally correct code often exhibits variations-such as differences in variable names, comments, or formatting-that should not influence its correctness. Yet, whether LLM judges can reliably handle these variations remains unclear. We present the first comprehensive study of this issue, defining six types of potential bias in code evaluation and revealing their systematic impact on LLM judges. Across five programming languages and multiple LLMs, we empirically demonstrate that all tested LLM judges are susceptible to both positive and negative biases, resulting in inflated or unfairly low scores. Moreover, we observe that LLM judges remain vulnerable to these biases even when prompted to generate test cases before scoring, highlighting the need for more robust code evaluation methods.
title Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
topic Computation and Language
Software Engineering
url https://arxiv.org/abs/2505.16222