Comparing Developer and LLM Biases in Code Evaluation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mittal, Aditya, Shar, Ryan, Wu, Zichu, Agarwal, Shyam, Wu, Tongshuang, Donahue, Chris, Talwalkar, Ameet, Chi, Wayne, Chen, Valerie
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918501475680256
author Mittal, Aditya
Shar, Ryan
Wu, Zichu
Agarwal, Shyam
Wu, Tongshuang
Donahue, Chris
Talwalkar, Ameet
Chi, Wayne
Chen, Valerie
author_facet Mittal, Aditya
Shar, Ryan
Wu, Zichu
Agarwal, Shyam
Wu, Tongshuang
Donahue, Chris
Talwalkar, Ameet
Chi, Wayne
Chen, Valerie
contents As LLMs are increasingly used as judges in code applications, they should be evaluated in realistic interactive settings that capture partial context and ambiguous intent. We present TRACE (Tool for Rubric Analysis in Code Evaluation), a framework that evaluates LLM judges' ability to predict human preferences and automatically extracts rubric items to reveal systematic biases in how humans and models weigh each item. Across three modalities -- chat-based programming, IDE autocompletion, and instructed code editing -- we use TRACE to measure how well LLM judges align with developer preferences. Among 13 different models, the best judges underperform human annotators by 12-23%. TRACE identifies 35 significant sources of misalignment between humans and judges across interaction modalities, the majority of which correspond to existing software engineering code quality criteria. For example, in chat-based coding, judges are biased towards longer code explanations while humans prefer shorter ones. We find significant misalignment on the majority of existing code quality dimensions, showing alignment gaps between LLM judges and human preference in realistic coding applications.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24586
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Comparing Developer and LLM Biases in Code Evaluation
Mittal, Aditya
Shar, Ryan
Wu, Zichu
Agarwal, Shyam
Wu, Tongshuang
Donahue, Chris
Talwalkar, Ameet
Chi, Wayne
Chen, Valerie
Software Engineering
Computation and Language
As LLMs are increasingly used as judges in code applications, they should be evaluated in realistic interactive settings that capture partial context and ambiguous intent. We present TRACE (Tool for Rubric Analysis in Code Evaluation), a framework that evaluates LLM judges' ability to predict human preferences and automatically extracts rubric items to reveal systematic biases in how humans and models weigh each item. Across three modalities -- chat-based programming, IDE autocompletion, and instructed code editing -- we use TRACE to measure how well LLM judges align with developer preferences. Among 13 different models, the best judges underperform human annotators by 12-23%. TRACE identifies 35 significant sources of misalignment between humans and judges across interaction modalities, the majority of which correspond to existing software engineering code quality criteria. For example, in chat-based coding, judges are biased towards longer code explanations while humans prefer shorter ones. We find significant misalignment on the majority of existing code quality dimensions, showing alignment gaps between LLM judges and human preference in realistic coding applications.
title Comparing Developer and LLM Biases in Code Evaluation
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2603.24586