Comparing Developer and LLM Biases in Code Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Mittal, Aditya, Shar, Ryan, Wu, Zichu, Agarwal, Shyam, Wu, Tongshuang, Donahue, Chris, Talwalkar, Ameet, Chi, Wayne, Chen, Valerie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
by: Chi, Wayne, et al.
Published: (2025)
by: Chi, Wayne, et al.
Published: (2025)
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
by: Chi, Wayne, et al.
Published: (2025)
by: Chi, Wayne, et al.
Published: (2025)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
by: Chen, Valerie, et al.
Published: (2025)
by: Chen, Valerie, et al.
Published: (2025)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
by: Chen, Valerie, et al.
Published: (2026)
by: Chen, Valerie, et al.
Published: (2026)
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
by: He, Keyu, et al.
Published: (2026)
by: He, Keyu, et al.
Published: (2026)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025)
by: Moon, Jiwon, et al.
Published: (2025)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025)
by: Yang, Chenyang, et al.
Published: (2025)
AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development
by: Agarwal, Shyam, et al.
Published: (2026)
by: Agarwal, Shyam, et al.
Published: (2026)
Evaluating and Achieving Controllable Code Completion in Code LLM
by: Zhang, Jiajun, et al.
Published: (2026)
by: Zhang, Jiajun, et al.
Published: (2026)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
by: Zhang, Chenchen, et al.
Published: (2025)
by: Zhang, Chenchen, et al.
Published: (2025)
The Impact of Element Ordering on LM Agent Performance
by: Chi, Wayne, et al.
Published: (2024)
by: Chi, Wayne, et al.
Published: (2024)
CodeAlignBench: Assessing Code Generation Models on Developer-Preferred Code Adjustments
by: Mehralian, Forough, et al.
Published: (2025)
by: Mehralian, Forough, et al.
Published: (2025)
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
by: Zhang, Yilin, et al.
Published: (2025)
by: Zhang, Yilin, et al.
Published: (2025)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
by: Wu, Jiarong, et al.
Published: (2025)
by: Wu, Jiarong, et al.
Published: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
by: Petrukha, Ivan, et al.
Published: (2025)
by: Petrukha, Ivan, et al.
Published: (2025)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
by: Guo, Jiaxun, et al.
Published: (2026)
by: Guo, Jiaxun, et al.
Published: (2026)
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
by: Wu, Jie JW, et al.
Published: (2025)
by: Wu, Jie JW, et al.
Published: (2025)
VersiCode: Towards Version-controllable Code Generation
by: Wu, Tongtong, et al.
Published: (2024)
by: Wu, Tongtong, et al.
Published: (2024)
Evaluation of Code LLMs on Geospatial Code Generation
by: Gramacki, Piotr, et al.
Published: (2024)
by: Gramacki, Piotr, et al.
Published: (2024)
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
by: Mu, Wenhan, et al.
Published: (2025)
by: Mu, Wenhan, et al.
Published: (2025)
RovoDev Code Reviewer: A Large-Scale Online Evaluation of LLM-based Code Review Automation at Atlassian
by: Tantithamthavorn, Kla, et al.
Published: (2026)
by: Tantithamthavorn, Kla, et al.
Published: (2026)
LocAgent: Graph-Guided LLM Agents for Code Localization
by: Chen, Zhaoling, et al.
Published: (2025)
by: Chen, Zhaoling, et al.
Published: (2025)
AlloyASG: Alloy Predicate Code Representation as a Compact Structurally Balanced Graph
by: Wu, Guanxuan, et al.
Published: (2024)
by: Wu, Guanxuan, et al.
Published: (2024)
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
by: Pysklo, Hubert M., et al.
Published: (2026)
by: Pysklo, Hubert M., et al.
Published: (2026)
Dataflow-Guided Retrieval Augmentation for Repository-Level Code Completion
by: Cheng, Wei, et al.
Published: (2024)
by: Cheng, Wei, et al.
Published: (2024)
CodeMirage: Hallucinations in Code Generated by Large Language Models
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
LLM Agents Improve Semantic Code Search
by: Jain, Sarthak, et al.
Published: (2024)
by: Jain, Sarthak, et al.
Published: (2024)
SIMCOPILOT: Evaluating Large Language Models for Copilot-Style Code Generation
by: Jiang, Mingchao, et al.
Published: (2025)
by: Jiang, Mingchao, et al.
Published: (2025)
Python Symbolic Execution with LLM-powered Code Generation
by: Wang, Wenhan, et al.
Published: (2024)
by: Wang, Wenhan, et al.
Published: (2024)
LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning
by: Dihan, Mahir Labib, et al.
Published: (2026)
by: Dihan, Mahir Labib, et al.
Published: (2026)
Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation
by: Wu, Hui, et al.
Published: (2026)
by: Wu, Hui, et al.
Published: (2026)
What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data Slicing
by: Yang, Chenyang, et al.
Published: (2024)
by: Yang, Chenyang, et al.
Published: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
M2rc-Eval: Massively Multilingual Repository-level Code Completion Evaluation
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Measuring LLM Code Generation Stability via Structural Entropy
by: Song, Yewei, et al.
Published: (2025)
by: Song, Yewei, et al.
Published: (2025)
Showing LLM-Generated Code Selectively Based on Confidence of LLMs
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Similar Items
-
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
by: Chi, Wayne, et al.
Published: (2025) -
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
by: Chi, Wayne, et al.
Published: (2025) -
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026) -
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
by: Chen, Valerie, et al.
Published: (2025) -
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
by: Chen, Valerie, et al.
Published: (2026)