Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | Dristi, Simantika Bhattacharjee, Dwyer, Matthew B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2026)
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2026)
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
by: Siddiq, Mohammed Latif, et al.
Published: (2024)
by: Siddiq, Mohammed Latif, et al.
Published: (2024)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
TOGLL: Correct and Strong Test Oracle Generation with LLMs
by: Hossain, Soneya Binta, et al.
Published: (2024)
by: Hossain, Soneya Binta, et al.
Published: (2024)
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
by: Ling, Lin, et al.
Published: (2024)
by: Ling, Lin, et al.
Published: (2024)
Doc2OracLL: Investigating the Impact of Documentation on LLM-based Test Oracle Generation
by: Hossain, Soneya Binta, et al.
Published: (2024)
by: Hossain, Soneya Binta, et al.
Published: (2024)
Mitigating Omitted Variable Bias in Empirical Software Engineering
by: Furia, Carlo A., et al.
Published: (2025)
by: Furia, Carlo A., et al.
Published: (2025)
From Bias To Improved Prompts: A Case Study of Bias Mitigation of Clone Detection Models
by: Chen, QiHong, et al.
Published: (2025)
by: Chen, QiHong, et al.
Published: (2025)
Generating Realistic, Diverse, and Fault-Revealing Inputs with Latent Space Interpolation for Testing Deep Neural Networks
by: Duan, Bin, et al.
Published: (2025)
by: Duan, Bin, et al.
Published: (2025)
Static Code Analyzer Recommendation via Preference Mining
by: Ge, Xiuting, et al.
Published: (2024)
by: Ge, Xiuting, et al.
Published: (2024)
Exploring Multi-Lingual Bias of Large Code Models in Code Generation
by: Wang, Chaozheng, et al.
Published: (2024)
by: Wang, Chaozheng, et al.
Published: (2024)
STADA: Specification-based Testing for Autonomous Driving Agents
by: Saha, Joy, et al.
Published: (2026)
by: Saha, Joy, et al.
Published: (2026)
Analyzing and Mitigating (with LLMs) the Security Misconfigurations of Helm Charts from Artifact Hub
by: Minna, Francesco, et al.
Published: (2024)
by: Minna, Francesco, et al.
Published: (2024)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
by: Du, Yongkang, et al.
Published: (2025)
by: Du, Yongkang, et al.
Published: (2025)
Generating Maximal Configurations and Their Variants Using Code Metrics
by: Yavuz, Tuba, et al.
Published: (2024)
by: Yavuz, Tuba, et al.
Published: (2024)
Social Bias in LLM-Generated Code: Benchmark and Mitigation
by: Rabbi, Fazle, et al.
Published: (2026)
by: Rabbi, Fazle, et al.
Published: (2026)
Can Code Evaluation Metrics Detect Code Plagiarism?
by: Ebrahim, Fahad, et al.
Published: (2026)
by: Ebrahim, Fahad, et al.
Published: (2026)
CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
ChatGPT for Code Refactoring: Analyzing Topics, Interaction, and Effective Prompts
by: AlOmar, Eman Abdullah, et al.
Published: (2025)
by: AlOmar, Eman Abdullah, et al.
Published: (2025)
Analyzing Dependency Distribution Changes Arising from Code Smell Interactions
by: Zhang, Zushuai, et al.
Published: (2025)
by: Zhang, Zushuai, et al.
Published: (2025)
Code-Survey: An LLM-Driven Methodology for Analyzing Large-Scale Codebases
by: Zheng, Yusheng, et al.
Published: (2024)
by: Zheng, Yusheng, et al.
Published: (2024)
Analyzing and Evaluating the Behavior of Git Diff and Merge
by: Glodny, Niels
Published: (2025)
by: Glodny, Niels
Published: (2025)
Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models
by: Xiao, Yisong, et al.
Published: (2025)
by: Xiao, Yisong, et al.
Published: (2025)
NRevisit: A Cognitive Behavioral Metric for Code Understandability Assessment
by: Hao, Gao, et al.
Published: (2025)
by: Hao, Gao, et al.
Published: (2025)
Software Code Quality Measurement: Implications from Metric Distributions
by: Jin, Siyuan, et al.
Published: (2023)
by: Jin, Siyuan, et al.
Published: (2023)
Integrating Code Metrics into Automated Documentation Generation for Computational Notebooks
by: Ghahfarokhi, Mojtaba Mostafavi, et al.
Published: (2026)
by: Ghahfarokhi, Mojtaba Mostafavi, et al.
Published: (2026)
Towards Understanding the Impact of Code Modifications on Software Quality Metrics
by: Karanikiotis, Thomas, et al.
Published: (2024)
by: Karanikiotis, Thomas, et al.
Published: (2024)
Who Introduces and Who Fixes? Analyzing Code Quality in Collaborative Student's Projects
by: Ferrao, Rafael Corsi, et al.
Published: (2025)
by: Ferrao, Rafael Corsi, et al.
Published: (2025)
Exploring Large Language Models for Analyzing and Improving Method Names in Scientific Code
by: Larsen, Gunnar, et al.
Published: (2025)
by: Larsen, Gunnar, et al.
Published: (2025)
PyGress: Tool for Analyzing the Progression of Code Proficiency in Python OSS Projects
by: Charatvaraphan, Rujiphart, et al.
Published: (2025)
by: Charatvaraphan, Rujiphart, et al.
Published: (2025)
Assessing Quality Metrics for Neural Reality Gap Input Mitigation in Autonomous Driving Testing
by: Lambertenghi, Stefano Carlo, et al.
Published: (2024)
by: Lambertenghi, Stefano Carlo, et al.
Published: (2024)
Hallucinations in Code Change to Natural Language Generation: Prevalence and Evaluation of Detection Metrics
by: Liu, Chunhua, et al.
Published: (2025)
by: Liu, Chunhua, et al.
Published: (2025)
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation
by: Yang, Guang, et al.
Published: (2025)
by: Yang, Guang, et al.
Published: (2025)
Mitigating Gender Bias in Code Large Language Models via Model Editing
by: Qin, Zhanyue, et al.
Published: (2024)
by: Qin, Zhanyue, et al.
Published: (2024)
How Reliable Are FOSS Popularity Metrics? Analyzing the Effort Required for Spoofing Common Software Popularity Metrics
by: Swierzy, Ben, et al.
Published: (2025)
by: Swierzy, Ben, et al.
Published: (2025)
Analyzing Prominent LLMs: An Empirical Study of Performance and Complexity in Solving LeetCode Problems
by: Guimaraes, Everton, et al.
Published: (2025)
by: Guimaraes, Everton, et al.
Published: (2025)
AI builds, We Analyze: An Empirical Study of AI-Generated Build Code Quality
by: Ghammam, Anwar, et al.
Published: (2026)
by: Ghammam, Anwar, et al.
Published: (2026)
FAIL: Analyzing Software Failures from the News Using LLMs
by: Anandayuvaraj, Dharun, et al.
Published: (2024)
by: Anandayuvaraj, Dharun, et al.
Published: (2024)
CodeScore: Evaluating Code Generation by Learning Code Execution
by: Dong, Yihong, et al.
Published: (2023)
by: Dong, Yihong, et al.
Published: (2023)
Similar Items
-
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2026) -
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
by: Siddiq, Mohammed Latif, et al.
Published: (2024) -
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023) -
TOGLL: Correct and Strong Test Oracle Generation with LLMs
by: Hossain, Soneya Binta, et al.
Published: (2024) -
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
by: Ling, Lin, et al.
Published: (2024)