The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Siddiq, Mohammed Latif, Dristi, Simantika, Saha, Joy, Santos, Joanna C. S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FRANC: A Lightweight Framework for High-Quality Code Generation
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2026)
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2026)
Large Language Models for Software Engineering: A Reproducibility Crisis
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2025)
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2025)
SALLM: Security Assessment of Generated Code
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
Using Large Language Models to Generate JUnit Tests: An Empirical Study
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
Assessing the Software Security Comprehension of Large Language Models
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
Security and Quality in LLM-Generated Code: A Multi-Language, Multi-Model Analysis
by: Kharma, Mohammed, et al.
Published: (2025)
by: Kharma, Mohammed, et al.
Published: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025)
by: Jiang, Yuancheng, et al.
Published: (2025)
Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
by: Yan, Kaiwen, et al.
Published: (2025)
by: Yan, Kaiwen, et al.
Published: (2025)
Enhancing Code Quality with Generative AI: Boosting Developer Warning Compliance
by: Chang, Hansen, et al.
Published: (2025)
by: Chang, Hansen, et al.
Published: (2025)
Assessing the Quality and Security of AI-Generated Code: A Quantitative Analysis
by: Sabra, Abbas, et al.
Published: (2025)
by: Sabra, Abbas, et al.
Published: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2026)
by: Orel, Daniil, et al.
Published: (2026)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
by: Sharifloo, Amir Molzam, et al.
Published: (2025)
by: Sharifloo, Amir Molzam, et al.
Published: (2025)
Correctness Assessment of Code Generated by Large Language Models Using Internal Representations
by: Bui, Tuan-Dung, et al.
Published: (2025)
by: Bui, Tuan-Dung, et al.
Published: (2025)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
by: Jahangirova, Gunel, et al.
Published: (2024)
by: Jahangirova, Gunel, et al.
Published: (2024)
Bug-Report-Driven Fault Localization: Industrial Benchmarking and Lesson Learned at ABB Robotics
by: Hall, Pernilla, et al.
Published: (2026)
by: Hall, Pernilla, et al.
Published: (2026)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
by: Huang, Jue, et al.
Published: (2026)
by: Huang, Jue, et al.
Published: (2026)
Combining Language and App UI Analysis for the Automated Assessment of Bug Reproduction Steps
by: Mahmud, Junayed, et al.
Published: (2025)
by: Mahmud, Junayed, et al.
Published: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
by: Raghavendra, Mohit, et al.
Published: (2026)
by: Raghavendra, Mohit, et al.
Published: (2026)
SemRep: Generative Code Representation Learning with Code Transformations
by: Li, Weichen, et al.
Published: (2026)
by: Li, Weichen, et al.
Published: (2026)
Think Anywhere in Code Generation
by: Jiang, Xue, et al.
Published: (2026)
by: Jiang, Xue, et al.
Published: (2026)
Code Roulette: How Prompt Variability Affects LLM Code Generation
by: Paleyes, Andrei, et al.
Published: (2025)
by: Paleyes, Andrei, et al.
Published: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
Fault Localization via Fine-tuning Large Language Models with Mutation Generated Stack Traces
by: Jambigi, Neetha, et al.
Published: (2025)
by: Jambigi, Neetha, et al.
Published: (2025)
Generating Realistic, Diverse, and Fault-Revealing Inputs with Latent Space Interpolation for Testing Deep Neural Networks
by: Duan, Bin, et al.
Published: (2025)
by: Duan, Bin, et al.
Published: (2025)
Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models
by: Haroon, Sabaat, et al.
Published: (2025)
by: Haroon, Sabaat, et al.
Published: (2025)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
LLM Performance for Code Generation on Noisy Tasks
by: Sendyka, Radzim, et al.
Published: (2025)
by: Sendyka, Radzim, et al.
Published: (2025)
Code Less, Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruning
by: Tsai, Yun-Da, et al.
Published: (2024)
by: Tsai, Yun-Da, et al.
Published: (2024)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
by: Zheng, Qinkai, et al.
Published: (2023)
by: Zheng, Qinkai, et al.
Published: (2023)
Leveraging Reviewer Experience in Code Review Comment Generation
by: Lin, Hong Yi, et al.
Published: (2024)
by: Lin, Hong Yi, et al.
Published: (2024)
StructCoder: Structure-Aware Transformer for Code Generation
by: Tipirneni, Sindhu, et al.
Published: (2022)
by: Tipirneni, Sindhu, et al.
Published: (2022)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
by: Qiu, Ruizhong, et al.
Published: (2024)
by: Qiu, Ruizhong, et al.
Published: (2024)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
A Feature-Driven Framework for Software Fault Prediction
by: Ghazi, Ahmad Nauman, et al.
Published: (2026)
by: Ghazi, Ahmad Nauman, et al.
Published: (2026)
Where's the Bug? Attention Probing for Scalable Fault Localization
by: Stein, Adam, et al.
Published: (2025)
by: Stein, Adam, et al.
Published: (2025)
Similar Items
-
FRANC: A Lightweight Framework for High-Quality Code Generation
by: Siddiq, Mohammed Latif, et al.
Published: (2023) -
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2026) -
Large Language Models for Software Engineering: A Reproducibility Crisis
by: Siddiq, Mohammed Latif, et al.
Published: (2025) -
Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2025) -
SALLM: Security Assessment of Generated Code
by: Siddiq, Mohammed Latif, et al.
Published: (2023)