The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Siddiq, Mohammed Latif, Dristi, Simantika, Saha, Joy, Santos, Joanna C. S. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FRANC: A Lightweight Framework for High-Quality Code Generation
por: Siddiq, Mohammed Latif, et al.
Publicado: (2023)
por: Siddiq, Mohammed Latif, et al.
Publicado: (2023)
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
por: Dristi, Simantika Bhattacharjee, et al.
Publicado: (2026)
por: Dristi, Simantika Bhattacharjee, et al.
Publicado: (2026)
Large Language Models for Software Engineering: A Reproducibility Crisis
por: Siddiq, Mohammed Latif, et al.
Publicado: (2025)
por: Siddiq, Mohammed Latif, et al.
Publicado: (2025)
Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
por: Dristi, Simantika Bhattacharjee, et al.
Publicado: (2025)
por: Dristi, Simantika Bhattacharjee, et al.
Publicado: (2025)
SALLM: Security Assessment of Generated Code
por: Siddiq, Mohammed Latif, et al.
Publicado: (2023)
por: Siddiq, Mohammed Latif, et al.
Publicado: (2023)
Using Large Language Models to Generate JUnit Tests: An Empirical Study
por: Siddiq, Mohammed Latif, et al.
Publicado: (2023)
por: Siddiq, Mohammed Latif, et al.
Publicado: (2023)
Assessing the Software Security Comprehension of Large Language Models
por: Siddiq, Mohammed Latif, et al.
Publicado: (2025)
por: Siddiq, Mohammed Latif, et al.
Publicado: (2025)
An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems
por: Siddiq, Mohammed Latif, et al.
Publicado: (2026)
por: Siddiq, Mohammed Latif, et al.
Publicado: (2026)
Security and Quality in LLM-Generated Code: A Multi-Language, Multi-Model Analysis
por: Kharma, Mohammed, et al.
Publicado: (2025)
por: Kharma, Mohammed, et al.
Publicado: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
por: Jiang, Yuancheng, et al.
Publicado: (2025)
por: Jiang, Yuancheng, et al.
Publicado: (2025)
Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub
por: Siddiq, Mohammed Latif, et al.
Publicado: (2026)
por: Siddiq, Mohammed Latif, et al.
Publicado: (2026)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
por: Yan, Kaiwen, et al.
Publicado: (2025)
por: Yan, Kaiwen, et al.
Publicado: (2025)
Enhancing Code Quality with Generative AI: Boosting Developer Warning Compliance
por: Chang, Hansen, et al.
Publicado: (2025)
por: Chang, Hansen, et al.
Publicado: (2025)
Assessing the Quality and Security of AI-Generated Code: A Quantitative Analysis
por: Sabra, Abbas, et al.
Publicado: (2025)
por: Sabra, Abbas, et al.
Publicado: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
por: Orel, Daniil, et al.
Publicado: (2026)
por: Orel, Daniil, et al.
Publicado: (2026)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
por: Sharifloo, Amir Molzam, et al.
Publicado: (2025)
por: Sharifloo, Amir Molzam, et al.
Publicado: (2025)
Correctness Assessment of Code Generated by Large Language Models Using Internal Representations
por: Bui, Tuan-Dung, et al.
Publicado: (2025)
por: Bui, Tuan-Dung, et al.
Publicado: (2025)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
por: Jahangirova, Gunel, et al.
Publicado: (2024)
por: Jahangirova, Gunel, et al.
Publicado: (2024)
Bug-Report-Driven Fault Localization: Industrial Benchmarking and Lesson Learned at ABB Robotics
por: Hall, Pernilla, et al.
Publicado: (2026)
por: Hall, Pernilla, et al.
Publicado: (2026)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
por: Huang, Jue, et al.
Publicado: (2026)
por: Huang, Jue, et al.
Publicado: (2026)
Combining Language and App UI Analysis for the Automated Assessment of Bug Reproduction Steps
por: Mahmud, Junayed, et al.
Publicado: (2025)
por: Mahmud, Junayed, et al.
Publicado: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
por: Raghavendra, Mohit, et al.
Publicado: (2026)
por: Raghavendra, Mohit, et al.
Publicado: (2026)
SemRep: Generative Code Representation Learning with Code Transformations
por: Li, Weichen, et al.
Publicado: (2026)
por: Li, Weichen, et al.
Publicado: (2026)
Think Anywhere in Code Generation
por: Jiang, Xue, et al.
Publicado: (2026)
por: Jiang, Xue, et al.
Publicado: (2026)
Code Roulette: How Prompt Variability Affects LLM Code Generation
por: Paleyes, Andrei, et al.
Publicado: (2025)
por: Paleyes, Andrei, et al.
Publicado: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
por: Galimzyanov, Timur, et al.
Publicado: (2024)
por: Galimzyanov, Timur, et al.
Publicado: (2024)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
por: Daghighfarsoodeh, Alireza, et al.
Publicado: (2025)
por: Daghighfarsoodeh, Alireza, et al.
Publicado: (2025)
Fault Localization via Fine-tuning Large Language Models with Mutation Generated Stack Traces
por: Jambigi, Neetha, et al.
Publicado: (2025)
por: Jambigi, Neetha, et al.
Publicado: (2025)
Generating Realistic, Diverse, and Fault-Revealing Inputs with Latent Space Interpolation for Testing Deep Neural Networks
por: Duan, Bin, et al.
Publicado: (2025)
por: Duan, Bin, et al.
Publicado: (2025)
Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models
por: Haroon, Sabaat, et al.
Publicado: (2025)
por: Haroon, Sabaat, et al.
Publicado: (2025)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
por: Xie, Zichen, et al.
Publicado: (2026)
por: Xie, Zichen, et al.
Publicado: (2026)
LLM Performance for Code Generation on Noisy Tasks
por: Sendyka, Radzim, et al.
Publicado: (2025)
por: Sendyka, Radzim, et al.
Publicado: (2025)
Code Less, Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruning
por: Tsai, Yun-Da, et al.
Publicado: (2024)
por: Tsai, Yun-Da, et al.
Publicado: (2024)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
por: Zheng, Qinkai, et al.
Publicado: (2023)
por: Zheng, Qinkai, et al.
Publicado: (2023)
Leveraging Reviewer Experience in Code Review Comment Generation
por: Lin, Hong Yi, et al.
Publicado: (2024)
por: Lin, Hong Yi, et al.
Publicado: (2024)
StructCoder: Structure-Aware Transformer for Code Generation
por: Tipirneni, Sindhu, et al.
Publicado: (2022)
por: Tipirneni, Sindhu, et al.
Publicado: (2022)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
por: Qiu, Ruizhong, et al.
Publicado: (2024)
por: Qiu, Ruizhong, et al.
Publicado: (2024)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
A Feature-Driven Framework for Software Fault Prediction
por: Ghazi, Ahmad Nauman, et al.
Publicado: (2026)
por: Ghazi, Ahmad Nauman, et al.
Publicado: (2026)
Where's the Bug? Attention Probing for Scalable Fault Localization
por: Stein, Adam, et al.
Publicado: (2025)
por: Stein, Adam, et al.
Publicado: (2025)
Ejemplares similares
-
FRANC: A Lightweight Framework for High-Quality Code Generation
por: Siddiq, Mohammed Latif, et al.
Publicado: (2023) -
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
por: Dristi, Simantika Bhattacharjee, et al.
Publicado: (2026) -
Large Language Models for Software Engineering: A Reproducibility Crisis
por: Siddiq, Mohammed Latif, et al.
Publicado: (2025) -
Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
por: Dristi, Simantika Bhattacharjee, et al.
Publicado: (2025) -
SALLM: Security Assessment of Generated Code
por: Siddiq, Mohammed Latif, et al.
Publicado: (2023)