Quality In, Quality Out: Investigating Training Data's Role in AI Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Improta, Cristina, Tufano, Rosalia, Liguori, Pietro, Cotroneo, Domenico, Bavota, Gabriele |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human-Written vs. AI-Generated Code: A Large-Scale Study of Defects, Vulnerabilities, and Complexity
by: Cotroneo, Domenico, et al.
Published: (2025)
by: Cotroneo, Domenico, et al.
Published: (2025)
Studying Quality Improvements Recommended via Manual and Automated Code Review
by: Crupi, Giuseppe, et al.
Published: (2026)
by: Crupi, Giuseppe, et al.
Published: (2026)
Automating Code Review: A Systematic Literature Review
by: Tufano, Rosalia, et al.
Published: (2025)
by: Tufano, Rosalia, et al.
Published: (2025)
AI Code Generators for Security: Friend or Foe?
by: Natella, Roberto, et al.
Published: (2024)
by: Natella, Roberto, et al.
Published: (2024)
Automating the Correctness Assessment of AI-generated Code for Security Contexts
by: Cotroneo, Domenico, et al.
Published: (2023)
by: Cotroneo, Domenico, et al.
Published: (2023)
Improving Code Generation via Small Language Model-as-a-judge
by: Crupi, Giuseppe, et al.
Published: (2026)
by: Crupi, Giuseppe, et al.
Published: (2026)
SEART Data Hub: Streamlining Large-Scale Source Code Mining and Pre-Processing
by: Dabić, Ozren, et al.
Published: (2024)
by: Dabić, Ozren, et al.
Published: (2024)
Enhancing AI-based Generation of Software Exploits with Contextual Information
by: Liguori, Pietro, et al.
Published: (2024)
by: Liguori, Pietro, et al.
Published: (2024)
Neural Fault Injection: Generating Software Faults from Natural Language
by: Cotroneo, Domenico, et al.
Published: (2024)
by: Cotroneo, Domenico, et al.
Published: (2024)
DeVAIC: A Tool for Security Assessment of AI-generated Code
by: Cotroneo, Domenico, et al.
Published: (2024)
by: Cotroneo, Domenico, et al.
Published: (2024)
On the Generalizability of Transformer Models to Code Completions of Different Lengths
by: Cooper, Nathan, et al.
Published: (2025)
by: Cooper, Nathan, et al.
Published: (2025)
Leveraging Reward Models for Guiding Code Review Comment Generation
by: Sghaier, Oussama Ben, et al.
Published: (2025)
by: Sghaier, Oussama Ben, et al.
Published: (2025)
Towards Summarizing Code Snippets Using Pre-Trained Transformers
by: Mastropaolo, Antonio, et al.
Published: (2024)
by: Mastropaolo, Antonio, et al.
Published: (2024)
Code Review Automation: Strengths and Weaknesses of the State of the Art
by: Tufano, Rosalia, et al.
Published: (2024)
by: Tufano, Rosalia, et al.
Published: (2024)
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
by: Crupi, Giuseppe, et al.
Published: (2025)
by: Crupi, Giuseppe, et al.
Published: (2025)
Guidelines to Prompt Large Language Models for Code Generation: An Empirical Characterization
by: Midolo, Alessandro, et al.
Published: (2026)
by: Midolo, Alessandro, et al.
Published: (2026)
Detecting Stealthy Data Poisoning Attacks in AI Code Generators
by: Improta, Cristina
Published: (2025)
by: Improta, Cristina
Published: (2025)
Deep Learning-based Code Reviews: A Paradigm Shift or a Double-Edged Sword?
by: Tufano, Rosalia, et al.
Published: (2024)
by: Tufano, Rosalia, et al.
Published: (2024)
PyResBugs: A Dataset of Residual Python Bugs for Natural Language-Driven Fault Injection
by: Cotroneo, Domenico, et al.
Published: (2025)
by: Cotroneo, Domenico, et al.
Published: (2025)
Developers and Generative AI: A Study of Self-Admitted Usage in Open Source Projects
by: Tufano, Rosalia, et al.
Published: (2026)
by: Tufano, Rosalia, et al.
Published: (2026)
What Makes Software Bugs Escape Testing? Evidence from a Large-Scale Empirical Study
by: Cotroneo, Domenico, et al.
Published: (2026)
by: Cotroneo, Domenico, et al.
Published: (2026)
Unveiling ChatGPT's Usage in Open Source Projects: A Mining-based Study
by: Tufano, Rosalia, et al.
Published: (2024)
by: Tufano, Rosalia, et al.
Published: (2024)
Vulnerabilities in AI Code Generators: Exploring Targeted Data Poisoning Attacks
by: Cotroneo, Domenico, et al.
Published: (2023)
by: Cotroneo, Domenico, et al.
Published: (2023)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
Enhancing Code Generation for Low-Resource Languages: No Silver Bullet
by: Giagnorio, Alessandro, et al.
Published: (2025)
by: Giagnorio, Alessandro, et al.
Published: (2025)
Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
by: Giagnorio, Alessandro, et al.
Published: (2025)
by: Giagnorio, Alessandro, et al.
Published: (2025)
Deep Learning-based Code Completion: On the Impact on Performance of Contextual Information
by: Ciniselli, Matteo, et al.
Published: (2025)
by: Ciniselli, Matteo, et al.
Published: (2025)
Enhancing Robustness of AI Offensive Code Generators via Data Augmentation
by: Improta, Cristina, et al.
Published: (2023)
by: Improta, Cristina, et al.
Published: (2023)
Why Personalizing Deep Learning-Based Code Completion Tools Matters
by: Giagnorio, Alessandro, et al.
Published: (2025)
by: Giagnorio, Alessandro, et al.
Published: (2025)
On the Generalizability of Deep Learning-based Code Completion Across Programming Language Versions
by: Ciniselli, Matteo, et al.
Published: (2024)
by: Ciniselli, Matteo, et al.
Published: (2024)
Investigating Execution-Aware Language Models for Code Optimization
by: Di Menna, Federico, et al.
Published: (2025)
by: Di Menna, Federico, et al.
Published: (2025)
How the Training Procedure Impacts the Performance of Deep Learning-based Vulnerability Patching
by: Mastropaolo, Antonio, et al.
Published: (2024)
by: Mastropaolo, Antonio, et al.
Published: (2024)
Enhancing Code Annotation Reliability: Generative AI's Role in Comment Quality Assessment Models
by: Killivalavan, Seetharam, et al.
Published: (2024)
by: Killivalavan, Seetharam, et al.
Published: (2024)
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
by: Steenhoek, Benjamin, et al.
Published: (2023)
by: Steenhoek, Benjamin, et al.
Published: (2023)
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
by: Steenhoek, Benjamin, et al.
Published: (2024)
by: Steenhoek, Benjamin, et al.
Published: (2024)
Will It Break in Production? Metric-Driven Prediction of Residual Defects in Python Systems
by: De Rosa, Giuseppe, et al.
Published: (2026)
by: De Rosa, Giuseppe, et al.
Published: (2026)
On the Quality of AI-Generated Source Code Comments: A Comprehensive Evaluation
by: Guelman, Ian, et al.
Published: (2024)
by: Guelman, Ian, et al.
Published: (2024)
AI builds, We Analyze: An Empirical Study of AI-Generated Build Code Quality
by: Ghammam, Anwar, et al.
Published: (2026)
by: Ghammam, Anwar, et al.
Published: (2026)
CGP-Tuning: Structure-Aware Soft Prompt Tuning for Code Vulnerability Detection
by: Feng, Ruijun, et al.
Published: (2025)
by: Feng, Ruijun, et al.
Published: (2025)
SATORI: Static Test Oracle Generation for REST APIs
by: Alonso, Juan C., et al.
Published: (2025)
by: Alonso, Juan C., et al.
Published: (2025)
Similar Items
-
Human-Written vs. AI-Generated Code: A Large-Scale Study of Defects, Vulnerabilities, and Complexity
by: Cotroneo, Domenico, et al.
Published: (2025) -
Studying Quality Improvements Recommended via Manual and Automated Code Review
by: Crupi, Giuseppe, et al.
Published: (2026) -
Automating Code Review: A Systematic Literature Review
by: Tufano, Rosalia, et al.
Published: (2025) -
AI Code Generators for Security: Friend or Foe?
by: Natella, Roberto, et al.
Published: (2024) -
Automating the Correctness Assessment of AI-generated Code for Security Contexts
by: Cotroneo, Domenico, et al.
Published: (2023)