Designing Empirical Studies on LLM-Based Code Generation: Towards a Reference Framework
Fuente:
arXiv
Guardado en:
| Autores principales: | Nascimento, Nathalia, Guimaraes, Everton, Alencar, Paulo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM4DS: Evaluating Large Language Models for Data Science Code Generation
por: Nascimento, Nathalia, et al.
Publicado: (2024)
por: Nascimento, Nathalia, et al.
Publicado: (2024)
Analyzing Prominent LLMs: An Empirical Study of Performance and Complexity in Solving LeetCode Problems
por: Guimaraes, Everton, et al.
Publicado: (2025)
por: Guimaraes, Everton, et al.
Publicado: (2025)
CodeSift: An LLM-Based Reference-Less Framework for Automatic Code Validation
por: Aggarwal, Pooja, et al.
Publicado: (2024)
por: Aggarwal, Pooja, et al.
Publicado: (2024)
Learn to Code Sustainably: An Empirical Study on LLM-based Green Code Generation
por: Vartziotis, Tina, et al.
Publicado: (2024)
por: Vartziotis, Tina, et al.
Publicado: (2024)
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
por: Imran, Mia Mohammad, et al.
Publicado: (2025)
por: Imran, Mia Mohammad, et al.
Publicado: (2025)
Automated Non-Functional Requirements Generation in Software Engineering with Large Language Models: A Comparative Study
por: Almonte, Jomar Thomas, et al.
Publicado: (2025)
por: Almonte, Jomar Thomas, et al.
Publicado: (2025)
AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Dependency Gaps in LLM-Based Coding Agents
por: Vangala, Bhanu Prakash, et al.
Publicado: (2025)
por: Vangala, Bhanu Prakash, et al.
Publicado: (2025)
Hallucination in LLM-Based Code Generation: An Automotive Case Study
por: Pavel, Marc, et al.
Publicado: (2025)
por: Pavel, Marc, et al.
Publicado: (2025)
Bugs in Large Language Models Generated Code: An Empirical Study
por: Tambon, Florian, et al.
Publicado: (2024)
por: Tambon, Florian, et al.
Publicado: (2024)
On the Effectiveness of LLMs for Manual Test Verifications
por: Peixoto, Myron David Lucena Campos, et al.
Publicado: (2024)
por: Peixoto, Myron David Lucena Campos, et al.
Publicado: (2024)
LLM-Based Robustness Testing of Microservice Applications: An Empirical Study
por: Tigulla, Hrushitha Goud, et al.
Publicado: (2026)
por: Tigulla, Hrushitha Goud, et al.
Publicado: (2026)
CodeCircuit: Toward Inferring LLM-Generated Code Correctness via Attribution Graphs
por: He, Yicheng, et al.
Publicado: (2026)
por: He, Yicheng, et al.
Publicado: (2026)
An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
por: Gandhi, Shubham, et al.
Publicado: (2025)
por: Gandhi, Shubham, et al.
Publicado: (2025)
Beyond Autoregression: An Empirical Study of Diffusion Large Language Models for Code Generation
por: Li, Chengze, et al.
Publicado: (2025)
por: Li, Chengze, et al.
Publicado: (2025)
A Performance Study of LLM-Generated Code on Leetcode
por: Coignion, Tristan, et al.
Publicado: (2024)
por: Coignion, Tristan, et al.
Publicado: (2024)
Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
por: Chakroborti, Apu Kumar, et al.
Publicado: (2025)
por: Chakroborti, Apu Kumar, et al.
Publicado: (2025)
An Empirical Evaluation of LLM-Based Approaches for Code Vulnerability Detection: RAG, SFT, and Dual-Agent Systems
por: Saju, Md Hasan, et al.
Publicado: (2026)
por: Saju, Md Hasan, et al.
Publicado: (2026)
An Empirical Study on Self-correcting Large Language Models for Data Science Code Generation
por: Quoc, Thai Tang, et al.
Publicado: (2024)
por: Quoc, Thai Tang, et al.
Publicado: (2024)
An Empirical Study of Knowledge Distillation for Code Understanding Tasks
por: Wang, Ruiqi, et al.
Publicado: (2025)
por: Wang, Ruiqi, et al.
Publicado: (2025)
Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
por: Acharya, Jagrit, et al.
Publicado: (2025)
por: Acharya, Jagrit, et al.
Publicado: (2025)
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
por: Yu, Jiongchi, et al.
Publicado: (2025)
por: Yu, Jiongchi, et al.
Publicado: (2025)
Towards Responsible Generative AI: A Reference Architecture for Designing Foundation Model based Agents
por: Lu, Qinghua, et al.
Publicado: (2023)
por: Lu, Qinghua, et al.
Publicado: (2023)
How Do Agents Perform Code Optimization? An Empirical Study
por: Peng, Huiyun, et al.
Publicado: (2025)
por: Peng, Huiyun, et al.
Publicado: (2025)
Agentic Frameworks for Reasoning Tasks: An Empirical Study
por: Rasheed, Zeeshan, et al.
Publicado: (2026)
por: Rasheed, Zeeshan, et al.
Publicado: (2026)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
por: d'Aloisio, Giordano, et al.
Publicado: (2024)
por: d'Aloisio, Giordano, et al.
Publicado: (2024)
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
por: Akli, Amal, et al.
Publicado: (2026)
por: Akli, Amal, et al.
Publicado: (2026)
Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
por: Majgaonkar, Oorja, et al.
Publicado: (2025)
por: Majgaonkar, Oorja, et al.
Publicado: (2025)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
por: Li, Lehui, et al.
Publicado: (2026)
por: Li, Lehui, et al.
Publicado: (2026)
An Empirical Study on Capability of Large Language Models in Understanding Code Semantics
por: Nguyen, Thu-Trang, et al.
Publicado: (2024)
por: Nguyen, Thu-Trang, et al.
Publicado: (2024)
AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers
por: Lin, Zijie, et al.
Publicado: (2025)
por: Lin, Zijie, et al.
Publicado: (2025)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
por: Meng, Xiangxin, et al.
Publicado: (2024)
por: Meng, Xiangxin, et al.
Publicado: (2024)
Investigating The Smells of LLM Generated Code
por: Paul, Debalina Ghosh, et al.
Publicado: (2025)
por: Paul, Debalina Ghosh, et al.
Publicado: (2025)
Is Vibe Coding the Future? An Empirical Assessment of LLM Generated Codes for Construction Safety
por: Uddin, S M Jamil
Publicado: (2026)
por: Uddin, S M Jamil
Publicado: (2026)
Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study
por: Bansal, Kaushal
Publicado: (2026)
por: Bansal, Kaushal
Publicado: (2026)
Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
por: Dong, Zeming, et al.
Publicado: (2023)
por: Dong, Zeming, et al.
Publicado: (2023)
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
por: Dong, Zeming, et al.
Publicado: (2024)
por: Dong, Zeming, et al.
Publicado: (2024)
Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence
por: Geruslu, Vehid, et al.
Publicado: (2026)
por: Geruslu, Vehid, et al.
Publicado: (2026)
Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies
por: Angermeir, Florian, et al.
Publicado: (2025)
por: Angermeir, Florian, et al.
Publicado: (2025)
Towards Specification-Driven LLM-Based Generation of Embedded Automotive Software
por: Patil, Minal Suresh, et al.
Publicado: (2024)
por: Patil, Minal Suresh, et al.
Publicado: (2024)
Do Prompt Patterns Affect Code Quality? A First Empirical Assessment of ChatGPT-Generated Code
por: Della Porta, Antonio, et al.
Publicado: (2025)
por: Della Porta, Antonio, et al.
Publicado: (2025)
Ejemplares similares
-
LLM4DS: Evaluating Large Language Models for Data Science Code Generation
por: Nascimento, Nathalia, et al.
Publicado: (2024) -
Analyzing Prominent LLMs: An Empirical Study of Performance and Complexity in Solving LeetCode Problems
por: Guimaraes, Everton, et al.
Publicado: (2025) -
CodeSift: An LLM-Based Reference-Less Framework for Automatic Code Validation
por: Aggarwal, Pooja, et al.
Publicado: (2024) -
Learn to Code Sustainably: An Empirical Study on LLM-based Green Code Generation
por: Vartziotis, Tina, et al.
Publicado: (2024) -
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
por: Imran, Mia Mohammad, et al.
Publicado: (2025)