Towards Advancing Code Generation with Large Language Models: A Research Roadmap

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jin, Haolin, Chen, Huaming, Lu, Qinghua, Zhu, Liming
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929682378653696
author Jin, Haolin
Chen, Huaming
Lu, Qinghua
Zhu, Liming
author_facet Jin, Haolin
Chen, Huaming
Lu, Qinghua
Zhu, Liming
contents Recently, we have witnessed the rapid development of large language models, which have demonstrated excellent capabilities in the downstream task of code generation. However, despite their potential, LLM-based code generation still faces numerous technical and evaluation challenges, particularly when embedded in real-world development. In this paper, we present our vision for current research directions, and provide an in-depth analysis of existing studies on this task. We propose a six-layer vision framework that categorizes code generation process into distinct phases, namely Input Phase, Orchestration Phase, Development Phase, and Validation Phase. Additionally, we outline our vision workflow, which reflects on the currently prevalent frameworks. We systematically analyse the challenges faced by large language models, including those LLM-based agent frameworks, in code generation tasks. With these, we offer various perspectives and actionable recommendations in this area. Our aim is to provide guidelines for improving the reliability, robustness and usability of LLM-based code generation systems. Ultimately, this work seeks to address persistent challenges and to provide practical suggestions for a more pragmatic LLM-based solution for future code generation endeavors.
format Preprint
id arxiv_https___arxiv_org_abs_2501_11354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Advancing Code Generation with Large Language Models: A Research Roadmap
Jin, Haolin
Chen, Huaming
Lu, Qinghua
Zhu, Liming
Software Engineering
Artificial Intelligence
Recently, we have witnessed the rapid development of large language models, which have demonstrated excellent capabilities in the downstream task of code generation. However, despite their potential, LLM-based code generation still faces numerous technical and evaluation challenges, particularly when embedded in real-world development. In this paper, we present our vision for current research directions, and provide an in-depth analysis of existing studies on this task. We propose a six-layer vision framework that categorizes code generation process into distinct phases, namely Input Phase, Orchestration Phase, Development Phase, and Validation Phase. Additionally, we outline our vision workflow, which reflects on the currently prevalent frameworks. We systematically analyse the challenges faced by large language models, including those LLM-based agent frameworks, in code generation tasks. With these, we offer various perspectives and actionable recommendations in this area. Our aim is to provide guidelines for improving the reliability, robustness and usability of LLM-based code generation systems. Ultimately, this work seeks to address persistent challenges and to provide practical suggestions for a more pragmatic LLM-based solution for future code generation endeavors.
title Towards Advancing Code Generation with Large Language Models: A Research Roadmap
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2501.11354