Large Language Models (LLMs) for Source Code Analysis: applications, models and datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jelodar, Hamed, Meymani, Mohammad, Razavi-Far, Roozbeh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913751461003264
author Jelodar, Hamed
Meymani, Mohammad
Razavi-Far, Roozbeh
author_facet Jelodar, Hamed
Meymani, Mohammad
Razavi-Far, Roozbeh
contents Large language models (LLMs) and transformer-based architectures are increasingly utilized for source code analysis. As software systems grow in complexity, integrating LLMs into code analysis workflows becomes essential for enhancing efficiency, accuracy, and automation. This paper explores the role of LLMs for different code analysis tasks, focusing on three key aspects: 1) what they can analyze and their applications, 2) what models are used and 3) what datasets are used, and the challenges they face. Regarding the goal of this research, we investigate scholarly articles that explore the use of LLMs for source code analysis to uncover research developments, current trends, and the intellectual structure of this emerging field. Additionally, we summarize limitations and highlight essential tools, datasets, and key challenges, which could be valuable for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models (LLMs) for Source Code Analysis: applications, models and datasets
Jelodar, Hamed
Meymani, Mohammad
Razavi-Far, Roozbeh
Software Engineering
Artificial Intelligence
Computation and Language
Large language models (LLMs) and transformer-based architectures are increasingly utilized for source code analysis. As software systems grow in complexity, integrating LLMs into code analysis workflows becomes essential for enhancing efficiency, accuracy, and automation. This paper explores the role of LLMs for different code analysis tasks, focusing on three key aspects: 1) what they can analyze and their applications, 2) what models are used and 3) what datasets are used, and the challenges they face. Regarding the goal of this research, we investigate scholarly articles that explore the use of LLMs for source code analysis to uncover research developments, current trends, and the intellectual structure of this emerging field. Additionally, we summarize limitations and highlight essential tools, datasets, and key challenges, which could be valuable for future work.
title Large Language Models (LLMs) for Source Code Analysis: applications, models and datasets
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.17502