The Rise of Language Models in Mining Software Repositories: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Romero-Arjona, Miguel, Barakat, Saman, Sánchez, Ana B., Segura, Sergio
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915905496154112
author Romero-Arjona, Miguel
Barakat, Saman
Sánchez, Ana B.
Segura, Sergio
author_facet Romero-Arjona, Miguel
Barakat, Saman
Sánchez, Ana B.
Segura, Sergio
contents The Mining Software Repositories (MSR) field focuses on analysing the rich data contained in software repositories to derive actionable insights into software processes and products. Mining repositories at scale requires techniques capable of handling large volumes of heterogeneous data, a challenge for which language models (LMs) are increasingly well-suited. Since the advent of Transformer-based architectures, LMs have been rapidly adopted across a wide range of MSR tasks. This article presents a comprehensive survey of the use of LMs in MSR, based on an analysis of 85 papers. We examine how LMs are applied, the types of artefacts analysed, which models are used, how their adoption has evolved over time, and the extent to which studies support reproducibility and reuse. Building on this analysis, we propose a taxonomy of LM applications in MSR, identify key trends shaping the field, and highlight open challenges alongside actionable directions for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00787
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Rise of Language Models in Mining Software Repositories: A Survey
Romero-Arjona, Miguel
Barakat, Saman
Sánchez, Ana B.
Segura, Sergio
Software Engineering
The Mining Software Repositories (MSR) field focuses on analysing the rich data contained in software repositories to derive actionable insights into software processes and products. Mining repositories at scale requires techniques capable of handling large volumes of heterogeneous data, a challenge for which language models (LMs) are increasingly well-suited. Since the advent of Transformer-based architectures, LMs have been rapidly adopted across a wide range of MSR tasks. This article presents a comprehensive survey of the use of LMs in MSR, based on an analysis of 85 papers. We examine how LMs are applied, the types of artefacts analysed, which models are used, how their adoption has evolved over time, and the extent to which studies support reproducibility and reuse. Building on this analysis, we propose a taxonomy of LM applications in MSR, identify key trends shaping the field, and highlight open challenges alongside actionable directions for future research.
title The Rise of Language Models in Mining Software Repositories: A Survey
topic Software Engineering
url https://arxiv.org/abs/2604.00787