A Methodological Framework for LLM-Based Mining of Software Repositories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De Martino, Vincenzo, Castaño, Joel, Palomba, Fabio, Franch, Xavier, Martínez-Fernández, Silverio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912532239745024
author De Martino, Vincenzo
Castaño, Joel
Palomba, Fabio
Franch, Xavier
Martínez-Fernández, Silverio
author_facet De Martino, Vincenzo
Castaño, Joel
Palomba, Fabio
Franch, Xavier
Martínez-Fernández, Silverio
contents Large Language Models (LLMs) are increasingly used in software engineering research, offering new opportunities for automating repository mining tasks. However, despite their growing popularity, the methodological integration of LLMs into Mining Software Repositories (MSR) remains poorly understood. Existing studies tend to focus on specific capabilities or performance benchmarks, providing limited insight into how researchers utilize LLMs across the full research pipeline. To address this gap, we conduct a mixed-method study that combines a rapid review and questionnaire survey in the field of LLM4MSR. We investigate (1) the approaches and (2) the threats that affect the empirical rigor of researchers involved in this field. Our findings reveal 15 methodological approaches, nine main threats, and 25 mitigation strategies. Building on these findings, we present PRIMES 2.0, a refined empirical framework organized into six stages, comprising 23 methodological substeps, each mapped to specific threats and corresponding mitigation strategies, providing prescriptive and adaptive support throughout the lifecycle of LLM-based MSR studies. Our work contributes to establishing a more transparent and reproducible foundation for LLM-based MSR research.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02233
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Methodological Framework for LLM-Based Mining of Software Repositories
De Martino, Vincenzo
Castaño, Joel
Palomba, Fabio
Franch, Xavier
Martínez-Fernández, Silverio
Software Engineering
Large Language Models (LLMs) are increasingly used in software engineering research, offering new opportunities for automating repository mining tasks. However, despite their growing popularity, the methodological integration of LLMs into Mining Software Repositories (MSR) remains poorly understood. Existing studies tend to focus on specific capabilities or performance benchmarks, providing limited insight into how researchers utilize LLMs across the full research pipeline. To address this gap, we conduct a mixed-method study that combines a rapid review and questionnaire survey in the field of LLM4MSR. We investigate (1) the approaches and (2) the threats that affect the empirical rigor of researchers involved in this field. Our findings reveal 15 methodological approaches, nine main threats, and 25 mitigation strategies. Building on these findings, we present PRIMES 2.0, a refined empirical framework organized into six stages, comprising 23 methodological substeps, each mapped to specific threats and corresponding mitigation strategies, providing prescriptive and adaptive support throughout the lifecycle of LLM-based MSR studies. Our work contributes to establishing a more transparent and reproducible foundation for LLM-based MSR research.
title A Methodological Framework for LLM-Based Mining of Software Repositories
topic Software Engineering
url https://arxiv.org/abs/2508.02233