On the Creation of Representative Samples of Software Repositories

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gorostidi, June, Ait, Adem, Cabot, Jordi, Izquierdo, Javier Luis Cánovas
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910628127440896
author Gorostidi, June
Ait, Adem
Cabot, Jordi
Izquierdo, Javier Luis Cánovas
author_facet Gorostidi, June
Ait, Adem
Cabot, Jordi
Izquierdo, Javier Luis Cánovas
contents Software repositories is one of the sources of data in Empirical Software Engineering, primarily in the Mining Software Repositories field, aimed at extracting knowledge from the dynamics and practice of software projects. With the emergence of social coding platforms such as GitHub, researchers have now access to millions of software repositories to use as source data for their studies. With this massive amount of data, sampling techniques are needed to create more manageable datasets. The creation of these datasets is a crucial step, and researchers have to carefully select the repositories to create representative samples according to a set of variables of interest. However, current sampling methods are often based on random selection or rely on variables which may not be related to the research study (e.g., popularity or activity). In this paper, we present a methodology for creating representative samples of software repositories, where such representativeness is properly aligned with both the characteristics of the population of repositories and the requirements of the empirical study. We illustrate our approach with use cases based on Hugging Face repositories.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00639
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Creation of Representative Samples of Software Repositories
Gorostidi, June
Ait, Adem
Cabot, Jordi
Izquierdo, Javier Luis Cánovas
Software Engineering
Software repositories is one of the sources of data in Empirical Software Engineering, primarily in the Mining Software Repositories field, aimed at extracting knowledge from the dynamics and practice of software projects. With the emergence of social coding platforms such as GitHub, researchers have now access to millions of software repositories to use as source data for their studies. With this massive amount of data, sampling techniques are needed to create more manageable datasets. The creation of these datasets is a crucial step, and researchers have to carefully select the repositories to create representative samples according to a set of variables of interest. However, current sampling methods are often based on random selection or rely on variables which may not be related to the research study (e.g., popularity or activity). In this paper, we present a methodology for creating representative samples of software repositories, where such representativeness is properly aligned with both the characteristics of the population of repositories and the requirements of the empirical study. We illustrate our approach with use cases based on Hugging Face repositories.
title On the Creation of Representative Samples of Software Repositories
topic Software Engineering
url https://arxiv.org/abs/2410.00639