MADE-WIC: Multiple Annotated Datasets for Exploring Weaknesses In Code

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mock, Moritz, Melegati, Jorge, Kretschmann, Max, Ferreyra, Nicolás E. Díaz, Russo, Barbara
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910800565764096
author Mock, Moritz
Melegati, Jorge
Kretschmann, Max
Ferreyra, Nicolás E. Díaz
Russo, Barbara
author_facet Mock, Moritz
Melegati, Jorge
Kretschmann, Max
Ferreyra, Nicolás E. Díaz
Russo, Barbara
contents In this paper, we present MADE-WIC, a large dataset of functions and their comments with multiple annotations for technical debt and code weaknesses leveraging different state-of-the-art approaches. It contains about 860K code functions and more than 2.7M related comments from 12 open-source projects. To the best of our knowledge, no such dataset is publicly available. MADE-WIC aims to provide researchers with a curated dataset on which to test and compare tools designed for the detection of code weaknesses and technical debt. As we have fused existing datasets, researchers have the possibility to evaluate the performance of their tools by also controlling the bias related to the annotation definition and dataset construction. The demonstration video can be retrieved at https://www.youtube.com/watch?v=GaQodPrcb6E.
format Preprint
id arxiv_https___arxiv_org_abs_2408_05163
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MADE-WIC: Multiple Annotated Datasets for Exploring Weaknesses In Code
Mock, Moritz
Melegati, Jorge
Kretschmann, Max
Ferreyra, Nicolás E. Díaz
Russo, Barbara
Software Engineering
In this paper, we present MADE-WIC, a large dataset of functions and their comments with multiple annotations for technical debt and code weaknesses leveraging different state-of-the-art approaches. It contains about 860K code functions and more than 2.7M related comments from 12 open-source projects. To the best of our knowledge, no such dataset is publicly available. MADE-WIC aims to provide researchers with a curated dataset on which to test and compare tools designed for the detection of code weaknesses and technical debt. As we have fused existing datasets, researchers have the possibility to evaluate the performance of their tools by also controlling the bias related to the annotation definition and dataset construction. The demonstration video can be retrieved at https://www.youtube.com/watch?v=GaQodPrcb6E.
title MADE-WIC: Multiple Annotated Datasets for Exploring Weaknesses In Code
topic Software Engineering
url https://arxiv.org/abs/2408.05163