High-quality data augmentation for code comment classification

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Borsani, Thomas, Rosani, Andrea, Di Fatta, Giuseppe
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911402386522112
author Borsani, Thomas
Rosani, Andrea
Di Fatta, Giuseppe
author_facet Borsani, Thomas
Rosani, Andrea
Di Fatta, Giuseppe
contents Code comments serve a crucial role in software development for documenting functionality, clarifying design choices, and assisting with issue tracking. They capture developers' insights about the surrounding source code, serving as an essential resource for both human comprehension and automated analysis. Nevertheless, since comments are in natural language, they present challenges for machine-based code understanding. To address this, recent studies have applied natural language processing (NLP) and deep learning techniques to classify comments according to developers' intentions. However, existing datasets for this task suffer from size limitations and class imbalance, as they rely on manual annotations and may not accurately represent the distribution of comments in real-world codebases. To overcome this issue, we introduce new synthetic oversampling and augmentation techniques based on high-quality data generation to enhance the NLBSE'26 challenge datasets. Our Synthetic Quality Oversampling Technique and Augmentation Technique (Q-SYNTH) yield promising results, improving the base classifier by $2.56\%$.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19383
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle High-quality data augmentation for code comment classification
Borsani, Thomas
Rosani, Andrea
Di Fatta, Giuseppe
Software Engineering
Machine Learning
Code comments serve a crucial role in software development for documenting functionality, clarifying design choices, and assisting with issue tracking. They capture developers' insights about the surrounding source code, serving as an essential resource for both human comprehension and automated analysis. Nevertheless, since comments are in natural language, they present challenges for machine-based code understanding. To address this, recent studies have applied natural language processing (NLP) and deep learning techniques to classify comments according to developers' intentions. However, existing datasets for this task suffer from size limitations and class imbalance, as they rely on manual annotations and may not accurately represent the distribution of comments in real-world codebases. To overcome this issue, we introduce new synthetic oversampling and augmentation techniques based on high-quality data generation to enhance the NLBSE'26 challenge datasets. Our Synthetic Quality Oversampling Technique and Augmentation Technique (Q-SYNTH) yield promising results, improving the base classifier by $2.56\%$.
title High-quality data augmentation for code comment classification
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2601.19383