Measuring the effectiveness of code review comments in GitHub repositories: A machine learning approach

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rahman, Shadikur, Koana, Umme Ayman, Shanto, Hasibul Karim, Akter, Mahmuda, Roy, Chitra, Ismael, Aras M.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908498777866240
author Rahman, Shadikur
Koana, Umme Ayman
Shanto, Hasibul Karim
Akter, Mahmuda
Roy, Chitra
Ismael, Aras M.
author_facet Rahman, Shadikur
Koana, Umme Ayman
Shanto, Hasibul Karim
Akter, Mahmuda
Roy, Chitra
Ismael, Aras M.
contents This paper illustrates an empirical study of the working efficiency of machine learning techniques in classifying code review text by semantic meaning. The code review comments from the source control repository in GitHub were extracted for development activity from the existing year for three open-source projects. Apart from that, programmers need to be aware of their code and point out their errors. In that case, it is a must to classify the sentiment polarity of the code review comments to avoid an error. We manually labelled 13557 code review comments generated by three open source projects in GitHub during the existing year. In order to recognize the sentiment polarity (or sentiment orientation) of code reviews, we use seven machine learning algorithms and compare those results to find the better ones. Among those Linear Support Vector Classifier(SVC) classifier technique achieves higher accuracy than others. This study will help programmers to make any solution based on code reviews by avoiding misconceptions.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16053
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring the effectiveness of code review comments in GitHub repositories: A machine learning approach
Rahman, Shadikur
Koana, Umme Ayman
Shanto, Hasibul Karim
Akter, Mahmuda
Roy, Chitra
Ismael, Aras M.
Software Engineering
This paper illustrates an empirical study of the working efficiency of machine learning techniques in classifying code review text by semantic meaning. The code review comments from the source control repository in GitHub were extracted for development activity from the existing year for three open-source projects. Apart from that, programmers need to be aware of their code and point out their errors. In that case, it is a must to classify the sentiment polarity of the code review comments to avoid an error. We manually labelled 13557 code review comments generated by three open source projects in GitHub during the existing year. In order to recognize the sentiment polarity (or sentiment orientation) of code reviews, we use seven machine learning algorithms and compare those results to find the better ones. Among those Linear Support Vector Classifier(SVC) classifier technique achieves higher accuracy than others. This study will help programmers to make any solution based on code reviews by avoiding misconceptions.
title Measuring the effectiveness of code review comments in GitHub repositories: A machine learning approach
topic Software Engineering
url https://arxiv.org/abs/2508.16053