Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nguyen, Linh, Liu, Chunhua, Lin, Hong Yi, Thongtanunam, Patanamon
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918124458082304
author Nguyen, Linh
Liu, Chunhua
Lin, Hong Yi
Thongtanunam, Patanamon
author_facet Nguyen, Linh
Liu, Chunhua
Lin, Hong Yi
Thongtanunam, Patanamon
contents Code review is a crucial practice in software development. As code review nowadays is lightweight, various issues can be identified, and sometimes, they can be trivial. Research has investigated automated approaches to classify review comments to gauge the effectiveness of code reviews. However, previous studies have primarily relied on supervised machine learning, which requires extensive manual annotation to train the models effectively. To address this limitation, we explore the potential of using Large Language Models (LLMs) to classify code review comments. We assess the performance of LLMs to classify 17 categories of code review comments. Our results show that LLMs can classify code review comments, outperforming the state-of-the-art approach using a trained deep learning model. In particular, LLMs achieve better accuracy in classifying the five most useful categories, which the state-of-the-art approach struggles with due to low training examples. Rather than relying solely on a specific small training data distribution, our results show that LLMs provide balanced performance across high- and low-frequency categories. These results suggest that the LLMs could offer a scalable solution for code review analytics to improve the effectiveness of the code review process.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
Nguyen, Linh
Liu, Chunhua
Lin, Hong Yi
Thongtanunam, Patanamon
Software Engineering
Artificial Intelligence
Code review is a crucial practice in software development. As code review nowadays is lightweight, various issues can be identified, and sometimes, they can be trivial. Research has investigated automated approaches to classify review comments to gauge the effectiveness of code reviews. However, previous studies have primarily relied on supervised machine learning, which requires extensive manual annotation to train the models effectively. To address this limitation, we explore the potential of using Large Language Models (LLMs) to classify code review comments. We assess the performance of LLMs to classify 17 categories of code review comments. Our results show that LLMs can classify code review comments, outperforming the state-of-the-art approach using a trained deep learning model. In particular, LLMs achieve better accuracy in classifying the five most useful categories, which the state-of-the-art approach struggles with due to low training examples. Rather than relying solely on a specific small training data distribution, our results show that LLMs provide balanced performance across high- and low-frequency categories. These results suggest that the LLMs could offer a scalable solution for code review analytics to improve the effectiveness of the code review process.
title Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2508.09832