What Types of Code Review Comments Do Developers Most Frequently Resolve?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goldman, Saul, Lin, Hong Yi, Pasuksmit, Jirat, Thongtanunam, Patanamon, Tantithamthavorn, Kla, Wang, Zhe, Zhang, Ray, Behnaz, Ali, Jiang, Fan, Siers, Michael, Jiang, Ryan, Buller, Mike, Jeong, Minwoo, Wu, Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911195429076992
author Goldman, Saul
Lin, Hong Yi
Pasuksmit, Jirat
Thongtanunam, Patanamon
Tantithamthavorn, Kla
Wang, Zhe
Zhang, Ray
Behnaz, Ali
Jiang, Fan
Siers, Michael
Jiang, Ryan
Buller, Mike
Jeong, Minwoo
Wu, Ming
author_facet Goldman, Saul
Lin, Hong Yi
Pasuksmit, Jirat
Thongtanunam, Patanamon
Tantithamthavorn, Kla
Wang, Zhe
Zhang, Ray
Behnaz, Ali
Jiang, Fan
Siers, Michael
Jiang, Ryan
Buller, Mike
Jeong, Minwoo
Wu, Ming
contents Large language model (LLM)-powered code review automation tools have been introduced to generate code review comments. However, not all generated comments will drive code changes. Understanding what types of generated review comments are likely to trigger code changes is crucial for identifying those that are actionable. In this paper, we set out to investigate (1) the types of review comments written by humans and LLMs, and (2) the types of generated comments that are most frequently resolved by developers. To do so, we developed an LLM-as-a-Judge to automatically classify review comments based on our own taxonomy of five categories. Our empirical study confirms that (1) the LLM reviewer and human reviewers exhibit distinct strengths and weaknesses depending on the project context, and (2) readability, bugs, and maintainability-related comments had higher resolution rates than those focused on code design. These results suggest that a substantial proportion of LLM-generated comments are actionable and can be resolved by developers. Our work highlights the complementarity between LLM and human reviewers and offers suggestions to improve the practical effectiveness of LLM-powered code review tools.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Types of Code Review Comments Do Developers Most Frequently Resolve?
Goldman, Saul
Lin, Hong Yi
Pasuksmit, Jirat
Thongtanunam, Patanamon
Tantithamthavorn, Kla
Wang, Zhe
Zhang, Ray
Behnaz, Ali
Jiang, Fan
Siers, Michael
Jiang, Ryan
Buller, Mike
Jeong, Minwoo
Wu, Ming
Software Engineering
Large language model (LLM)-powered code review automation tools have been introduced to generate code review comments. However, not all generated comments will drive code changes. Understanding what types of generated review comments are likely to trigger code changes is crucial for identifying those that are actionable. In this paper, we set out to investigate (1) the types of review comments written by humans and LLMs, and (2) the types of generated comments that are most frequently resolved by developers. To do so, we developed an LLM-as-a-Judge to automatically classify review comments based on our own taxonomy of five categories. Our empirical study confirms that (1) the LLM reviewer and human reviewers exhibit distinct strengths and weaknesses depending on the project context, and (2) readability, bugs, and maintainability-related comments had higher resolution rates than those focused on code design. These results suggest that a substantial proportion of LLM-generated comments are actionable and can be resolved by developers. Our work highlights the complementarity between LLM and human reviewers and offers suggestions to improve the practical effectiveness of LLM-powered code review tools.
title What Types of Code Review Comments Do Developers Most Frequently Resolve?
topic Software Engineering
url https://arxiv.org/abs/2510.05450