DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Junyi, Li, Xiaojia, Hua, Zihan, Yu, Lei, Cheng, Shiqi, Yang, Li, Zhang, Fengjun, Zuo, Chun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916583761248256
author Lu, Junyi
Li, Xiaojia
Hua, Zihan
Yu, Lei
Cheng, Shiqi
Yang, Li
Zhang, Fengjun
Zuo, Chun
author_facet Lu, Junyi
Li, Xiaojia
Hua, Zihan
Yu, Lei
Cheng, Shiqi
Yang, Li
Zhang, Fengjun
Zuo, Chun
contents Code review is a vital but demanding aspect of software development, generating significant interest in automating review comments. Traditional evaluation methods for these comments, primarily based on text similarity, face two major challenges: inconsistent reliability of human-authored comments in open-source projects and the weak correlation of text similarity with objectives like enhancing code quality and detecting defects. This study empirically analyzes benchmark comments using a novel set of criteria informed by prior research and developer interviews. We then similarly revisit the evaluation of existing methodologies. Our evaluation framework, DeepCRCEval, integrates human evaluators and Large Language Models (LLMs) for a comprehensive reassessment of current techniques based on the criteria set. Besides, we also introduce an innovative and efficient baseline, LLM-Reviewer, leveraging the few-shot learning capabilities of LLMs for a target-oriented comparison. Our research highlights the limitations of text similarity metrics, finding that less than 10% of benchmark comments are high quality for automation. In contrast, DeepCRCEval effectively distinguishes between high and low-quality comments, proving to be a more reliable evaluation mechanism. Incorporating LLM evaluators into DeepCRCEval significantly boosts efficiency, reducing time and cost by 88.78% and 90.32%, respectively. Furthermore, LLM-Reviewer demonstrates significant potential of focusing task real targets in comment generation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18291
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation
Lu, Junyi
Li, Xiaojia
Hua, Zihan
Yu, Lei
Cheng, Shiqi
Yang, Li
Zhang, Fengjun
Zuo, Chun
Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
Code review is a vital but demanding aspect of software development, generating significant interest in automating review comments. Traditional evaluation methods for these comments, primarily based on text similarity, face two major challenges: inconsistent reliability of human-authored comments in open-source projects and the weak correlation of text similarity with objectives like enhancing code quality and detecting defects. This study empirically analyzes benchmark comments using a novel set of criteria informed by prior research and developer interviews. We then similarly revisit the evaluation of existing methodologies. Our evaluation framework, DeepCRCEval, integrates human evaluators and Large Language Models (LLMs) for a comprehensive reassessment of current techniques based on the criteria set. Besides, we also introduce an innovative and efficient baseline, LLM-Reviewer, leveraging the few-shot learning capabilities of LLMs for a target-oriented comparison. Our research highlights the limitations of text similarity metrics, finding that less than 10% of benchmark comments are high quality for automation. In contrast, DeepCRCEval effectively distinguishes between high and low-quality comments, proving to be a more reliable evaluation mechanism. Incorporating LLM evaluators into DeepCRCEval significantly boosts efficiency, reducing time and cost by 88.78% and 90.32%, respectively. Furthermore, LLM-Reviewer demonstrates significant potential of focusing task real targets in comment generation.
title DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation
topic Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.18291