AI-Mediated Code Comment Improvement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dhakal, Maria, Su, Chia-Yi, Wallace, Robert, Fakhimi, Chris, Bansal, Aakash, Li, Toby, Huang, Yu, McMillan, Collin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910943178391552
author Dhakal, Maria
Su, Chia-Yi
Wallace, Robert
Fakhimi, Chris
Bansal, Aakash
Li, Toby
Huang, Yu
McMillan, Collin
author_facet Dhakal, Maria
Su, Chia-Yi
Wallace, Robert
Fakhimi, Chris
Bansal, Aakash
Li, Toby
Huang, Yu
McMillan, Collin
contents This paper describes an approach to improve code comments along different quality axes by rewriting those comments with customized Artificial Intelligence (AI)-based tools. We conduct an empirical study followed by grounded theory qualitative analysis to determine the quality axes to improve. Then we propose a procedure using a Large Language Model (LLM) to rewrite existing code comments along the quality axes. We implement our procedure using GPT-4o, then distil the results into a smaller model capable of being run in-house, so users can maintain data custody. We evaluate both our approach using GPT-4o and the distilled model versions. We show in an evaluation how our procedure improves code comments along the quality axes. We release all data and source code in an online repository for reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AI-Mediated Code Comment Improvement
Dhakal, Maria
Su, Chia-Yi
Wallace, Robert
Fakhimi, Chris
Bansal, Aakash
Li, Toby
Huang, Yu
McMillan, Collin
Software Engineering
Artificial Intelligence
Programming Languages
This paper describes an approach to improve code comments along different quality axes by rewriting those comments with customized Artificial Intelligence (AI)-based tools. We conduct an empirical study followed by grounded theory qualitative analysis to determine the quality axes to improve. Then we propose a procedure using a Large Language Model (LLM) to rewrite existing code comments along the quality axes. We implement our procedure using GPT-4o, then distil the results into a smaller model capable of being run in-house, so users can maintain data custody. We evaluate both our approach using GPT-4o and the distilled model versions. We show in an evaluation how our procedure improves code comments along the quality axes. We release all data and source code in an online repository for reproducibility.
title AI-Mediated Code Comment Improvement
topic Software Engineering
Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2505.09021