Structure-Aware Corpus Construction and User-Perception-Aligned Metrics for Large-Language-Model Code Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Dengfeng, Zhai, Jucai, Jiang, Xiaoguang, Li, Ziqun, Yu, Qianjin, Liu, Feng, Ye, Rui, Liu, Huang, Yang, Zhiguo, Du, Yongsheng, Tan, Fang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912382109876224
author Liu, Dengfeng
Zhai, Jucai
Jiang, Xiaoguang
Li, Ziqun
Yu, Qianjin
Liu, Feng
Ye, Rui
Liu, Huang
Yang, Zhiguo
Du, Yongsheng
Tan, Fang
author_facet Liu, Dengfeng
Zhai, Jucai
Jiang, Xiaoguang
Li, Ziqun
Yu, Qianjin
Liu, Feng
Ye, Rui
Liu, Huang
Yang, Zhiguo
Du, Yongsheng
Tan, Fang
contents Code completion technology based on large language model has significantly improved the development efficiency of programmers. However, in practical applications, there remains a gap between current commonly used code completion evaluation metrics and users' actual perception. To address this issue, we propose two evaluation metrics for code completion tasks--LCP and ROUGE-LCP, from the perspective of probabilistic modeling. Furthermore, to tackle the lack of effective structural semantic modeling and cross-module dependency information in LLMs for repository-level code completion scenarios, we propose a data processing method based on a Structure-Preserving and Semantically-Reordered Code Graph (SPSR-Graph). Through theoretical analysis and experimental validation, we demonstrate the superiority of the proposed evaluation metrics in terms of user perception consistency, as well as the effectiveness of the data processing method in enhancing model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13073
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structure-Aware Corpus Construction and User-Perception-Aligned Metrics for Large-Language-Model Code Completion
Liu, Dengfeng
Zhai, Jucai
Jiang, Xiaoguang
Li, Ziqun
Yu, Qianjin
Liu, Feng
Ye, Rui
Liu, Huang
Yang, Zhiguo
Du, Yongsheng
Tan, Fang
Software Engineering
Artificial Intelligence
Code completion technology based on large language model has significantly improved the development efficiency of programmers. However, in practical applications, there remains a gap between current commonly used code completion evaluation metrics and users' actual perception. To address this issue, we propose two evaluation metrics for code completion tasks--LCP and ROUGE-LCP, from the perspective of probabilistic modeling. Furthermore, to tackle the lack of effective structural semantic modeling and cross-module dependency information in LLMs for repository-level code completion scenarios, we propose a data processing method based on a Structure-Preserving and Semantically-Reordered Code Graph (SPSR-Graph). Through theoretical analysis and experimental validation, we demonstrate the superiority of the proposed evaluation metrics in terms of user perception consistency, as well as the effectiveness of the data processing method in enhancing model performance.
title Structure-Aware Corpus Construction and User-Perception-Aligned Metrics for Large-Language-Model Code Completion
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2505.13073