Defect Prediction with Content-based Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pham, Hung Viet, Nguyen, Tung Thanh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917788218556416
author Pham, Hung Viet
Nguyen, Tung Thanh
author_facet Pham, Hung Viet
Nguyen, Tung Thanh
contents Traditional defect prediction approaches often use metrics that measure the complexity of the design or implementing code of a software system, such as the number of lines of code in a source file. In this paper, we explore a different approach based on content of source code. Our key assumption is that source code of a software system contains information about its technical aspects and those aspects might have different levels of defect-proneness. Thus, content-based features such as words, topics, data types, and package names extracted from a source code file could be used to predict its defects. We have performed an extensive empirical evaluation and found that: i) such content-based features have higher predictive power than code complexity metrics and ii) the use of feature selection, reduction, and combination further improves the prediction performance.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18365
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Defect Prediction with Content-based Features
Pham, Hung Viet
Nguyen, Tung Thanh
Software Engineering
Computation and Language
Machine Learning
Traditional defect prediction approaches often use metrics that measure the complexity of the design or implementing code of a software system, such as the number of lines of code in a source file. In this paper, we explore a different approach based on content of source code. Our key assumption is that source code of a software system contains information about its technical aspects and those aspects might have different levels of defect-proneness. Thus, content-based features such as words, topics, data types, and package names extracted from a source code file could be used to predict its defects. We have performed an extensive empirical evaluation and found that: i) such content-based features have higher predictive power than code complexity metrics and ii) the use of feature selection, reduction, and combination further improves the prediction performance.
title Defect Prediction with Content-based Features
topic Software Engineering
Computation and Language
Machine Learning
url https://arxiv.org/abs/2409.18365