How to Get Insight from Disagreeing Post-hoc Explanations in ML-based Defect Prediction Methods? An Empirical Study on Defect Predictors

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Anonymous
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902297523519488
author Anonymous
author_facet Anonymous
contents <p>Machine learning (ML)-based defect prediction models can help practitioners to identify bug-prone modules in large software projects. Identifying bug-prone modules improves resource allocation, lowers maintenance costs, raises software quality, and ensures dependable secure products. However, such defect predictors might not be accepted by practitioners due to a lack of interpretability. Therefore, post-hoc explanation methods such as LIME, SHAP, and BreakDown have gained popularity. These explanation techniques offer insights into the decision-making of ML models by ranking features in order of importance. However, the post-hoc explainability of ML techniques is novel to the Software Engineering (SE) Community; hence, it is unclear whether such methods help practitioners make better decisions regarding software maintenance. Furthermore, recent user studies show that data scientists often employ multiple post-hoc explainers to understand the decision of a single model due to a lack of ground truth datasets. The different techniques approximate the behavior of the model to explain the causes of the disagreement, and because of this disagreement, the usage of the post-hoc method is often confusing for practitioners. In this study, we first investigate disagreements among three explainers: LIME, SHAP, and BreakDown for software defect prediction. Second, we attempt to identify the types of disagreement that occur more frequently than others. Finally, we surveyed 74 practitioners to hear whether they agreed with our findings. The proposition is to aggregate post-hoc explanations, reducing emphasis on disagreements and highlighting areas of agreement. According to the survey responses, 90% (approximately) of the participants confirmed the validation of the proposed aggregation method. Our novel method bridges the gap of disagreements, opening doors for software engineers to extract valuable insights from multiple explanations.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_14752639
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle How to Get Insight from Disagreeing Post-hoc Explanations in ML-based Defect Prediction Methods? An Empirical Study on Defect Predictors
Anonymous
Software Engineering
Defect Prediction
Explainability
Empirical Research
LIME
SHAP
BreakDown
Software Maintenance
<p>Machine learning (ML)-based defect prediction models can help practitioners to identify bug-prone modules in large software projects. Identifying bug-prone modules improves resource allocation, lowers maintenance costs, raises software quality, and ensures dependable secure products. However, such defect predictors might not be accepted by practitioners due to a lack of interpretability. Therefore, post-hoc explanation methods such as LIME, SHAP, and BreakDown have gained popularity. These explanation techniques offer insights into the decision-making of ML models by ranking features in order of importance. However, the post-hoc explainability of ML techniques is novel to the Software Engineering (SE) Community; hence, it is unclear whether such methods help practitioners make better decisions regarding software maintenance. Furthermore, recent user studies show that data scientists often employ multiple post-hoc explainers to understand the decision of a single model due to a lack of ground truth datasets. The different techniques approximate the behavior of the model to explain the causes of the disagreement, and because of this disagreement, the usage of the post-hoc method is often confusing for practitioners. In this study, we first investigate disagreements among three explainers: LIME, SHAP, and BreakDown for software defect prediction. Second, we attempt to identify the types of disagreement that occur more frequently than others. Finally, we surveyed 74 practitioners to hear whether they agreed with our findings. The proposition is to aggregate post-hoc explanations, reducing emphasis on disagreements and highlighting areas of agreement. According to the survey responses, 90% (approximately) of the participants confirmed the validation of the proposed aggregation method. Our novel method bridges the gap of disagreements, opening doors for software engineers to extract valuable insights from multiple explanations.</p>
title How to Get Insight from Disagreeing Post-hoc Explanations in ML-based Defect Prediction Methods? An Empirical Study on Defect Predictors
topic Software Engineering
Defect Prediction
Explainability
Empirical Research
LIME
SHAP
BreakDown
Software Maintenance
url https://doi.org/10.5281/zenodo.14752639