Accurate estimation of feature importance faithfulness for tree models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Gajewski, Mateusz, Karczmarz, Adam, Rapicki, Mateusz, Sankowski, Piotr
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912244074283008
author Gajewski, Mateusz
Karczmarz, Adam
Rapicki, Mateusz
Sankowski, Piotr
author_facet Gajewski, Mateusz
Karczmarz, Adam
Rapicki, Mateusz
Sankowski, Piotr
contents In this paper, we consider a perturbation-based metric of predictive faithfulness of feature rankings (or attributions) that we call PGI squared. When applied to decision tree-based regression models, the metric can be computed accurately and efficiently for arbitrary independent feature perturbation distributions. In particular, the computation does not involve Monte Carlo sampling that has been typically used for computing similar metrics and which is inherently prone to inaccuracies. Moreover, we propose a method of ranking features by their importance for the tree model's predictions based on PGI squared. Our experiments indicate that in some respects, the method may identify the globally important features better than the state-of-the-art SHAP explainer
format Preprint
id arxiv_https___arxiv_org_abs_2404_03426
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accurate estimation of feature importance faithfulness for tree models
Gajewski, Mateusz
Karczmarz, Adam
Rapicki, Mateusz
Sankowski, Piotr
Machine Learning
In this paper, we consider a perturbation-based metric of predictive faithfulness of feature rankings (or attributions) that we call PGI squared. When applied to decision tree-based regression models, the metric can be computed accurately and efficiently for arbitrary independent feature perturbation distributions. In particular, the computation does not involve Monte Carlo sampling that has been typically used for computing similar metrics and which is inherently prone to inaccuracies. Moreover, we propose a method of ranking features by their importance for the tree model's predictions based on PGI squared. Our experiments indicate that in some respects, the method may identify the globally important features better than the state-of-the-art SHAP explainer
title Accurate estimation of feature importance faithfulness for tree models
topic Machine Learning
url https://arxiv.org/abs/2404.03426