Assessing high-order effects in feature importance via predictability decomposition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ontivero-Ortega, Marlis, Faes, Luca, Cortes, Jesus M, Marinazzo, Daniele, Stramaglia, Sebastiano
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929756780363776
author Ontivero-Ortega, Marlis
Faes, Luca
Cortes, Jesus M
Marinazzo, Daniele
Stramaglia, Sebastiano
author_facet Ontivero-Ortega, Marlis
Faes, Luca
Cortes, Jesus M
Marinazzo, Daniele
Stramaglia, Sebastiano
contents Leveraging the large body of work devoted in recent years to describe redundancy and synergy in multivariate interactions among random variables, we propose a novel approach to quantify cooperative effects in feature importance, one of the most used techniques for explainable artificial intelligence. In particular, we propose an adaptive version of a well-known metric of feature importance, named Leave One Covariate Out (LOCO), to disentangle high-order effects involving a given input feature in regression problems. LOCO is the reduction of the prediction error when the feature under consideration is added to the set of all the features used for regression. Instead of calculating the LOCO using all the features at hand, as in its standard version, our method searches for the multiplet of features that maximize LOCO and for the one that minimize it. This provides a decomposition of the LOCO as the sum of a two-body component and higher-order components (redundant and synergistic), also highlighting the features that contribute to building these high-order effects alongside the driving feature. We report the application to proton/pion discrimination from simulated detector measures by GEANT.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09964
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Assessing high-order effects in feature importance via predictability decomposition
Ontivero-Ortega, Marlis
Faes, Luca
Cortes, Jesus M
Marinazzo, Daniele
Stramaglia, Sebastiano
Data Analysis, Statistics and Probability
Machine Learning
Leveraging the large body of work devoted in recent years to describe redundancy and synergy in multivariate interactions among random variables, we propose a novel approach to quantify cooperative effects in feature importance, one of the most used techniques for explainable artificial intelligence. In particular, we propose an adaptive version of a well-known metric of feature importance, named Leave One Covariate Out (LOCO), to disentangle high-order effects involving a given input feature in regression problems. LOCO is the reduction of the prediction error when the feature under consideration is added to the set of all the features used for regression. Instead of calculating the LOCO using all the features at hand, as in its standard version, our method searches for the multiplet of features that maximize LOCO and for the one that minimize it. This provides a decomposition of the LOCO as the sum of a two-body component and higher-order components (redundant and synergistic), also highlighting the features that contribute to building these high-order effects alongside the driving feature. We report the application to proton/pion discrimination from simulated detector measures by GEANT.
title Assessing high-order effects in feature importance via predictability decomposition
topic Data Analysis, Statistics and Probability
Machine Learning
url https://arxiv.org/abs/2412.09964