Ultra-marginal Feature Importance: Learning from Data with Causal Guarantees

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Janssen, Joseph, Guan, Vincent, Robeva, Elina
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916474390577152
author Janssen, Joseph
Guan, Vincent
Robeva, Elina
author_facet Janssen, Joseph
Guan, Vincent
Robeva, Elina
contents Scientists frequently prioritize learning from data rather than training the best possible model; however, research in machine learning often prioritizes the latter. Marginal contribution feature importance (MCI) was developed to break this trend by providing a useful framework for quantifying the relationships in data. In this work, we aim to improve upon the theoretical properties, performance, and runtime of MCI by introducing ultra-marginal feature importance (UMFI), which uses dependence removal techniques from the AI fairness literature as its foundation. We first propose axioms for feature importance methods that seek to explain the causal and associative relationships in data, and we prove that UMFI satisfies these axioms under basic assumptions. We then show on real and simulated data that UMFI performs better than MCI, especially in the presence of correlated interactions and unrelated features, while partially learning the structure of the causal graph and reducing the exponential runtime of MCI to super-linear.
format Preprint
id arxiv_https___arxiv_org_abs_2204_09938
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Ultra-marginal Feature Importance: Learning from Data with Causal Guarantees
Janssen, Joseph
Guan, Vincent
Robeva, Elina
Machine Learning
Information Theory
Applications
Scientists frequently prioritize learning from data rather than training the best possible model; however, research in machine learning often prioritizes the latter. Marginal contribution feature importance (MCI) was developed to break this trend by providing a useful framework for quantifying the relationships in data. In this work, we aim to improve upon the theoretical properties, performance, and runtime of MCI by introducing ultra-marginal feature importance (UMFI), which uses dependence removal techniques from the AI fairness literature as its foundation. We first propose axioms for feature importance methods that seek to explain the causal and associative relationships in data, and we prove that UMFI satisfies these axioms under basic assumptions. We then show on real and simulated data that UMFI performs better than MCI, especially in the presence of correlated interactions and unrelated features, while partially learning the structure of the causal graph and reducing the exponential runtime of MCI to super-linear.
title Ultra-marginal Feature Importance: Learning from Data with Causal Guarantees
topic Machine Learning
Information Theory
Applications
url https://arxiv.org/abs/2204.09938