Provable Recovery of Locally Important Signed Features and Interactions from Random Forest

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vuk, Kata, Ihlo, Nicolas Alexander, Behr, Merle
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912759673782272
author Vuk, Kata
Ihlo, Nicolas Alexander
Behr, Merle
author_facet Vuk, Kata
Ihlo, Nicolas Alexander
Behr, Merle
contents Feature and Interaction Importance (FII) methods are essential in supervised learning for assessing the relevance of input variables and their interactions in complex prediction models. In many domains, such as personalized medicine, local interpretations for individual predictions are often required, rather than global scores summarizing overall feature importance. Random Forests (RFs) are widely used in these settings, and existing interpretability methods typically exploit tree structures and split statistics to provide model-specific insights. However, theoretical understanding of local FII methods for RF remains limited, making it unclear how to interpret high importance scores for individual predictions. We propose a novel, local, model-specific FII method that identifies frequent co-occurrences of features along decision paths, combining global patterns with those observed on paths specific to a given test point. We prove that our method consistently recovers the true local signal features and their interactions under a Locally Spike Sparse (LSS) model and also identifies whether large or small feature values drive a prediction. We illustrate the usefulness of our method and theoretical results through simulation studies and a real-world data example.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Provable Recovery of Locally Important Signed Features and Interactions from Random Forest
Vuk, Kata
Ihlo, Nicolas Alexander
Behr, Merle
Machine Learning
Methodology
Feature and Interaction Importance (FII) methods are essential in supervised learning for assessing the relevance of input variables and their interactions in complex prediction models. In many domains, such as personalized medicine, local interpretations for individual predictions are often required, rather than global scores summarizing overall feature importance. Random Forests (RFs) are widely used in these settings, and existing interpretability methods typically exploit tree structures and split statistics to provide model-specific insights. However, theoretical understanding of local FII methods for RF remains limited, making it unclear how to interpret high importance scores for individual predictions. We propose a novel, local, model-specific FII method that identifies frequent co-occurrences of features along decision paths, combining global patterns with those observed on paths specific to a given test point. We prove that our method consistently recovers the true local signal features and their interactions under a Locally Spike Sparse (LSS) model and also identifies whether large or small feature values drive a prediction. We illustrate the usefulness of our method and theoretical results through simulation studies and a real-world data example.
title Provable Recovery of Locally Important Signed Features and Interactions from Random Forest
topic Machine Learning
Methodology
url https://arxiv.org/abs/2512.11081