MMD-based Variable Importance for Distributional Random Forest

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bénard, Clément, Näf, Jeffrey, Josse, Julie
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909105680023552
author Bénard, Clément
Näf, Jeffrey
Josse, Julie
author_facet Bénard, Clément
Näf, Jeffrey
Josse, Julie
contents Distributional Random Forest (DRF) is a flexible forest-based method to estimate the full conditional distribution of a multivariate output of interest given input variables. In this article, we introduce a variable importance algorithm for DRFs, based on the well-established drop and relearn principle and MMD distance. While traditional importance measures only detect variables with an influence on the output mean, our algorithm detects variables impacting the output distribution more generally. We show that the introduced importance measure is consistent, exhibits high empirical performance on both real and simulated data, and outperforms competitors. In particular, our algorithm is highly efficient to select variables through recursive feature elimination, and can therefore provide small sets of variables to build accurate estimates of conditional output distributions.
format Preprint
id arxiv_https___arxiv_org_abs_2310_12115
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MMD-based Variable Importance for Distributional Random Forest
Bénard, Clément
Näf, Jeffrey
Josse, Julie
Machine Learning
Methodology
Distributional Random Forest (DRF) is a flexible forest-based method to estimate the full conditional distribution of a multivariate output of interest given input variables. In this article, we introduce a variable importance algorithm for DRFs, based on the well-established drop and relearn principle and MMD distance. While traditional importance measures only detect variables with an influence on the output mean, our algorithm detects variables impacting the output distribution more generally. We show that the introduced importance measure is consistent, exhibits high empirical performance on both real and simulated data, and outperforms competitors. In particular, our algorithm is highly efficient to select variables through recursive feature elimination, and can therefore provide small sets of variables to build accurate estimates of conditional output distributions.
title MMD-based Variable Importance for Distributional Random Forest
topic Machine Learning
Methodology
url https://arxiv.org/abs/2310.12115