Leveraging Predictive Equivalence in Decision Trees

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: McTavish, Hayden, Boner, Zachery, Donnelly, Jon, Seltzer, Margo, Rudin, Cynthia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917010378588160
author McTavish, Hayden
Boner, Zachery
Donnelly, Jon
Seltzer, Margo
Rudin, Cynthia
author_facet McTavish, Hayden
Boner, Zachery
Donnelly, Jon
Seltzer, Margo
Rudin, Cynthia
contents Decision trees are widely used for interpretable machine learning due to their clearly structured reasoning process. However, this structure belies a challenge we refer to as predictive equivalence: a given tree's decision boundary can be represented by many different decision trees. The presence of models with identical decision boundaries but different evaluation processes makes model selection challenging. The models will have different variable importance and behave differently in the presence of missing values, but most optimization procedures will arbitrarily choose one such model to return. We present a boolean logical representation of decision trees that does not exhibit predictive equivalence and is faithful to the underlying decision boundary. We apply our representation to several downstream machine learning tasks. Using our representation, we show that decision trees are surprisingly robust to test-time missingness of feature values; we address predictive equivalence's impact on quantifying variable importance; and we present an algorithm to optimize the cost of reaching predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Predictive Equivalence in Decision Trees
McTavish, Hayden
Boner, Zachery
Donnelly, Jon
Seltzer, Margo
Rudin, Cynthia
Machine Learning
Decision trees are widely used for interpretable machine learning due to their clearly structured reasoning process. However, this structure belies a challenge we refer to as predictive equivalence: a given tree's decision boundary can be represented by many different decision trees. The presence of models with identical decision boundaries but different evaluation processes makes model selection challenging. The models will have different variable importance and behave differently in the presence of missing values, but most optimization procedures will arbitrarily choose one such model to return. We present a boolean logical representation of decision trees that does not exhibit predictive equivalence and is faithful to the underlying decision boundary. We apply our representation to several downstream machine learning tasks. Using our representation, we show that decision trees are surprisingly robust to test-time missingness of feature values; we address predictive equivalence's impact on quantifying variable importance; and we present an algorithm to optimize the cost of reaching predictions.
title Leveraging Predictive Equivalence in Decision Trees
topic Machine Learning
url https://arxiv.org/abs/2506.14143