Autoencoding Random Forests

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vu, Binh Duc, Kapar, Jan, Wright, Marvin, Watson, David S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918289514430464
author Vu, Binh Duc
Kapar, Jan
Wright, Marvin
Watson, David S.
author_facet Vu, Binh Duc
Kapar, Jan
Wright, Marvin
Watson, David S.
contents We propose a principled method for autoencoding with random forests. Our strategy builds on foundational results from nonparametric statistics and spectral graph theory to learn a low-dimensional embedding of the model that optimally represents relationships in the data. We provide exact and approximate solutions to the decoding problem via constrained optimization, split relabeling, and nearest neighbors regression. These methods effectively invert the compression pipeline, establishing a map from the embedding space back to the input space using splits learned by the ensemble's constituent trees. The resulting decoders are universally consistent under common regularity assumptions. The procedure works with supervised or unsupervised models, providing a window into conditional or joint distributions. We demonstrate various applications of this autoencoder, including powerful new tools for visualization, compression, clustering, and denoising. Experiments illustrate the ease and utility of our method in a wide range of settings, including tabular, image, and genomic data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21441
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Autoencoding Random Forests
Vu, Binh Duc
Kapar, Jan
Wright, Marvin
Watson, David S.
Machine Learning
Artificial Intelligence
We propose a principled method for autoencoding with random forests. Our strategy builds on foundational results from nonparametric statistics and spectral graph theory to learn a low-dimensional embedding of the model that optimally represents relationships in the data. We provide exact and approximate solutions to the decoding problem via constrained optimization, split relabeling, and nearest neighbors regression. These methods effectively invert the compression pipeline, establishing a map from the embedding space back to the input space using splits learned by the ensemble's constituent trees. The resulting decoders are universally consistent under common regularity assumptions. The procedure works with supervised or unsupervised models, providing a window into conditional or joint distributions. We demonstrate various applications of this autoencoder, including powerful new tools for visualization, compression, clustering, and denoising. Experiments illustrate the ease and utility of our method in a wide range of settings, including tabular, image, and genomic data.
title Autoencoding Random Forests
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.21441