Vectorized Adaptive Histograms for Sparse Oblique Forests

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lubonja, Ariel, Yoon, Jungsang, Xu, Haoyin, Wan, Yue, Xu, Yilin, Stotz, Richard, Guillame-Bert, Mathieu, Vogelstein, Joshua T., Burns, Randal
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912932105814016
author Lubonja, Ariel
Yoon, Jungsang
Xu, Haoyin
Wan, Yue
Xu, Yilin
Stotz, Richard
Guillame-Bert, Mathieu
Vogelstein, Joshua T.
Burns, Randal
author_facet Lubonja, Ariel
Yoon, Jungsang
Xu, Haoyin
Wan, Yue
Xu, Yilin
Stotz, Richard
Guillame-Bert, Mathieu
Vogelstein, Joshua T.
Burns, Randal
contents Classification using sparse oblique random forests provides guarantees on uncertainty and confidence while controlling for specific error types. However, they use more data and more compute than other tree ensembles because they create deep trees and need to sort or histogram linear combinations of data at runtime. We provide a method for dynamically switching between histograms and sorting to find the best split. We further optimize histogram construction using vector intrinsics. Evaluating this on large datasets, our optimizations speedup training by 1.7-2.5x compared to existing oblique forests and 1.5-2x compared to standard random forests. We also provide a GPU and hybrid CPU-GPU implementation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_00326
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vectorized Adaptive Histograms for Sparse Oblique Forests
Lubonja, Ariel
Yoon, Jungsang
Xu, Haoyin
Wan, Yue
Xu, Yilin
Stotz, Richard
Guillame-Bert, Mathieu
Vogelstein, Joshua T.
Burns, Randal
Machine Learning
Distributed, Parallel, and Cluster Computing
Performance
I.2.6; D.1.3
Classification using sparse oblique random forests provides guarantees on uncertainty and confidence while controlling for specific error types. However, they use more data and more compute than other tree ensembles because they create deep trees and need to sort or histogram linear combinations of data at runtime. We provide a method for dynamically switching between histograms and sorting to find the best split. We further optimize histogram construction using vector intrinsics. Evaluating this on large datasets, our optimizations speedup training by 1.7-2.5x compared to existing oblique forests and 1.5-2x compared to standard random forests. We also provide a GPU and hybrid CPU-GPU implementation.
title Vectorized Adaptive Histograms for Sparse Oblique Forests
topic Machine Learning
Distributed, Parallel, and Cluster Computing
Performance
I.2.6; D.1.3
url https://arxiv.org/abs/2603.00326