RETENTION: Resource-Efficient Tree-Based Ensemble Model Acceleration with Content-Addressable Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Yi-Chun, Tsai, Chieh-Lin, Chang, Yuan-Hao, Slimani, Camélia, Boukhobza, Jalil, Kuo, Tei-Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917547891228672
author Liao, Yi-Chun
Tsai, Chieh-Lin
Chang, Yuan-Hao
Slimani, Camélia
Boukhobza, Jalil
Kuo, Tei-Wei
author_facet Liao, Yi-Chun
Tsai, Chieh-Lin
Chang, Yuan-Hao
Slimani, Camélia
Boukhobza, Jalil
Kuo, Tei-Wei
contents Although deep learning has demonstrated remarkable capability in learning from unstructured data, modern tree-based ensemble models remain superior in extracting relevant information and learning from structured datasets. While several efforts have been made to accelerate tree-based models, the inherent characteristics of the models pose significant challenges for conventional accelerators. Recent research leveraging content-addressable memory (CAM) offers a promising solution for accelerating tree-based models, yet existing designs suffer from excessive memory consumption and low utilization. This work addresses these challenges by introducing RETENTION, an end-to-end framework that significantly reduces CAM capacity requirement for tree-based model inference. We propose an iterative pruning algorithm with a novel pruning criterion tailored for bagging-based models (e.g., Random Forest), which minimizes model complexity while ensuring controlled accuracy degradation. Additionally, we present a tree mapping scheme that incorporates two innovative data placement strategies to alleviate the memory redundancy caused by the widespread use of don't care states in CAM. Experimental results show that implementing the tree mapping scheme alone reduces CAM capacity requirement by $1.46\times$ to $21.30 \times$, while the full RETENTION framework achieves $4.35\times$ to $207.12\times$ reduction with less than 3\% accuracy loss. These results demonstrate that RETENTION is highly effective in minimizing CAM resource demand, providing a resource-efficient direction for tree-based model acceleration.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05994
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RETENTION: Resource-Efficient Tree-Based Ensemble Model Acceleration with Content-Addressable Memory
Liao, Yi-Chun
Tsai, Chieh-Lin
Chang, Yuan-Hao
Slimani, Camélia
Boukhobza, Jalil
Kuo, Tei-Wei
Machine Learning
Hardware Architecture
Emerging Technologies
Although deep learning has demonstrated remarkable capability in learning from unstructured data, modern tree-based ensemble models remain superior in extracting relevant information and learning from structured datasets. While several efforts have been made to accelerate tree-based models, the inherent characteristics of the models pose significant challenges for conventional accelerators. Recent research leveraging content-addressable memory (CAM) offers a promising solution for accelerating tree-based models, yet existing designs suffer from excessive memory consumption and low utilization. This work addresses these challenges by introducing RETENTION, an end-to-end framework that significantly reduces CAM capacity requirement for tree-based model inference. We propose an iterative pruning algorithm with a novel pruning criterion tailored for bagging-based models (e.g., Random Forest), which minimizes model complexity while ensuring controlled accuracy degradation. Additionally, we present a tree mapping scheme that incorporates two innovative data placement strategies to alleviate the memory redundancy caused by the widespread use of don't care states in CAM. Experimental results show that implementing the tree mapping scheme alone reduces CAM capacity requirement by $1.46\times$ to $21.30 \times$, while the full RETENTION framework achieves $4.35\times$ to $207.12\times$ reduction with less than 3\% accuracy loss. These results demonstrate that RETENTION is highly effective in minimizing CAM resource demand, providing a resource-efficient direction for tree-based model acceleration.
title RETENTION: Resource-Efficient Tree-Based Ensemble Model Acceleration with Content-Addressable Memory
topic Machine Learning
Hardware Architecture
Emerging Technologies
url https://arxiv.org/abs/2506.05994