Regularisation of CART trees by summation of $p$-values

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Engler, Nils, Lindholm, Mathias, Lindskog, Filip, Nazar, Taariq
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909872384114688
author Engler, Nils
Lindholm, Mathias
Lindskog, Filip
Nazar, Taariq
author_facet Engler, Nils
Lindholm, Mathias
Lindskog, Filip
Nazar, Taariq
contents The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the selected CART regression tree is not a deterministic function of the data. Moreover, the cross-validation procedure may become time consuming and result in inefficient use of training data. We propose a simple deterministic in-sample method that can be used for stopping the growing of a CART regression tree based on node-wise statistical tests. This testing procedure is derived using a connection to change point detection, where the null hypothesis corresponds to no signal. The suggested $p$-value based procedure allows us to consider covariate vectors of arbitrary dimension and allows us to bound the $p$-value of an entire tree from above. Further, we show that the test detects a not too weak signal with a high probability, given a not too small sample size. We illustrate our methodology and the asymptotic results on both simulated and real world data. Additionally, we illustrate how the $p$-value based method can be used to construct a deterministic piece-wise constant auto-calibrated predictor based on a given black-box predictor.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18769
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Regularisation of CART trees by summation of $p$-values
Engler, Nils
Lindholm, Mathias
Lindskog, Filip
Nazar, Taariq
Methodology
Statistics Theory
62F03, 62J99
The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the selected CART regression tree is not a deterministic function of the data. Moreover, the cross-validation procedure may become time consuming and result in inefficient use of training data. We propose a simple deterministic in-sample method that can be used for stopping the growing of a CART regression tree based on node-wise statistical tests. This testing procedure is derived using a connection to change point detection, where the null hypothesis corresponds to no signal. The suggested $p$-value based procedure allows us to consider covariate vectors of arbitrary dimension and allows us to bound the $p$-value of an entire tree from above. Further, we show that the test detects a not too weak signal with a high probability, given a not too small sample size. We illustrate our methodology and the asymptotic results on both simulated and real world data. Additionally, we illustrate how the $p$-value based method can be used to construct a deterministic piece-wise constant auto-calibrated predictor based on a given black-box predictor.
title Regularisation of CART trees by summation of $p$-values
topic Methodology
Statistics Theory
62F03, 62J99
url https://arxiv.org/abs/2505.18769