Handling Missing Data in Probabilistic Regression Trees: Methods and Implementation in R

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prass, Taiane Schaedler, Neimaier, Alisson Silva, Pumi, Guilherme
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916990017339392
author Prass, Taiane Schaedler
Neimaier, Alisson Silva
Pumi, Guilherme
author_facet Prass, Taiane Schaedler
Neimaier, Alisson Silva
Pumi, Guilherme
contents Probabilistic Regression Trees (PRTrees) generalize traditional decision trees by incorporating probability functions that associate each data point with different regions of the tree, providing smooth decisions and continuous responses. This paper introduces an adaptation of PRTrees capable of handling missing values in covariates through three distinct approaches: (i) a uniform probability method, (ii) a partial observation approach, and (iii) a dimension-reduced smoothing technique. The proposed methods preserve the interpretability properties of PRTrees while extending their applicability to incomplete datasets. Simulation studies under MCAR conditions demonstrate the relative performance of each approach, including comparisons with traditional regression trees on smooth function estimation tasks. The proposed methods, together with the original version, have been developed in R with highly optimized routines and are distributed in the PRTree package, publicly available on CRAN. In this paper we also present and discuss the main functionalities of the PRTree package, providing researchers and practitioners with new tools for incomplete data analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03634
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Handling Missing Data in Probabilistic Regression Trees: Methods and Implementation in R
Prass, Taiane Schaedler
Neimaier, Alisson Silva
Pumi, Guilherme
Methodology
Machine Learning
62G08, 62D10, 62H30, 62J02, 65C60
Probabilistic Regression Trees (PRTrees) generalize traditional decision trees by incorporating probability functions that associate each data point with different regions of the tree, providing smooth decisions and continuous responses. This paper introduces an adaptation of PRTrees capable of handling missing values in covariates through three distinct approaches: (i) a uniform probability method, (ii) a partial observation approach, and (iii) a dimension-reduced smoothing technique. The proposed methods preserve the interpretability properties of PRTrees while extending their applicability to incomplete datasets. Simulation studies under MCAR conditions demonstrate the relative performance of each approach, including comparisons with traditional regression trees on smooth function estimation tasks. The proposed methods, together with the original version, have been developed in R with highly optimized routines and are distributed in the PRTree package, publicly available on CRAN. In this paper we also present and discuss the main functionalities of the PRTree package, providing researchers and practitioners with new tools for incomplete data analysis.
title Handling Missing Data in Probabilistic Regression Trees: Methods and Implementation in R
topic Methodology
Machine Learning
62G08, 62D10, 62H30, 62J02, 65C60
url https://arxiv.org/abs/2510.03634