Handling Missing Data in Probabilistic Regression Trees: Methods and Implementation in R
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916990017339392 |
|---|---|
| author | Prass, Taiane Schaedler Neimaier, Alisson Silva Pumi, Guilherme |
| author_facet | Prass, Taiane Schaedler Neimaier, Alisson Silva Pumi, Guilherme |
| contents | Probabilistic Regression Trees (PRTrees) generalize traditional decision trees by incorporating probability functions that associate each data point with different regions of the tree, providing smooth decisions and continuous responses. This paper introduces an adaptation of PRTrees capable of handling missing values in covariates through three distinct approaches: (i) a uniform probability method, (ii) a partial observation approach, and (iii) a dimension-reduced smoothing technique. The proposed methods preserve the interpretability properties of PRTrees while extending their applicability to incomplete datasets. Simulation studies under MCAR conditions demonstrate the relative performance of each approach, including comparisons with traditional regression trees on smooth function estimation tasks. The proposed methods, together with the original version, have been developed in R with highly optimized routines and are distributed in the PRTree package, publicly available on CRAN. In this paper we also present and discuss the main functionalities of the PRTree package, providing researchers and practitioners with new tools for incomplete data analysis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_03634 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Handling Missing Data in Probabilistic Regression Trees: Methods and Implementation in R Prass, Taiane Schaedler Neimaier, Alisson Silva Pumi, Guilherme Methodology Machine Learning 62G08, 62D10, 62H30, 62J02, 65C60 Probabilistic Regression Trees (PRTrees) generalize traditional decision trees by incorporating probability functions that associate each data point with different regions of the tree, providing smooth decisions and continuous responses. This paper introduces an adaptation of PRTrees capable of handling missing values in covariates through three distinct approaches: (i) a uniform probability method, (ii) a partial observation approach, and (iii) a dimension-reduced smoothing technique. The proposed methods preserve the interpretability properties of PRTrees while extending their applicability to incomplete datasets. Simulation studies under MCAR conditions demonstrate the relative performance of each approach, including comparisons with traditional regression trees on smooth function estimation tasks. The proposed methods, together with the original version, have been developed in R with highly optimized routines and are distributed in the PRTree package, publicly available on CRAN. In this paper we also present and discuss the main functionalities of the PRTree package, providing researchers and practitioners with new tools for incomplete data analysis. |
| title | Handling Missing Data in Probabilistic Regression Trees: Methods and Implementation in R |
| topic | Methodology Machine Learning 62G08, 62D10, 62H30, 62J02, 65C60 |
| url | https://arxiv.org/abs/2510.03634 |