Fast Bayesian Basis Selection for Functional Data Representation with Correlated Errors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: da Cruz, Ana Carolina, de Souza, Camila P. E., Sousa, Pedro H. T. O.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915580222636032
author da Cruz, Ana Carolina
de Souza, Camila P. E.
Sousa, Pedro H. T. O.
author_facet da Cruz, Ana Carolina
de Souza, Camila P. E.
Sousa, Pedro H. T. O.
contents Functional data analysis finds widespread application across various fields. While functional data are intrinsically infinite-dimensional, in practice, they are observed only at a finite set of points, typically over a dense grid. As a result, smoothing techniques are often used to approximate the observed data as functions. In this work, we propose a novel Bayesian approach for selecting basis functions for smoothing one or multiple curves simultaneously. Our method differentiates from other Bayesian approaches in two key ways: (i) by accounting for correlated errors and (ii) by developing a variational Expectation-Maximization (VEM) algorithm, which is faster than Markov chain Monte Carlo (MCMC) methods such as Gibbs sampling. Simulation studies demonstrate that our method effectively identifies the true underlying structure of the data across various scenarios, and it is applicable to different types of functional data. Our VEM algorithm not only recovers the basis coefficients and the correct set of basis functions but also estimates the existing within-curve correlation. When applied to the motorcycle, LIDAR (LIght Detection And Ranging) experiment and Canadian weather datasets, our method demonstrates comparable, and in some cases superior, performance in terms of adjusted R2 compared to regression splines, smoothing splines, least absolute shrinkage and selection operator (LASSO) and Bayesian LASSO. Our proposed method is implemented in R and codes are available at https://github.com/acarolcruz/VB-Bases-Selection
format Preprint
id arxiv_https___arxiv_org_abs_2405_20758
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fast Bayesian Basis Selection for Functional Data Representation with Correlated Errors
da Cruz, Ana Carolina
de Souza, Camila P. E.
Sousa, Pedro H. T. O.
Methodology
Functional data analysis finds widespread application across various fields. While functional data are intrinsically infinite-dimensional, in practice, they are observed only at a finite set of points, typically over a dense grid. As a result, smoothing techniques are often used to approximate the observed data as functions. In this work, we propose a novel Bayesian approach for selecting basis functions for smoothing one or multiple curves simultaneously. Our method differentiates from other Bayesian approaches in two key ways: (i) by accounting for correlated errors and (ii) by developing a variational Expectation-Maximization (VEM) algorithm, which is faster than Markov chain Monte Carlo (MCMC) methods such as Gibbs sampling. Simulation studies demonstrate that our method effectively identifies the true underlying structure of the data across various scenarios, and it is applicable to different types of functional data. Our VEM algorithm not only recovers the basis coefficients and the correct set of basis functions but also estimates the existing within-curve correlation. When applied to the motorcycle, LIDAR (LIght Detection And Ranging) experiment and Canadian weather datasets, our method demonstrates comparable, and in some cases superior, performance in terms of adjusted R2 compared to regression splines, smoothing splines, least absolute shrinkage and selection operator (LASSO) and Bayesian LASSO. Our proposed method is implemented in R and codes are available at https://github.com/acarolcruz/VB-Bases-Selection
title Fast Bayesian Basis Selection for Functional Data Representation with Correlated Errors
topic Methodology
url https://arxiv.org/abs/2405.20758