Missing data imputation for noisy time-series data and applications in healthcare

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le, Lien P., Thi, Xuan-Hien Nguyen, Nguyen, Thu, Riegler, Michael A., Halvorsen, Pål, Nguyen, Binh T.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909429983608832
author Le, Lien P.
Thi, Xuan-Hien Nguyen
Nguyen, Thu
Riegler, Michael A.
Halvorsen, Pål
Nguyen, Binh T.
author_facet Le, Lien P.
Thi, Xuan-Hien Nguyen
Nguyen, Thu
Riegler, Michael A.
Halvorsen, Pål
Nguyen, Binh T.
contents Healthcare time series data is vital for monitoring patient activity but often contains noise and missing values due to various reasons such as sensor errors or data interruptions. Imputation, i.e., filling in the missing values, is a common way to deal with this issue. In this study, we compare imputation methods, including Multiple Imputation with Random Forest (MICE-RF) and advanced deep learning approaches (SAITS, BRITS, Transformer) for noisy, missing time series data in terms of MAE, F1-score, AUC, and MCC, across missing data rates (10 % - 80 %). Our results show that MICE-RF can effectively impute missing data compared to deep learning methods and the improvement in classification of data imputed indicates that imputation can have denoising effects. Therefore, using an imputation algorithm on time series with missing data can, at the same time, offer denoising effects.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11164
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Missing data imputation for noisy time-series data and applications in healthcare
Le, Lien P.
Thi, Xuan-Hien Nguyen
Nguyen, Thu
Riegler, Michael A.
Halvorsen, Pål
Nguyen, Binh T.
Machine Learning
Applications
Healthcare time series data is vital for monitoring patient activity but often contains noise and missing values due to various reasons such as sensor errors or data interruptions. Imputation, i.e., filling in the missing values, is a common way to deal with this issue. In this study, we compare imputation methods, including Multiple Imputation with Random Forest (MICE-RF) and advanced deep learning approaches (SAITS, BRITS, Transformer) for noisy, missing time series data in terms of MAE, F1-score, AUC, and MCC, across missing data rates (10 % - 80 %). Our results show that MICE-RF can effectively impute missing data compared to deep learning methods and the improvement in classification of data imputed indicates that imputation can have denoising effects. Therefore, using an imputation algorithm on time series with missing data can, at the same time, offer denoising effects.
title Missing data imputation for noisy time-series data and applications in healthcare
topic Machine Learning
Applications
url https://arxiv.org/abs/2412.11164