DeepIFSAC: Deep Imputation of Missing Values Using Feature and Sample Attention within Contrastive Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kowsar, Ibna, Rabbani, Shourav B., Hou, Yina, Samad, Manar D.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917967628861440
author Kowsar, Ibna
Rabbani, Shourav B.
Hou, Yina
Samad, Manar D.
author_facet Kowsar, Ibna
Rabbani, Shourav B.
Hou, Yina
Samad, Manar D.
contents Missing values of varying patterns and rates in real-world tabular data pose a significant challenge in developing reliable data-driven models. The most commonly used statistical and machine learning methods for missing value imputation may be ineffective when the missing rate is high and not random. This paper explores row and column attention in tabular data as between-feature and between-sample attention in a novel framework to reconstruct missing values. The proposed method uses CutMix data augmentation within a contrastive learning framework to improve the uncertainty of missing value estimation. The performance and generalizability of trained imputation models are evaluated in set-aside test data folds with missing values. The proposed framework is compared with 11 state-of-the-art statistical, machine learning, and deep imputation methods using 12 diverse tabular data sets. The average performance rank of our proposed method demonstrates its superiority over the state-of-the-art methods for missing rates between 10% and 90% and three missing value types, especially when the missing values are not random. The quality of the imputed data using our proposed method is compared in a downstream patient classification task using real-world electronic health records. This paper highlights the heterogeneity of tabular data sets to recommend imputation methods based on missing value types and data characteristics.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10910
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepIFSAC: Deep Imputation of Missing Values Using Feature and Sample Attention within Contrastive Framework
Kowsar, Ibna
Rabbani, Shourav B.
Hou, Yina
Samad, Manar D.
Machine Learning
Missing values of varying patterns and rates in real-world tabular data pose a significant challenge in developing reliable data-driven models. The most commonly used statistical and machine learning methods for missing value imputation may be ineffective when the missing rate is high and not random. This paper explores row and column attention in tabular data as between-feature and between-sample attention in a novel framework to reconstruct missing values. The proposed method uses CutMix data augmentation within a contrastive learning framework to improve the uncertainty of missing value estimation. The performance and generalizability of trained imputation models are evaluated in set-aside test data folds with missing values. The proposed framework is compared with 11 state-of-the-art statistical, machine learning, and deep imputation methods using 12 diverse tabular data sets. The average performance rank of our proposed method demonstrates its superiority over the state-of-the-art methods for missing rates between 10% and 90% and three missing value types, especially when the missing values are not random. The quality of the imputed data using our proposed method is compared in a downstream patient classification task using real-world electronic health records. This paper highlights the heterogeneity of tabular data sets to recommend imputation methods based on missing value types and data characteristics.
title DeepIFSAC: Deep Imputation of Missing Values Using Feature and Sample Attention within Contrastive Framework
topic Machine Learning
url https://arxiv.org/abs/2501.10910