Machine Learning Based Missing Values Imputation in Categorical Datasets

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ishaq, Muhammad, Zahir, Sana, Iftikhar, Laila, Bulbul, Mohammad Farhad, Rho, Seungmin, Lee, Mi Young
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912023310237696
author Ishaq, Muhammad
Zahir, Sana
Iftikhar, Laila
Bulbul, Mohammad Farhad
Rho, Seungmin
Lee, Mi Young
author_facet Ishaq, Muhammad
Zahir, Sana
Iftikhar, Laila
Bulbul, Mohammad Farhad
Rho, Seungmin
Lee, Mi Young
contents In order to predict and fill in the gaps in categorical datasets, this research looked into the use of machine learning algorithms. The emphasis was on ensemble models constructed using the Error Correction Output Codes framework, including models based on SVM and KNN as well as a hybrid classifier that combines models based on SVM, KNN,and MLP. Three diverse datasets, the CPU, Hypothyroid, and Breast Cancer datasets were employed to validate these algorithms. Results indicated that these machine learning techniques provided substantial performance in predicting and completing missing data, with the effectiveness varying based on the specific dataset and missing data pattern. Compared to solo models, ensemble models that made use of the ECOC framework significantly improved prediction accuracy and robustness. Deep learning for missing data imputation has obstacles despite these encouraging results, including the requirement for large amounts of labeled data and the possibility of overfitting. Subsequent research endeavors ought to evaluate the feasibility and efficacy of deep learning algorithms in the context of the imputation of missing data.
format Preprint
id arxiv_https___arxiv_org_abs_2306_06338
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Machine Learning Based Missing Values Imputation in Categorical Datasets
Ishaq, Muhammad
Zahir, Sana
Iftikhar, Laila
Bulbul, Mohammad Farhad
Rho, Seungmin
Lee, Mi Young
Machine Learning
In order to predict and fill in the gaps in categorical datasets, this research looked into the use of machine learning algorithms. The emphasis was on ensemble models constructed using the Error Correction Output Codes framework, including models based on SVM and KNN as well as a hybrid classifier that combines models based on SVM, KNN,and MLP. Three diverse datasets, the CPU, Hypothyroid, and Breast Cancer datasets were employed to validate these algorithms. Results indicated that these machine learning techniques provided substantial performance in predicting and completing missing data, with the effectiveness varying based on the specific dataset and missing data pattern. Compared to solo models, ensemble models that made use of the ECOC framework significantly improved prediction accuracy and robustness. Deep learning for missing data imputation has obstacles despite these encouraging results, including the requirement for large amounts of labeled data and the possibility of overfitting. Subsequent research endeavors ought to evaluate the feasibility and efficacy of deep learning algorithms in the context of the imputation of missing data.
title Machine Learning Based Missing Values Imputation in Categorical Datasets
topic Machine Learning
url https://arxiv.org/abs/2306.06338