Distributed Record Linkage in Healthcare Data with Apache Spark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Heydari, Mohammad, Sarshar, Reza, Soltanshahi, Mohammad Ali
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911836404711424
author Heydari, Mohammad
Sarshar, Reza
Soltanshahi, Mohammad Ali
author_facet Heydari, Mohammad
Sarshar, Reza
Soltanshahi, Mohammad Ali
contents Healthcare data is a valuable resource for research, analysis, and decision-making in the medical field. However, healthcare data is often fragmented and distributed across various sources, making it challenging to combine and analyze effectively. Record linkage, also known as data matching, is a crucial step in integrating and cleaning healthcare data to ensure data quality and accuracy. Apache Spark, a powerful open-source distributed big data processing framework, provides a robust platform for performing record linkage tasks with the aid of its machine learning library. In this study, we developed a new distributed data-matching model based on the Apache Spark Machine Learning library. To ensure the correct functioning of our model, the validation phase has been performed on the training data. The main challenge is data imbalance because a large amount of data is labeled false, and a small number of records are labeled true. By utilizing SVM and Regression algorithms, our results demonstrate that research data was neither over-fitted nor under-fitted, and this shows that our distributed model works well on the data.
format Preprint
id arxiv_https___arxiv_org_abs_2404_07939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Distributed Record Linkage in Healthcare Data with Apache Spark
Heydari, Mohammad
Sarshar, Reza
Soltanshahi, Mohammad Ali
Distributed, Parallel, and Cluster Computing
Machine Learning
Healthcare data is a valuable resource for research, analysis, and decision-making in the medical field. However, healthcare data is often fragmented and distributed across various sources, making it challenging to combine and analyze effectively. Record linkage, also known as data matching, is a crucial step in integrating and cleaning healthcare data to ensure data quality and accuracy. Apache Spark, a powerful open-source distributed big data processing framework, provides a robust platform for performing record linkage tasks with the aid of its machine learning library. In this study, we developed a new distributed data-matching model based on the Apache Spark Machine Learning library. To ensure the correct functioning of our model, the validation phase has been performed on the training data. The main challenge is data imbalance because a large amount of data is labeled false, and a small number of records are labeled true. By utilizing SVM and Regression algorithms, our results demonstrate that research data was neither over-fitted nor under-fitted, and this shows that our distributed model works well on the data.
title Distributed Record Linkage in Healthcare Data with Apache Spark
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2404.07939