Saved in:
Bibliographic Details
Main Authors: Dr. J. Jebamalar Tamilselvi, Dr. K. Sutha, Dr. S. Sweetlin Susilabai, Mr. Raman Raguraman, Dr. Surya Susan Thomas
Format: Recurso digital
Language:
Published: Zenodo 2025
Online Access:https://doi.org/10.5281/zenodo.15664496
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866902334911545344
author Dr. J. Jebamalar Tamilselvi
Dr. K. Sutha
Dr. S. Sweetlin Susilabai
Mr. Raman Raguraman
Dr. Surya Susan Thomas
author_facet Dr. J. Jebamalar Tamilselvi
Dr. K. Sutha
Dr. S. Sweetlin Susilabai
Mr. Raman Raguraman
Dr. Surya Susan Thomas
contents <p>Effective data management is essential in Big Data Analytics, where vast amounts of data are collected, processed, and analyzed to extract valuable insights. Duplicate data presents significant challenges, including increased storage costs, diminished analytical efficiency, and compromised data quality, all of which can lead to erroneous decision-making. Conventional deduplication techniques based on exact matching or simplistic heuristics often fall short in addressing data variability, semantic equivalence, and the massive scale typical of big data. This research proposes a token-based framework enhanced by a hybrid machine learning approach, integrating tokenbased strategies with both supervised and unsupervised learning techniques to accurately detect and eliminate duplicate data in Big Data Analytics environments.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_15664496
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle A Token-Based Framework for Detecting and Eliminating Duplicate Data in Big Data Analytics Using Machine Learning Techniques
Dr. J. Jebamalar Tamilselvi
Dr. K. Sutha
Dr. S. Sweetlin Susilabai
Mr. Raman Raguraman
Dr. Surya Susan Thomas
<p>Effective data management is essential in Big Data Analytics, where vast amounts of data are collected, processed, and analyzed to extract valuable insights. Duplicate data presents significant challenges, including increased storage costs, diminished analytical efficiency, and compromised data quality, all of which can lead to erroneous decision-making. Conventional deduplication techniques based on exact matching or simplistic heuristics often fall short in addressing data variability, semantic equivalence, and the massive scale typical of big data. This research proposes a token-based framework enhanced by a hybrid machine learning approach, integrating tokenbased strategies with both supervised and unsupervised learning techniques to accurately detect and eliminate duplicate data in Big Data Analytics environments.</p>
title A Token-Based Framework for Detecting and Eliminating Duplicate Data in Big Data Analytics Using Machine Learning Techniques
url https://doi.org/10.5281/zenodo.15664496