Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Recurso digital |
| Language: | |
| Published: |
Zenodo
2025
|
| Online Access: | https://doi.org/10.5281/zenodo.15664496 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866902334911545344 |
|---|---|
| author | Dr. J. Jebamalar Tamilselvi Dr. K. Sutha Dr. S. Sweetlin Susilabai Mr. Raman Raguraman Dr. Surya Susan Thomas |
| author_facet | Dr. J. Jebamalar Tamilselvi Dr. K. Sutha Dr. S. Sweetlin Susilabai Mr. Raman Raguraman Dr. Surya Susan Thomas |
| contents | <p>Effective data management is essential in Big Data Analytics, where vast amounts of data are collected, processed, and analyzed to extract valuable insights. Duplicate data presents significant challenges, including increased storage costs, diminished analytical efficiency, and compromised data quality, all of which can lead to erroneous decision-making. Conventional deduplication techniques based on exact matching or simplistic heuristics often fall short in addressing data variability, semantic equivalence, and the massive scale typical of big data. This research proposes a token-based framework enhanced by a hybrid machine learning approach, integrating tokenbased strategies with both supervised and unsupervised learning techniques to accurately detect and eliminate duplicate data in Big Data Analytics environments.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_15664496 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | A Token-Based Framework for Detecting and Eliminating Duplicate Data in Big Data Analytics Using Machine Learning Techniques Dr. J. Jebamalar Tamilselvi Dr. K. Sutha Dr. S. Sweetlin Susilabai Mr. Raman Raguraman Dr. Surya Susan Thomas <p>Effective data management is essential in Big Data Analytics, where vast amounts of data are collected, processed, and analyzed to extract valuable insights. Duplicate data presents significant challenges, including increased storage costs, diminished analytical efficiency, and compromised data quality, all of which can lead to erroneous decision-making. Conventional deduplication techniques based on exact matching or simplistic heuristics often fall short in addressing data variability, semantic equivalence, and the massive scale typical of big data. This research proposes a token-based framework enhanced by a hybrid machine learning approach, integrating tokenbased strategies with both supervised and unsupervised learning techniques to accurately detect and eliminate duplicate data in Big Data Analytics environments.</p> |
| title | A Token-Based Framework for Detecting and Eliminating Duplicate Data in Big Data Analytics Using Machine Learning Techniques |
| url | https://doi.org/10.5281/zenodo.15664496 |