Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908681285664768 |
|---|---|
| author | Eduardo, González Trigueros Jesús Alejandro, Alonso Sánchez Emilio, Muñoz Rivera Jaqueline, Peñarán Prieto Mariana Natalia, Mendoza González Camila |
| author_facet | Eduardo, González Trigueros Jesús Alejandro, Alonso Sánchez Emilio, Muñoz Rivera Jaqueline, Peñarán Prieto Mariana Natalia, Mendoza González Camila |
| contents | This study analyzes the impact of heterogeneity ("Variety") in Big Data by comparing classification strategies across structured (Epsilon) and unstructured (Rest-Mex, IMDB) domains. A dual methodology was implemented: evolutionary and Bayesian hyperparameter optimization (Genetic Algorithms, Optuna) in Python for numerical data, and distributed processing in Apache Spark for massive textual corpora. The results reveal a "complexity paradox": in high-dimensional spaces, optimized linear models (SVM, Logistic Regression) outperformed deep architectures and Gradient Boosting. Conversely, in text-based domains, the constraints of distributed fine-tuning led to overfitting in complex models, whereas robust feature engineering -- specifically Transformer-based embeddings (ROBERTa) and Bayesian Target Encoding -- enabled simpler models to generalize effectively. This work provides a unified framework for algorithm selection based on data nature and infrastructure constraints. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_00298 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains Eduardo, González Trigueros Jesús Alejandro, Alonso Sánchez Emilio, Muñoz Rivera Jaqueline, Peñarán Prieto Mariana Natalia, Mendoza González Camila Machine Learning Computation and Language Distributed, Parallel, and Cluster Computing I.2.6; H.2.8; I.5.2 This study analyzes the impact of heterogeneity ("Variety") in Big Data by comparing classification strategies across structured (Epsilon) and unstructured (Rest-Mex, IMDB) domains. A dual methodology was implemented: evolutionary and Bayesian hyperparameter optimization (Genetic Algorithms, Optuna) in Python for numerical data, and distributed processing in Apache Spark for massive textual corpora. The results reveal a "complexity paradox": in high-dimensional spaces, optimized linear models (SVM, Logistic Regression) outperformed deep architectures and Gradient Boosting. Conversely, in text-based domains, the constraints of distributed fine-tuning led to overfitting in complex models, whereas robust feature engineering -- specifically Transformer-based embeddings (ROBERTa) and Bayesian Target Encoding -- enabled simpler models to generalize effectively. This work provides a unified framework for algorithm selection based on data nature and infrastructure constraints. |
| title | Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains |
| topic | Machine Learning Computation and Language Distributed, Parallel, and Cluster Computing I.2.6; H.2.8; I.5.2 |
| url | https://arxiv.org/abs/2512.00298 |