Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Eduardo, González Trigueros Jesús, Alejandro, Alonso Sánchez, Emilio, Muñoz Rivera, Jaqueline, Peñarán Prieto Mariana, Natalia, Mendoza González Camila
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908681285664768
author Eduardo, González Trigueros Jesús
Alejandro, Alonso Sánchez
Emilio, Muñoz Rivera
Jaqueline, Peñarán Prieto Mariana
Natalia, Mendoza González Camila
author_facet Eduardo, González Trigueros Jesús
Alejandro, Alonso Sánchez
Emilio, Muñoz Rivera
Jaqueline, Peñarán Prieto Mariana
Natalia, Mendoza González Camila
contents This study analyzes the impact of heterogeneity ("Variety") in Big Data by comparing classification strategies across structured (Epsilon) and unstructured (Rest-Mex, IMDB) domains. A dual methodology was implemented: evolutionary and Bayesian hyperparameter optimization (Genetic Algorithms, Optuna) in Python for numerical data, and distributed processing in Apache Spark for massive textual corpora. The results reveal a "complexity paradox": in high-dimensional spaces, optimized linear models (SVM, Logistic Regression) outperformed deep architectures and Gradient Boosting. Conversely, in text-based domains, the constraints of distributed fine-tuning led to overfitting in complex models, whereas robust feature engineering -- specifically Transformer-based embeddings (ROBERTa) and Bayesian Target Encoding -- enabled simpler models to generalize effectively. This work provides a unified framework for algorithm selection based on data nature and infrastructure constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains
Eduardo, González Trigueros Jesús
Alejandro, Alonso Sánchez
Emilio, Muñoz Rivera
Jaqueline, Peñarán Prieto Mariana
Natalia, Mendoza González Camila
Machine Learning
Computation and Language
Distributed, Parallel, and Cluster Computing
I.2.6; H.2.8; I.5.2
This study analyzes the impact of heterogeneity ("Variety") in Big Data by comparing classification strategies across structured (Epsilon) and unstructured (Rest-Mex, IMDB) domains. A dual methodology was implemented: evolutionary and Bayesian hyperparameter optimization (Genetic Algorithms, Optuna) in Python for numerical data, and distributed processing in Apache Spark for massive textual corpora. The results reveal a "complexity paradox": in high-dimensional spaces, optimized linear models (SVM, Logistic Regression) outperformed deep architectures and Gradient Boosting. Conversely, in text-based domains, the constraints of distributed fine-tuning led to overfitting in complex models, whereas robust feature engineering -- specifically Transformer-based embeddings (ROBERTa) and Bayesian Target Encoding -- enabled simpler models to generalize effectively. This work provides a unified framework for algorithm selection based on data nature and infrastructure constraints.
title Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains
topic Machine Learning
Computation and Language
Distributed, Parallel, and Cluster Computing
I.2.6; H.2.8; I.5.2
url https://arxiv.org/abs/2512.00298