Blog Data Showdown: Machine Learning vs Neuro-Symbolic Models for Gender Classification

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sinshaw, Natnael Tilahun, He, Mengmei, Bahiru, Tadesse K., Mohapatra, Sudhir Kumar
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911403409932288
author Sinshaw, Natnael Tilahun
He, Mengmei
Bahiru, Tadesse K.
Mohapatra, Sudhir Kumar
author_facet Sinshaw, Natnael Tilahun
He, Mengmei
Bahiru, Tadesse K.
Mohapatra, Sudhir Kumar
contents Text classification problems, such as gender classification from a blog, have been a well-matured research area that has been well studied using machine learning algorithms. It has several application domains in market analysis, customer recommendation, and recommendation systems. This study presents a comparative analysis of the widely used machine learning algorithms, namely Support Vector Machines (SVM), Naive Bayes (NB), Logistic Regression (LR), AdaBoost, XGBoost, and an SVM variant (SVM_R) with neuro-symbolic AI (NeSy). The paper also explores the effect of text representations such as TF-IDF, the Universal Sentence Encoder (USE), and RoBERTa. Additionally, various feature extraction techniques, including Chi-Square, Mutual Information, and Principal Component Analysis, are explored. Building on these, we introduce a comparative analysis of the machine learning and deep learning approaches in comparison to the NeSy. The experimental results show that the use of the NeSy approach matched strong MLP results despite a limited dataset. Future work on this research will expand the knowledge base, the scope of embedding types, and the hyperparameter configuration to further study the effectiveness of the NeSy approach.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Blog Data Showdown: Machine Learning vs Neuro-Symbolic Models for Gender Classification
Sinshaw, Natnael Tilahun
He, Mengmei
Bahiru, Tadesse K.
Mohapatra, Sudhir Kumar
Machine Learning
Text classification problems, such as gender classification from a blog, have been a well-matured research area that has been well studied using machine learning algorithms. It has several application domains in market analysis, customer recommendation, and recommendation systems. This study presents a comparative analysis of the widely used machine learning algorithms, namely Support Vector Machines (SVM), Naive Bayes (NB), Logistic Regression (LR), AdaBoost, XGBoost, and an SVM variant (SVM_R) with neuro-symbolic AI (NeSy). The paper also explores the effect of text representations such as TF-IDF, the Universal Sentence Encoder (USE), and RoBERTa. Additionally, various feature extraction techniques, including Chi-Square, Mutual Information, and Principal Component Analysis, are explored. Building on these, we introduce a comparative analysis of the machine learning and deep learning approaches in comparison to the NeSy. The experimental results show that the use of the NeSy approach matched strong MLP results despite a limited dataset. Future work on this research will expand the knowledge base, the scope of embedding types, and the hyperparameter configuration to further study the effectiveness of the NeSy approach.
title Blog Data Showdown: Machine Learning vs Neuro-Symbolic Models for Gender Classification
topic Machine Learning
url https://arxiv.org/abs/2512.16687