A Deep Learning Approach to Language-independent Gender Prediction on Twitter

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hashempour, Reyhaneh, Plank, Barbara, Villavicencio, Aline, de Amorim, Renato Cordeiro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909408914571264
author Hashempour, Reyhaneh
Plank, Barbara
Villavicencio, Aline
de Amorim, Renato Cordeiro
author_facet Hashempour, Reyhaneh
Plank, Barbara
Villavicencio, Aline
de Amorim, Renato Cordeiro
contents This work presents a set of experiments conducted to predict the gender of Twitter users based on language-independent features extracted from the text of the users' tweets. The experiments were performed on a version of TwiSty dataset including tweets written by the users of six different languages: Portuguese, French, Dutch, English, German, and Italian. Logistic regression (LR), and feed-forward neural networks (FFNN) with back-propagation were used to build models in two different settings: Inter-Lingual (IL) and Cross-Lingual (CL). In the IL setting, the training and testing were performed on the same language whereas in the CL, Italian and German datasets were set aside and only used as test sets and the rest were combined to compose training and development sets. In the IL, the highest accuracy score belongs to LR whereas in the CL, FFNN with three hidden layers yields the highest score. The results show that neural network based models underperform traditional models when the size of the training set is small; however, they beat traditional models by a non-trivial margin, when they are fed with large enough data. Finally, the feature analysis confirms that men and women have different writing styles independent of their language.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19733
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Deep Learning Approach to Language-independent Gender Prediction on Twitter
Hashempour, Reyhaneh
Plank, Barbara
Villavicencio, Aline
de Amorim, Renato Cordeiro
Computation and Language
This work presents a set of experiments conducted to predict the gender of Twitter users based on language-independent features extracted from the text of the users' tweets. The experiments were performed on a version of TwiSty dataset including tweets written by the users of six different languages: Portuguese, French, Dutch, English, German, and Italian. Logistic regression (LR), and feed-forward neural networks (FFNN) with back-propagation were used to build models in two different settings: Inter-Lingual (IL) and Cross-Lingual (CL). In the IL setting, the training and testing were performed on the same language whereas in the CL, Italian and German datasets were set aside and only used as test sets and the rest were combined to compose training and development sets. In the IL, the highest accuracy score belongs to LR whereas in the CL, FFNN with three hidden layers yields the highest score. The results show that neural network based models underperform traditional models when the size of the training set is small; however, they beat traditional models by a non-trivial margin, when they are fed with large enough data. Finally, the feature analysis confirms that men and women have different writing styles independent of their language.
title A Deep Learning Approach to Language-independent Gender Prediction on Twitter
topic Computation and Language
url https://arxiv.org/abs/2411.19733