Dataset Creation and Baseline Models for Sexism Detection in Hausa

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Muhammad, Fatima Adam, Hassan, Shamsuddeen Muhammad, Inuwa-Dutse, Isa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912679134756864
author Muhammad, Fatima Adam
Hassan, Shamsuddeen Muhammad
Inuwa-Dutse, Isa
author_facet Muhammad, Fatima Adam
Hassan, Shamsuddeen Muhammad
Inuwa-Dutse, Isa
contents Sexism reinforces gender inequality and social exclusion by perpetuating stereotypes, bias, and discriminatory norms. Noting how online platforms enable various forms of sexism to thrive, there is a growing need for effective sexism detection and mitigation strategies. While computational approaches to sexism detection are widespread in high-resource languages, progress remains limited in low-resource languages where limited linguistic resources and cultural differences affect how sexism is expressed and perceived. This study introduces the first Hausa sexism detection dataset, developed through community engagement, qualitative coding, and data augmentation. For cultural nuances and linguistic representation, we conducted a two-stage user study (n=66) involving native speakers to explore how sexism is defined and articulated in everyday discourse. We further experiment with both traditional machine learning classifiers and pre-trained multilingual language models and evaluating the effectiveness few-shot learning in detecting sexism in Hausa. Our findings highlight challenges in capturing cultural nuance, particularly with clarification-seeking and idiomatic expressions, and reveal a tendency for many false positives in such cases.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dataset Creation and Baseline Models for Sexism Detection in Hausa
Muhammad, Fatima Adam
Hassan, Shamsuddeen Muhammad
Inuwa-Dutse, Isa
Computation and Language
Artificial Intelligence
Sexism reinforces gender inequality and social exclusion by perpetuating stereotypes, bias, and discriminatory norms. Noting how online platforms enable various forms of sexism to thrive, there is a growing need for effective sexism detection and mitigation strategies. While computational approaches to sexism detection are widespread in high-resource languages, progress remains limited in low-resource languages where limited linguistic resources and cultural differences affect how sexism is expressed and perceived. This study introduces the first Hausa sexism detection dataset, developed through community engagement, qualitative coding, and data augmentation. For cultural nuances and linguistic representation, we conducted a two-stage user study (n=66) involving native speakers to explore how sexism is defined and articulated in everyday discourse. We further experiment with both traditional machine learning classifiers and pre-trained multilingual language models and evaluating the effectiveness few-shot learning in detecting sexism in Hausa. Our findings highlight challenges in capturing cultural nuance, particularly with clarification-seeking and idiomatic expressions, and reveal a tendency for many false positives in such cases.
title Dataset Creation and Baseline Models for Sexism Detection in Hausa
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.27038