ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914587169783808 |
|---|---|
| author | Zaghouani, Wajdi Ibrahim, Shimaa Amer Bessghaier, Mabrouka Bouamor, Houda |
| author_facet | Zaghouani, Wajdi Ibrahim, Shimaa Amer Bessghaier, Mabrouka Bouamor, Houda |
| contents | We present ArabDiscrim, a decade-long lexical resource and corpus of 293K public Arabic Facebook posts (2014--2024) discussing racism and discrimination. Unlike existing Twitter-centric datasets, ArabDiscrim integrates platform-native engagement signals, including reactions, shares, comments, and page metadata, enabling joint analysis of language and audience response. The resource includes 200 curated terms (100 racism-related and 100 discrimination-related) with morphological regex families (13+ inflections per lemma), and 20 discrimination axes capturing identity-based grounds for unequal treatment. It also provides explicit attribution patterns. Released under a restricted research-use license for ethical compliance with platform terms, ArabDiscrim supports weak supervision, axis-aware sampling, and platform ecology research. By bridging lexical depth and ecological validity, it establishes a foundation for fairness-oriented, platform-aware Arabic NLP. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_22081 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination Zaghouani, Wajdi Ibrahim, Shimaa Amer Bessghaier, Mabrouka Bouamor, Houda Computation and Language We present ArabDiscrim, a decade-long lexical resource and corpus of 293K public Arabic Facebook posts (2014--2024) discussing racism and discrimination. Unlike existing Twitter-centric datasets, ArabDiscrim integrates platform-native engagement signals, including reactions, shares, comments, and page metadata, enabling joint analysis of language and audience response. The resource includes 200 curated terms (100 racism-related and 100 discrimination-related) with morphological regex families (13+ inflections per lemma), and 20 discrimination axes capturing identity-based grounds for unequal treatment. It also provides explicit attribution patterns. Released under a restricted research-use license for ethical compliance with platform terms, ArabDiscrim supports weak supervision, axis-aware sampling, and platform ecology research. By bridging lexical depth and ecological validity, it establishes a foundation for fairness-oriented, platform-aware Arabic NLP. |
| title | ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2605.22081 |