Who Wrote This? Identifying Machine vs Human-Generated Text in Hausa

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sani, Babangida, Soy, Aakansha, Imam, Sukairaj Hafiz, Mustapha, Ahmad, Aliyu, Lukman Jibril, Abdulmumin, Idris, Ahmad, Ibrahim Said, Muhammad, Shamsuddeen Hassan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912278444507136
author Sani, Babangida
Soy, Aakansha
Imam, Sukairaj Hafiz
Mustapha, Ahmad
Aliyu, Lukman Jibril
Abdulmumin, Idris
Ahmad, Ibrahim Said
Muhammad, Shamsuddeen Hassan
author_facet Sani, Babangida
Soy, Aakansha
Imam, Sukairaj Hafiz
Mustapha, Ahmad
Aliyu, Lukman Jibril
Abdulmumin, Idris
Ahmad, Ibrahim Said
Muhammad, Shamsuddeen Hassan
contents The advancement of large language models (LLMs) has allowed them to be proficient in various tasks, including content generation. However, their unregulated usage can lead to malicious activities such as plagiarism and generating and spreading fake news, especially for low-resource languages. Most existing machine-generated text detectors are trained on high-resource languages like English, French, etc. In this study, we developed the first large-scale detector that can distinguish between human- and machine-generated content in Hausa. We scrapped seven Hausa-language media outlets for the human-generated text and the Gemini-2.0 flash model to automatically generate the corresponding Hausa-language articles based on the human-generated article headlines. We fine-tuned four pre-trained Afri-centric models (AfriTeVa, AfriBERTa, AfroXLMR, and AfroXLMR-76L) on the resulting dataset and assessed their performance using accuracy and F1-score metrics. AfroXLMR achieved the highest performance with an accuracy of 99.23% and an F1 score of 99.21%, demonstrating its effectiveness for Hausa text detection. Our dataset is made publicly available to enable further research.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13101
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Who Wrote This? Identifying Machine vs Human-Generated Text in Hausa
Sani, Babangida
Soy, Aakansha
Imam, Sukairaj Hafiz
Mustapha, Ahmad
Aliyu, Lukman Jibril
Abdulmumin, Idris
Ahmad, Ibrahim Said
Muhammad, Shamsuddeen Hassan
Computation and Language
The advancement of large language models (LLMs) has allowed them to be proficient in various tasks, including content generation. However, their unregulated usage can lead to malicious activities such as plagiarism and generating and spreading fake news, especially for low-resource languages. Most existing machine-generated text detectors are trained on high-resource languages like English, French, etc. In this study, we developed the first large-scale detector that can distinguish between human- and machine-generated content in Hausa. We scrapped seven Hausa-language media outlets for the human-generated text and the Gemini-2.0 flash model to automatically generate the corresponding Hausa-language articles based on the human-generated article headlines. We fine-tuned four pre-trained Afri-centric models (AfriTeVa, AfriBERTa, AfroXLMR, and AfroXLMR-76L) on the resulting dataset and assessed their performance using accuracy and F1-score metrics. AfroXLMR achieved the highest performance with an accuracy of 99.23% and an F1 score of 99.21%, demonstrating its effectiveness for Hausa text detection. Our dataset is made publicly available to enable further research.
title Who Wrote This? Identifying Machine vs Human-Generated Text in Hausa
topic Computation and Language
url https://arxiv.org/abs/2503.13101