Text Categorization Can Enhance Domain-Agnostic Stopword Extraction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Turki, Houcemeddine, Etori, Naome A., Taieb, Mohamed Ali Hadj, Omotayo, Abdul-Hakeem, Emezue, Chris Chinenye, Aouicha, Mohamed Ben, Awokoya, Ayodele, Lawan, Falalu Ibrahim, Nixdorf, Doreen
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929221894406144
author Turki, Houcemeddine
Etori, Naome A.
Taieb, Mohamed Ali Hadj
Omotayo, Abdul-Hakeem
Emezue, Chris Chinenye
Aouicha, Mohamed Ben
Awokoya, Ayodele
Lawan, Falalu Ibrahim
Nixdorf, Doreen
author_facet Turki, Houcemeddine
Etori, Naome A.
Taieb, Mohamed Ali Hadj
Omotayo, Abdul-Hakeem
Emezue, Chris Chinenye
Aouicha, Mohamed Ben
Awokoya, Ayodele
Lawan, Falalu Ibrahim
Nixdorf, Doreen
contents This paper investigates the role of text categorization in streamlining stopword extraction in natural language processing (NLP), specifically focusing on nine African languages alongside French. By leveraging the MasakhaNEWS, African Stopwords Project, and MasakhaPOS datasets, our findings emphasize that text categorization effectively identifies domain-agnostic stopwords with over 80% detection success rate for most examined languages. Nevertheless, linguistic variances result in lower detection rates for certain languages. Interestingly, we find that while over 40% of stopwords are common across news categories, less than 15% are unique to a single category. Uncommon stopwords add depth to text but their classification as stopwords depends on context. Therefore combining statistical and linguistic approaches creates comprehensive stopword lists, highlighting the value of our hybrid method. This research enhances NLP for African languages and underscores the importance of text categorization in stopword extraction.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13398
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text Categorization Can Enhance Domain-Agnostic Stopword Extraction
Turki, Houcemeddine
Etori, Naome A.
Taieb, Mohamed Ali Hadj
Omotayo, Abdul-Hakeem
Emezue, Chris Chinenye
Aouicha, Mohamed Ben
Awokoya, Ayodele
Lawan, Falalu Ibrahim
Nixdorf, Doreen
Computation and Language
Machine Learning
This paper investigates the role of text categorization in streamlining stopword extraction in natural language processing (NLP), specifically focusing on nine African languages alongside French. By leveraging the MasakhaNEWS, African Stopwords Project, and MasakhaPOS datasets, our findings emphasize that text categorization effectively identifies domain-agnostic stopwords with over 80% detection success rate for most examined languages. Nevertheless, linguistic variances result in lower detection rates for certain languages. Interestingly, we find that while over 40% of stopwords are common across news categories, less than 15% are unique to a single category. Uncommon stopwords add depth to text but their classification as stopwords depends on context. Therefore combining statistical and linguistic approaches creates comprehensive stopword lists, highlighting the value of our hybrid method. This research enhances NLP for African languages and underscores the importance of text categorization in stopword extraction.
title Text Categorization Can Enhance Domain-Agnostic Stopword Extraction
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2401.13398