Stemming -- The Evolution and Current State with a Focus on Bangla

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paul, Abhijit, Farin, Mashiat Amin, Abdullah, Sharif Md., Kabir, Ahmedul, Masud, Zarif, Rayana, Shebuti
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908497406328832
author Paul, Abhijit
Farin, Mashiat Amin
Abdullah, Sharif Md.
Kabir, Ahmedul
Masud, Zarif
Rayana, Shebuti
author_facet Paul, Abhijit
Farin, Mashiat Amin
Abdullah, Sharif Md.
Kabir, Ahmedul
Masud, Zarif
Rayana, Shebuti
contents Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing step in language analysis, is essential for low-resource, highly-inflectional languages like Bangla, because it can reduce the complexity of algorithms and models by significantly reducing the number of words the algorithm needs to consider. This paper conducts a comprehensive survey of stemming approaches, emphasizing the importance of handling morphological variants effectively. While exploring the landscape of Bangla stemming, it becomes evident that there is a significant gap in the existing literature. The paper highlights the discontinuity from previous research and the scarcity of accessible implementations for replication. Furthermore, it critiques the evaluation methodologies, stressing the need for more relevant metrics. In the context of Bangla's rich morphology and diverse dialects, the paper acknowledges the challenges it poses. To address these challenges, the paper suggests directions for Bangla stemmer development. It concludes by advocating for robust Bangla stemmers and continued research in the field to enhance language analysis and processing.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stemming -- The Evolution and Current State with a Focus on Bangla
Paul, Abhijit
Farin, Mashiat Amin
Abdullah, Sharif Md.
Kabir, Ahmedul
Masud, Zarif
Rayana, Shebuti
Computation and Language
Information Retrieval
Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing step in language analysis, is essential for low-resource, highly-inflectional languages like Bangla, because it can reduce the complexity of algorithms and models by significantly reducing the number of words the algorithm needs to consider. This paper conducts a comprehensive survey of stemming approaches, emphasizing the importance of handling morphological variants effectively. While exploring the landscape of Bangla stemming, it becomes evident that there is a significant gap in the existing literature. The paper highlights the discontinuity from previous research and the scarcity of accessible implementations for replication. Furthermore, it critiques the evaluation methodologies, stressing the need for more relevant metrics. In the context of Bangla's rich morphology and diverse dialects, the paper acknowledges the challenges it poses. To address these challenges, the paper suggests directions for Bangla stemmer development. It concludes by advocating for robust Bangla stemmers and continued research in the field to enhance language analysis and processing.
title Stemming -- The Evolution and Current State with a Focus on Bangla
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2508.15711