XSTEM: An exemplar-based stemming algorithm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Baker, Kirk
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917682077499392
author Baker, Kirk
author_facet Baker, Kirk
contents Stemming is the process of reducing related words to a standard form by removing affixes from them. Existing algorithms vary with respect to their complexity, configurability, handling of unknown words, and ability to avoid under- and over-stemming. This paper presents a fast, simple, configurable, high-precision, high-recall stemming algorithm that combines the simplicity and performance of word-based lookup tables with the strong generalizability of rule-based methods to avert problems with out-of-vocabulary words.
format Preprint
id arxiv_https___arxiv_org_abs_2205_04355
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle XSTEM: An exemplar-based stemming algorithm
Baker, Kirk
Computation and Language
Stemming is the process of reducing related words to a standard form by removing affixes from them. Existing algorithms vary with respect to their complexity, configurability, handling of unknown words, and ability to avoid under- and over-stemming. This paper presents a fast, simple, configurable, high-precision, high-recall stemming algorithm that combines the simplicity and performance of word-based lookup tables with the strong generalizability of rule-based methods to avert problems with out-of-vocabulary words.
title XSTEM: An exemplar-based stemming algorithm
topic Computation and Language
url https://arxiv.org/abs/2205.04355