XSTEM: An exemplar-based stemming algorithm
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917682077499392 |
|---|---|
| author | Baker, Kirk |
| author_facet | Baker, Kirk |
| contents | Stemming is the process of reducing related words to a standard form by removing affixes from them. Existing algorithms vary with respect to their complexity, configurability, handling of unknown words, and ability to avoid under- and over-stemming. This paper presents a fast, simple, configurable, high-precision, high-recall stemming algorithm that combines the simplicity and performance of word-based lookup tables with the strong generalizability of rule-based methods to avert problems with out-of-vocabulary words. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2205_04355 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | XSTEM: An exemplar-based stemming algorithm Baker, Kirk Computation and Language Stemming is the process of reducing related words to a standard form by removing affixes from them. Existing algorithms vary with respect to their complexity, configurability, handling of unknown words, and ability to avoid under- and over-stemming. This paper presents a fast, simple, configurable, high-precision, high-recall stemming algorithm that combines the simplicity and performance of word-based lookup tables with the strong generalizability of rule-based methods to avert problems with out-of-vocabulary words. |
| title | XSTEM: An exemplar-based stemming algorithm |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2205.04355 |