Leveraging Transformer-Based Models for Predicting Inflection Classes of Words in an Endangered Sami Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alnajjar, Khalid, Hämäläinen, Mika, Rueter, Jack
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929576446263296
author Alnajjar, Khalid
Hämäläinen, Mika
Rueter, Jack
author_facet Alnajjar, Khalid
Hämäläinen, Mika
Rueter, Jack
contents This paper presents a methodology for training a transformer-based model to classify lexical and morphosyntactic features of Skolt Sami, an endangered Uralic language characterized by complex morphology. The goal of our approach is to create an effective system for understanding and analyzing Skolt Sami, given the limited data availability and linguistic intricacies inherent to the language. Our end-to-end pipeline includes data extraction, augmentation, and training a transformer-based model capable of predicting inflection classes. The motivation behind this work is to support language preservation and revitalization efforts for minority languages like Skolt Sami. Accurate classification not only helps improve the state of Finite-State Transducers (FSTs) by providing greater lexical coverage but also contributes to systematic linguistic documentation for researchers working with newly discovered words from literature and native speakers. Our model achieves an average weighted F1 score of 1.00 for POS classification and 0.81 for inflection class classification. The trained model and code will be released publicly to facilitate future research in endangered NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02556
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Transformer-Based Models for Predicting Inflection Classes of Words in an Endangered Sami Language
Alnajjar, Khalid
Hämäläinen, Mika
Rueter, Jack
Computation and Language
This paper presents a methodology for training a transformer-based model to classify lexical and morphosyntactic features of Skolt Sami, an endangered Uralic language characterized by complex morphology. The goal of our approach is to create an effective system for understanding and analyzing Skolt Sami, given the limited data availability and linguistic intricacies inherent to the language. Our end-to-end pipeline includes data extraction, augmentation, and training a transformer-based model capable of predicting inflection classes. The motivation behind this work is to support language preservation and revitalization efforts for minority languages like Skolt Sami. Accurate classification not only helps improve the state of Finite-State Transducers (FSTs) by providing greater lexical coverage but also contributes to systematic linguistic documentation for researchers working with newly discovered words from literature and native speakers. Our model achieves an average weighted F1 score of 1.00 for POS classification and 0.81 for inflection class classification. The trained model and code will be released publicly to facilitate future research in endangered NLP.
title Leveraging Transformer-Based Models for Predicting Inflection Classes of Words in an Endangered Sami Language
topic Computation and Language
url https://arxiv.org/abs/2411.02556