Surfing the modeling of PoS taggers in low-resource scenarios

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ferro, Manuel Vilares, Bilbao, Víctor M. Darriba, Ribadas-Pena, Francisco J., Gil, Jorge Graña
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913223048953856
author Ferro, Manuel Vilares
Bilbao, Víctor M. Darriba
Ribadas-Pena, Francisco J.
Gil, Jorge Graña
author_facet Ferro, Manuel Vilares
Bilbao, Víctor M. Darriba
Ribadas-Pena, Francisco J.
Gil, Jorge Graña
contents The recent trend towards the application of deep structured techniques has revealed the limits of huge models in natural language processing. This has reawakened the interest in traditional machine learning algorithms, which have proved still to be competitive in certain contexts, in particular low-resource settings. In parallel, model selection has become an essential task to boost performance at reasonable cost, even more so when we talk about processes involving domains where the training and/or computational resources are scarce. Against this backdrop, we evaluate the early estimation of learning curves as a practical mechanism for selecting the most appropriate model in scenarios characterized by the use of non-deep learners in resource-lean settings. On the basis of a formal approximation model previously evaluated under conditions of wide availability of training and validation resources, we study the reliability of such an approach in a different and much more demanding operationalenvironment. Using as case study the generation of PoS taggers for Galician, a language belonging to the Western Ibero-Romance group, the experimental results are consistent with our expectations.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02449
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Surfing the modeling of PoS taggers in low-resource scenarios
Ferro, Manuel Vilares
Bilbao, Víctor M. Darriba
Ribadas-Pena, Francisco J.
Gil, Jorge Graña
Computation and Language
Machine Learning
68, 68T50
The recent trend towards the application of deep structured techniques has revealed the limits of huge models in natural language processing. This has reawakened the interest in traditional machine learning algorithms, which have proved still to be competitive in certain contexts, in particular low-resource settings. In parallel, model selection has become an essential task to boost performance at reasonable cost, even more so when we talk about processes involving domains where the training and/or computational resources are scarce. Against this backdrop, we evaluate the early estimation of learning curves as a practical mechanism for selecting the most appropriate model in scenarios characterized by the use of non-deep learners in resource-lean settings. On the basis of a formal approximation model previously evaluated under conditions of wide availability of training and validation resources, we study the reliability of such an approach in a different and much more demanding operationalenvironment. Using as case study the generation of PoS taggers for Galician, a language belonging to the Western Ibero-Romance group, the experimental results are consistent with our expectations.
title Surfing the modeling of PoS taggers in low-resource scenarios
topic Computation and Language
Machine Learning
68, 68T50
url https://arxiv.org/abs/2402.02449