R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915556683153408 |
|---|---|
| author | Keita, Mamadou K. Homan, Christopher Diarra, Sebastien |
| author_facet | Keita, Mamadou K. Homan, Christopher Diarra, Sebastien |
| contents | We introduce the Rule-to-Tag (R2T) framework, a hybrid approach that integrates a multi-tiered system of linguistic rules directly into a neural network's training objective. R2T's novelty lies in its adaptive loss function, which includes a regularization term that teaches the model to handle out-of-vocabulary (OOV) words with principled uncertainty. We frame this work as a case study in a paradigm we call principled learning (PrL), where models are trained with explicit task constraints rather than on labeled examples alone. Our experiments on Zarma part-of-speech (POS) tagging show that the R2T-BiLSTM model, trained only on unlabeled text, achieves 98.2% accuracy, outperforming baselines like AfriBERTa fine-tuned on 300 labeled sentences. We further show that for more complex tasks like named entity recognition (NER), R2T serves as a powerful pre-training step; a model pre-trained with R2T and fine-tuned on just 50 labeled sentences outperformes a baseline trained on 300. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_13854 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging Keita, Mamadou K. Homan, Christopher Diarra, Sebastien Computation and Language Machine Learning We introduce the Rule-to-Tag (R2T) framework, a hybrid approach that integrates a multi-tiered system of linguistic rules directly into a neural network's training objective. R2T's novelty lies in its adaptive loss function, which includes a regularization term that teaches the model to handle out-of-vocabulary (OOV) words with principled uncertainty. We frame this work as a case study in a paradigm we call principled learning (PrL), where models are trained with explicit task constraints rather than on labeled examples alone. Our experiments on Zarma part-of-speech (POS) tagging show that the R2T-BiLSTM model, trained only on unlabeled text, achieves 98.2% accuracy, outperforming baselines like AfriBERTa fine-tuned on 300 labeled sentences. We further show that for more complex tasks like named entity recognition (NER), R2T serves as a powerful pre-training step; a model pre-trained with R2T and fine-tuned on just 50 labeled sentences outperformes a baseline trained on 300. |
| title | R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.13854 |