A First Context-Free Grammar Applied to Nawatl Corpora Augmentation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guzmán-Landa, Juan-José, Torres-Moreno, Juan-Manuel, Figueroa-Saavedra, Miguel, Quintana-Torres, Ligia, Avendaño-Garrido, Martha-Lorena, Ranger, Graham
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914077791485952
author Guzmán-Landa, Juan-José
Torres-Moreno, Juan-Manuel
Figueroa-Saavedra, Miguel
Quintana-Torres, Ligia
Avendaño-Garrido, Martha-Lorena
Ranger, Graham
author_facet Guzmán-Landa, Juan-José
Torres-Moreno, Juan-Manuel
Figueroa-Saavedra, Miguel
Quintana-Torres, Ligia
Avendaño-Garrido, Martha-Lorena
Ranger, Graham
contents In this article we introduce a context-free grammar (CFG) for the Nawatl language. Nawatl (or Nahuatl) is an Amerindian language of the $π$-language type, i.e. a language with few digital resources, in which the corpora available for machine learning are virtually non-existent. The objective here is to generate a significant number of grammatically correct artificial sentences, in order to increase the corpora available for language model training. We want to show that a grammar enables us significantly to expand a corpus in Nawatl which we call $π$-\textsc{yalli}. The corpus, thus enriched, enables us to train algorithms such as FastText and to evaluate them on sentence-level semantic tasks. Preliminary results show that by using the grammar, comparative improvements are achieved over some LLMs. However, it is observed that to achieve more significant improvement, grammars that model the Nawatl language even more effectively are required.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04945
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A First Context-Free Grammar Applied to Nawatl Corpora Augmentation
Guzmán-Landa, Juan-José
Torres-Moreno, Juan-Manuel
Figueroa-Saavedra, Miguel
Quintana-Torres, Ligia
Avendaño-Garrido, Martha-Lorena
Ranger, Graham
Computation and Language
Artificial Intelligence
In this article we introduce a context-free grammar (CFG) for the Nawatl language. Nawatl (or Nahuatl) is an Amerindian language of the $π$-language type, i.e. a language with few digital resources, in which the corpora available for machine learning are virtually non-existent. The objective here is to generate a significant number of grammatically correct artificial sentences, in order to increase the corpora available for language model training. We want to show that a grammar enables us significantly to expand a corpus in Nawatl which we call $π$-\textsc{yalli}. The corpus, thus enriched, enables us to train algorithms such as FastText and to evaluate them on sentence-level semantic tasks. Preliminary results show that by using the grammar, comparative improvements are achieved over some LLMs. However, it is observed that to achieve more significant improvement, grammars that model the Nawatl language even more effectively are required.
title A First Context-Free Grammar Applied to Nawatl Corpora Augmentation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.04945