$π$-yalli: un nouveau corpus pour le nahuatl

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Torres-Moreno, Juan-Manuel, Guzmán-Landa, Juan-José, Ranger, Graham, Garrido, Martha Lorena Avendaño, Figueroa-Saavedra, Miguel, Quintana-Torres, Ligia, González-Gallardo, Carlos-Emiliano, Pontes, Elvys Linhares, Morales, Patricia Velázquez, Jiménez, Luis-Gil Moreno
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917875539771392
author Torres-Moreno, Juan-Manuel
Guzmán-Landa, Juan-José
Ranger, Graham
Garrido, Martha Lorena Avendaño
Figueroa-Saavedra, Miguel
Quintana-Torres, Ligia
González-Gallardo, Carlos-Emiliano
Pontes, Elvys Linhares
Morales, Patricia Velázquez
Jiménez, Luis-Gil Moreno
author_facet Torres-Moreno, Juan-Manuel
Guzmán-Landa, Juan-José
Ranger, Graham
Garrido, Martha Lorena Avendaño
Figueroa-Saavedra, Miguel
Quintana-Torres, Ligia
González-Gallardo, Carlos-Emiliano
Pontes, Elvys Linhares
Morales, Patricia Velázquez
Jiménez, Luis-Gil Moreno
contents The NAHU$^2$ project is a Franco-Mexican collaboration aimed at building the $π$-YALLI corpus adapted to machine learning, which will subsequently be used to develop computer resources for the Nahuatl language. Nahuatl is a language with few computational resources, even though it is a living language spoken by around 2 million people. We have decided to build $π$-YALLI, a corpus that will enable to carry out research on Nahuatl in order to develop Language Models (LM), whether dynamic or not, which will make it possible to in turn enable the development of Natural Language Processing (NLP) tools such as: a) a grapheme unifier, b) a word segmenter, c) a POS grammatical analyser, d) a content-based Automatic Text Summarization; and possibly, e) a translator translator (probabilistic or learning-based).
format Preprint
id arxiv_https___arxiv_org_abs_2412_15821
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle $π$-yalli: un nouveau corpus pour le nahuatl
Torres-Moreno, Juan-Manuel
Guzmán-Landa, Juan-José
Ranger, Graham
Garrido, Martha Lorena Avendaño
Figueroa-Saavedra, Miguel
Quintana-Torres, Ligia
González-Gallardo, Carlos-Emiliano
Pontes, Elvys Linhares
Morales, Patricia Velázquez
Jiménez, Luis-Gil Moreno
Computation and Language
Artificial Intelligence
The NAHU$^2$ project is a Franco-Mexican collaboration aimed at building the $π$-YALLI corpus adapted to machine learning, which will subsequently be used to develop computer resources for the Nahuatl language. Nahuatl is a language with few computational resources, even though it is a living language spoken by around 2 million people. We have decided to build $π$-YALLI, a corpus that will enable to carry out research on Nahuatl in order to develop Language Models (LM), whether dynamic or not, which will make it possible to in turn enable the development of Natural Language Processing (NLP) tools such as: a) a grapheme unifier, b) a word segmenter, c) a POS grammatical analyser, d) a content-based Automatic Text Summarization; and possibly, e) a translator translator (probabilistic or learning-based).
title $π$-yalli: un nouveau corpus pour le nahuatl
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.15821