Tokenizer Adapted for the Nasa Yuwe Language
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Artículo científico |
| Sprache: | en |
| Veröffentlicht: |
Instituto Politécnico Nacional
2016
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866817184213237760 |
|---|---|
| author | Luz Marina Sierra Martínez |
| author_facet | Luz Marina Sierra Martínez |
| contents | Tokenizer Adapted for the Nasa Yuwe Language Luz Marina Sierra Martínez Carlos Alberto Cobos Lozada Juan Carlos Corrales Computación Nasa Yuwe language tokenizer for Nasa Yuwe Nasa indigenous community information retrieval for texts written in Nasa Yuwe In Colombia, ethnic and cultural diversity is conceived by the government to be a social right. Such diversity finds expression, among other ways, in a large number of indigenous languages, which have been kept alive for centuries. However, efforts toward conservation and preservation of these languages have generally fallen short. This is the case for the Nasa Yuwe language, spoken by the Nasa, or Páez, indigenous community, the status of which is endangered. Given such a predicament, the use of technology has been found to provide a strategic opportunity for adaptation, ownership, and development of Nasa Yuwe within the social and cultural environment of the Nasa people. The technology includes the use of computational techniques, which allow the exchange of information by means of IR activities. These encourage different, new possibilities for the Nasa people to be able to interact in Nasa Yuwe. It has therefore become necessary to adapt the stages of the IR process to this language. The current paper specifically presents a process for adapting a tokenizer to texts written in Nasa Yuwe. This involves making use of the precision-recall curve as an evaluation and comparison measure. The results presented allow appreciation of all stages in the process of adapting the standard tokenizer to produce the Nasa version, of the Nasa tokenizer and its results over texts written in Nasa Yuwe, and of the analysis of the precision-recall curve baseline in contrast to that of the Nasa tokenizer. 2016 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61547469005 en http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.20 |
| format | Artículo científico |
| id | redalyc_61547469005 |
| language | en |
| publishDate | 2016 |
| publisher | Instituto Politécnico Nacional |
| spellingShingle | Tokenizer Adapted for the Nasa Yuwe Language Luz Marina Sierra Martínez Computación Nasa Yuwe language tokenizer for Nasa Yuwe Nasa indigenous community information retrieval for texts written in Nasa Yuwe Tokenizer Adapted for the Nasa Yuwe Language Luz Marina Sierra Martínez Carlos Alberto Cobos Lozada Juan Carlos Corrales Computación Nasa Yuwe language tokenizer for Nasa Yuwe Nasa indigenous community information retrieval for texts written in Nasa Yuwe In Colombia, ethnic and cultural diversity is conceived by the government to be a social right. Such diversity finds expression, among other ways, in a large number of indigenous languages, which have been kept alive for centuries. However, efforts toward conservation and preservation of these languages have generally fallen short. This is the case for the Nasa Yuwe language, spoken by the Nasa, or Páez, indigenous community, the status of which is endangered. Given such a predicament, the use of technology has been found to provide a strategic opportunity for adaptation, ownership, and development of Nasa Yuwe within the social and cultural environment of the Nasa people. The technology includes the use of computational techniques, which allow the exchange of information by means of IR activities. These encourage different, new possibilities for the Nasa people to be able to interact in Nasa Yuwe. It has therefore become necessary to adapt the stages of the IR process to this language. The current paper specifically presents a process for adapting a tokenizer to texts written in Nasa Yuwe. This involves making use of the precision-recall curve as an evaluation and comparison measure. The results presented allow appreciation of all stages in the process of adapting the standard tokenizer to produce the Nasa version, of the Nasa tokenizer and its results over texts written in Nasa Yuwe, and of the analysis of the precision-recall curve baseline in contrast to that of the Nasa tokenizer. 2016 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61547469005 en http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.20 |
| title | Tokenizer Adapted for the Nasa Yuwe Language |
| topic | Computación Nasa Yuwe language tokenizer for Nasa Yuwe Nasa indigenous community information retrieval for texts written in Nasa Yuwe |
| url | https://www.redalyc.org/articulo.oa?id=61547469005 |