YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Olawole, Akindele Michael, Alabi, Jesujoba O., Sakpere, Aderonke Busayo, Adelani, David I.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913628791242752
author Olawole, Akindele Michael
Alabi, Jesujoba O.
Sakpere, Aderonke Busayo
Adelani, David I.
author_facet Olawole, Akindele Michael
Alabi, Jesujoba O.
Sakpere, Aderonke Busayo
Adelani, David I.
contents In this work, we present Yorùbá automatic diacritization (YAD) benchmark dataset for evaluating Yorùbá diacritization systems. In addition, we pre-train text-to-text transformer, T5 model for Yorùbá and showed that this model outperform several multilingually trained T5 models. Lastly, we showed that more data and larger models are better at diacritization for Yorùbá
format Preprint
id arxiv_https___arxiv_org_abs_2412_20218
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text
Olawole, Akindele Michael
Alabi, Jesujoba O.
Sakpere, Aderonke Busayo
Adelani, David I.
Computation and Language
In this work, we present Yorùbá automatic diacritization (YAD) benchmark dataset for evaluating Yorùbá diacritization systems. In addition, we pre-train text-to-text transformer, T5 model for Yorùbá and showed that this model outperform several multilingually trained T5 models. Lastly, we showed that more data and larger models are better at diacritization for Yorùbá
title YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text
topic Computation and Language
url https://arxiv.org/abs/2412.20218