Development of Pre-Trained Transformer-based Models for the Nepali Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thapa, Prajwal, Nyachhyon, Jinu, Sharma, Mridul, Bal, Bal Krishna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915450891272192
author Thapa, Prajwal
Nyachhyon, Jinu
Sharma, Mridul
Bal, Bal Krishna
author_facet Thapa, Prajwal
Nyachhyon, Jinu
Sharma, Mridul
Bal, Bal Krishna
contents Transformer-based pre-trained language models have dominated the field of Natural Language Processing (NLP) for quite some time now. However, the Nepali language, spoken by approximately 32 million people worldwide, remains significantly underrepresented in this domain. This underrepresentation is primarily attributed to the scarcity of monolingual data corpora and limited available resources for the Nepali language. While existing efforts have predominantly concentrated on basic encoder-based models, there is a notable gap in the exploration of decoder-based architectures. To address this gap, we have collected 27.5 GB of Nepali text data, approximately 2.4x larger than any previously available Nepali language corpus. Leveraging this data, we pre-trained three different models i.e., BERT, RoBERTa, and GPT-2, exclusively for the Nepali Language. Furthermore, we performed instruction tuning and explored its potential for monolingual Nepali data, providing a foundation for future research. Our models outperformed the existing best model by 2 points on Nep-gLUE benchmark, scoring 95.60 and also outperformed existing models on text generation tasks, demonstrating improvements in both understanding and generating Nepali text.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15734
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Development of Pre-Trained Transformer-based Models for the Nepali Language
Thapa, Prajwal
Nyachhyon, Jinu
Sharma, Mridul
Bal, Bal Krishna
Computation and Language
Machine Learning
Transformer-based pre-trained language models have dominated the field of Natural Language Processing (NLP) for quite some time now. However, the Nepali language, spoken by approximately 32 million people worldwide, remains significantly underrepresented in this domain. This underrepresentation is primarily attributed to the scarcity of monolingual data corpora and limited available resources for the Nepali language. While existing efforts have predominantly concentrated on basic encoder-based models, there is a notable gap in the exploration of decoder-based architectures. To address this gap, we have collected 27.5 GB of Nepali text data, approximately 2.4x larger than any previously available Nepali language corpus. Leveraging this data, we pre-trained three different models i.e., BERT, RoBERTa, and GPT-2, exclusively for the Nepali Language. Furthermore, we performed instruction tuning and explored its potential for monolingual Nepali data, providing a foundation for future research. Our models outperformed the existing best model by 2 points on Nep-gLUE benchmark, scoring 95.60 and also outperformed existing models on text generation tasks, demonstrating improvements in both understanding and generating Nepali text.
title Development of Pre-Trained Transformer-based Models for the Nepali Language
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.15734