MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910008111792128 |
|---|---|
| author | Yadav, Sumit Yadav, Raju Kumar Maskey, Utsav Kashyap, Gautam Siddharth Gautam, Ganesh Naseem, Usman |
| author_facet | Yadav, Sumit Yadav, Raju Kumar Maskey, Utsav Kashyap, Gautam Siddharth Gautam, Ganesh Naseem, Usman |
| contents | Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due to the scarcity of high-quality data and language-specific models. Maithili, despite being spoken by millions, lacks adequate computational resources, limiting its inclusion in digital and AI-driven applications. To address this gap, we introducemaiBERT, a BERT-based language model pre-trained specifically for Maithili using the Masked Language Modeling (MLM) technique. Our model is trained on a newly constructed Maithili corpus and evaluated through a news classification task. In our experiments, maiBERT achieved an accuracy of 87.02%, outperforming existing regional models like NepBERTa and HindiBERT, with a 0.13% overall accuracy gain and 5-7% improvement across various classes. We have open-sourced maiBERT on Hugging Face enabling further fine-tuning for downstream tasks such as sentiment analysis and Named Entity Recognition (NER). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15048 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language Yadav, Sumit Yadav, Raju Kumar Maskey, Utsav Kashyap, Gautam Siddharth Gautam, Ganesh Naseem, Usman Computation and Language Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due to the scarcity of high-quality data and language-specific models. Maithili, despite being spoken by millions, lacks adequate computational resources, limiting its inclusion in digital and AI-driven applications. To address this gap, we introducemaiBERT, a BERT-based language model pre-trained specifically for Maithili using the Masked Language Modeling (MLM) technique. Our model is trained on a newly constructed Maithili corpus and evaluated through a news classification task. In our experiments, maiBERT achieved an accuracy of 87.02%, outperforming existing regional models like NepBERTa and HindiBERT, with a 0.13% overall accuracy gain and 5-7% improvement across various classes. We have open-sourced maiBERT on Hugging Face enabling further fine-tuning for downstream tasks such as sentiment analysis and Named Entity Recognition (NER). |
| title | MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.15048 |