Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raihan, Md Nishat, Goswami, Dhiman, Mahmud, Antara
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911796876541952
author Raihan, Md Nishat
Goswami, Dhiman
Mahmud, Antara
author_facet Raihan, Md Nishat
Goswami, Dhiman
Mahmud, Antara
contents One of the most popular downstream tasks in the field of Natural Language Processing is text classification. Text classification tasks have become more daunting when the texts are code-mixed. Though they are not exposed to such text during pre-training, different BERT models have demonstrated success in tackling Code-Mixed NLP challenges. Again, in order to enhance their performance, Code-Mixed NLP models have depended on combining synthetic data with real-world data. It is crucial to understand how the BERT models' performance is impacted when they are pretrained using corresponding code-mixed languages. In this paper, we introduce Tri-Distil-BERT, a multilingual model pre-trained on Bangla, English, and Hindi, and Mixed-Distil-BERT, a model fine-tuned on code-mixed data. Both models are evaluated across multiple NLP tasks and demonstrate competitive performance against larger models like mBERT and XLM-R. Our two-tiered pre-training approach offers efficient alternatives for multilingual and code-mixed language understanding, contributing to advancements in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2309_10272
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi
Raihan, Md Nishat
Goswami, Dhiman
Mahmud, Antara
Computation and Language
One of the most popular downstream tasks in the field of Natural Language Processing is text classification. Text classification tasks have become more daunting when the texts are code-mixed. Though they are not exposed to such text during pre-training, different BERT models have demonstrated success in tackling Code-Mixed NLP challenges. Again, in order to enhance their performance, Code-Mixed NLP models have depended on combining synthetic data with real-world data. It is crucial to understand how the BERT models' performance is impacted when they are pretrained using corresponding code-mixed languages. In this paper, we introduce Tri-Distil-BERT, a multilingual model pre-trained on Bangla, English, and Hindi, and Mixed-Distil-BERT, a model fine-tuned on code-mixed data. Both models are evaluated across multiple NLP tasks and demonstrate competitive performance against larger models like mBERT and XLM-R. Our two-tiered pre-training approach offers efficient alternatives for multilingual and code-mixed language understanding, contributing to advancements in the field.
title Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi
topic Computation and Language
url https://arxiv.org/abs/2309.10272