BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kabir, Muhammad Rafsan, Nabil, Md. Mohibur Rahman, Khan, Mohammad Ashrafuzzaman
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909400891916288
author Kabir, Muhammad Rafsan
Nabil, Md. Mohibur Rahman
Khan, Mohammad Ashrafuzzaman
author_facet Kabir, Muhammad Rafsan
Nabil, Md. Mohibur Rahman
Khan, Mohammad Ashrafuzzaman
contents Sentence-level embedding is essential for various tasks that require understanding natural language. Many studies have explored such embeddings for high-resource languages like English. However, low-resource languages like Bengali (a language spoken by almost two hundred and thirty million people) are still under-explored. This work introduces two lightweight sentence transformers for the Bangla language, leveraging a novel cross-lingual knowledge distillation approach. This method distills knowledge from a pre-trained, high-performing English sentence transformer. Proposed models are evaluated across multiple downstream tasks, including paraphrase detection, semantic textual similarity (STS), and Bangla hate speech detection. The new method consistently outperformed existing Bangla sentence transformers. Moreover, the lightweight architecture and shorter inference time make the models highly suitable for deployment in resource-constrained environments, making them valuable for practical NLP applications in low-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15270
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques
Kabir, Muhammad Rafsan
Nabil, Md. Mohibur Rahman
Khan, Mohammad Ashrafuzzaman
Computation and Language
Machine Learning
Sentence-level embedding is essential for various tasks that require understanding natural language. Many studies have explored such embeddings for high-resource languages like English. However, low-resource languages like Bengali (a language spoken by almost two hundred and thirty million people) are still under-explored. This work introduces two lightweight sentence transformers for the Bangla language, leveraging a novel cross-lingual knowledge distillation approach. This method distills knowledge from a pre-trained, high-performing English sentence transformer. Proposed models are evaluated across multiple downstream tasks, including paraphrase detection, semantic textual similarity (STS), and Bangla hate speech detection. The new method consistently outperformed existing Bangla sentence transformers. Moreover, the lightweight architecture and shorter inference time make the models highly suitable for deployment in resource-constrained environments, making them valuable for practical NLP applications in low-resource languages.
title BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.15270