Non-Contextual BERT or FastText? A Comparative Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shanbhag, Abhay, Jadhav, Suramya, Thakurdesai, Amogh, Sinare, Ridhima, Joshi, Raviraj
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912238849228800
author Shanbhag, Abhay
Jadhav, Suramya
Thakurdesai, Amogh
Sinare, Ridhima
Joshi, Raviraj
author_facet Shanbhag, Abhay
Jadhav, Suramya
Thakurdesai, Amogh
Sinare, Ridhima
Joshi, Raviraj
contents Natural Language Processing (NLP) for low-resource languages, which lack large annotated datasets, faces significant challenges due to limited high-quality data and linguistic resources. The selection of embeddings plays a critical role in achieving strong performance in NLP tasks. While contextual BERT embeddings require a full forward pass, non-contextual BERT embeddings rely only on table lookup. Existing research has primarily focused on contextual BERT embeddings, leaving non-contextual embeddings largely unexplored. In this study, we analyze the effectiveness of non-contextual embeddings from BERT models (MuRIL and MahaBERT) and FastText models (IndicFT and MahaFT) for tasks such as news classification, sentiment analysis, and hate speech detection in one such low-resource language Marathi. We compare these embeddings with their contextual and compressed variants. Our findings indicate that non-contextual BERT embeddings extracted from the model's first embedding layer outperform FastText embeddings, presenting a promising alternative for low-resource NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17661
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Non-Contextual BERT or FastText? A Comparative Analysis
Shanbhag, Abhay
Jadhav, Suramya
Thakurdesai, Amogh
Sinare, Ridhima
Joshi, Raviraj
Computation and Language
Machine Learning
Natural Language Processing (NLP) for low-resource languages, which lack large annotated datasets, faces significant challenges due to limited high-quality data and linguistic resources. The selection of embeddings plays a critical role in achieving strong performance in NLP tasks. While contextual BERT embeddings require a full forward pass, non-contextual BERT embeddings rely only on table lookup. Existing research has primarily focused on contextual BERT embeddings, leaving non-contextual embeddings largely unexplored. In this study, we analyze the effectiveness of non-contextual embeddings from BERT models (MuRIL and MahaBERT) and FastText models (IndicFT and MahaFT) for tasks such as news classification, sentiment analysis, and hate speech detection in one such low-resource language Marathi. We compare these embeddings with their contextual and compressed variants. Our findings indicate that non-contextual BERT embeddings extracted from the model's first embedding layer outperform FastText embeddings, presenting a promising alternative for low-resource NLP.
title Non-Contextual BERT or FastText? A Comparative Analysis
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.17661