Comparative Study of Pre-Trained BERT and Large Language Models for Code-Mixed Named Entity Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shirke, Mayur, Shembade, Amey, Thorat, Pavan, Wagh, Madhushri, Joshi, Raviraj
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912567114334208
author Shirke, Mayur
Shembade, Amey
Thorat, Pavan
Wagh, Madhushri
Joshi, Raviraj
author_facet Shirke, Mayur
Shembade, Amey
Thorat, Pavan
Wagh, Madhushri
Joshi, Raviraj
contents Named Entity Recognition (NER) in code-mixed text, particularly Hindi-English (Hinglish), presents unique challenges due to informal structure, transliteration, and frequent language switching. This study conducts a comparative evaluation of code-mixed fine-tuned models and non-code-mixed multilingual models, along with zero-shot generative large language models (LLMs). Specifically, we evaluate HingBERT, HingMBERT, and HingRoBERTa (trained on code-mixed data), and BERT Base Cased, IndicBERT, RoBERTa and MuRIL (trained on non-code-mixed multilingual data). We also assess the performance of Google Gemini in a zero-shot setting using a modified version of the dataset with NER tags removed. All models are tested on a benchmark Hinglish NER dataset using Precision, Recall, and F1-score. Results show that code-mixed models, particularly HingRoBERTa and HingBERT-based fine-tuned models, outperform others - including closed-source LLMs like Google Gemini - due to domain-specific pretraining. Non-code-mixed models perform reasonably but show limited adaptability. Notably, Google Gemini exhibits competitive zero-shot performance, underlining the generalization strength of modern LLMs. This study provides key insights into the effectiveness of specialized versus generalized models for code-mixed NER tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02514
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comparative Study of Pre-Trained BERT and Large Language Models for Code-Mixed Named Entity Recognition
Shirke, Mayur
Shembade, Amey
Thorat, Pavan
Wagh, Madhushri
Joshi, Raviraj
Computation and Language
Machine Learning
Named Entity Recognition (NER) in code-mixed text, particularly Hindi-English (Hinglish), presents unique challenges due to informal structure, transliteration, and frequent language switching. This study conducts a comparative evaluation of code-mixed fine-tuned models and non-code-mixed multilingual models, along with zero-shot generative large language models (LLMs). Specifically, we evaluate HingBERT, HingMBERT, and HingRoBERTa (trained on code-mixed data), and BERT Base Cased, IndicBERT, RoBERTa and MuRIL (trained on non-code-mixed multilingual data). We also assess the performance of Google Gemini in a zero-shot setting using a modified version of the dataset with NER tags removed. All models are tested on a benchmark Hinglish NER dataset using Precision, Recall, and F1-score. Results show that code-mixed models, particularly HingRoBERTa and HingBERT-based fine-tuned models, outperform others - including closed-source LLMs like Google Gemini - due to domain-specific pretraining. Non-code-mixed models perform reasonably but show limited adaptability. Notably, Google Gemini exhibits competitive zero-shot performance, underlining the generalization strength of modern LLMs. This study provides key insights into the effectiveness of specialized versus generalized models for code-mixed NER tasks.
title Comparative Study of Pre-Trained BERT and Large Language Models for Code-Mixed Named Entity Recognition
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.02514