Fine-tuning Pre-trained Named Entity Recognition Models For Indian Languages

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bahad, Sankalp, Mishra, Pruthwik, Arora, Karunesh, Balabantaray, Rakesh Chandra, Sharma, Dipti Misra, Krishnamurthy, Parameswari
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911872458948608
author Bahad, Sankalp
Mishra, Pruthwik
Arora, Karunesh
Balabantaray, Rakesh Chandra
Sharma, Dipti Misra
Krishnamurthy, Parameswari
author_facet Bahad, Sankalp
Mishra, Pruthwik
Arora, Karunesh
Balabantaray, Rakesh Chandra
Sharma, Dipti Misra
Krishnamurthy, Parameswari
contents Named Entity Recognition (NER) is a useful component in Natural Language Processing (NLP) applications. It is used in various tasks such as Machine Translation, Summarization, Information Retrieval, and Question-Answering systems. The research on NER is centered around English and some other major languages, whereas limited attention has been given to Indian languages. We analyze the challenges and propose techniques that can be tailored for Multilingual Named Entity Recognition for Indian Languages. We present a human annotated named entity corpora of 40K sentences for 4 Indian languages from two of the major Indian language families. Additionally,we present a multilingual model fine-tuned on our dataset, which achieves an F1 score of 0.80 on our dataset on average. We achieve comparable performance on completely unseen benchmark datasets for Indian languages which affirms the usability of our model.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04829
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-tuning Pre-trained Named Entity Recognition Models For Indian Languages
Bahad, Sankalp
Mishra, Pruthwik
Arora, Karunesh
Balabantaray, Rakesh Chandra
Sharma, Dipti Misra
Krishnamurthy, Parameswari
Computation and Language
Named Entity Recognition (NER) is a useful component in Natural Language Processing (NLP) applications. It is used in various tasks such as Machine Translation, Summarization, Information Retrieval, and Question-Answering systems. The research on NER is centered around English and some other major languages, whereas limited attention has been given to Indian languages. We analyze the challenges and propose techniques that can be tailored for Multilingual Named Entity Recognition for Indian Languages. We present a human annotated named entity corpora of 40K sentences for 4 Indian languages from two of the major Indian language families. Additionally,we present a multilingual model fine-tuned on our dataset, which achieves an F1 score of 0.80 on our dataset on average. We achieve comparable performance on completely unseen benchmark datasets for Indian languages which affirms the usability of our model.
title Fine-tuning Pre-trained Named Entity Recognition Models For Indian Languages
topic Computation and Language
url https://arxiv.org/abs/2405.04829