Enhancing NER Performance in Low-Resource Pakistani Languages using Cross-Lingual Data Augmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ehsan, Toqeer, Solorio, Thamar
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916686036205568
author Ehsan, Toqeer
Solorio, Thamar
author_facet Ehsan, Toqeer
Solorio, Thamar
contents Named Entity Recognition (NER), a fundamental task in Natural Language Processing (NLP), has shown significant advancements for high-resource languages. However, due to a lack of annotated datasets and limited representation in Pre-trained Language Models (PLMs), it remains understudied and challenging for low-resource languages. To address these challenges, we propose a data augmentation technique that generates culturally plausible sentences and experiments on four low-resource Pakistani languages; Urdu, Shahmukhi, Sindhi, and Pashto. By fine-tuning multilingual masked Large Language Models (LLMs), our approach demonstrates significant improvements in NER performance for Shahmukhi and Pashto. We further explore the capability of generative LLMs for NER and data augmentation using few-shot learning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08792
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing NER Performance in Low-Resource Pakistani Languages using Cross-Lingual Data Augmentation
Ehsan, Toqeer
Solorio, Thamar
Computation and Language
Information Retrieval
Named Entity Recognition (NER), a fundamental task in Natural Language Processing (NLP), has shown significant advancements for high-resource languages. However, due to a lack of annotated datasets and limited representation in Pre-trained Language Models (PLMs), it remains understudied and challenging for low-resource languages. To address these challenges, we propose a data augmentation technique that generates culturally plausible sentences and experiments on four low-resource Pakistani languages; Urdu, Shahmukhi, Sindhi, and Pashto. By fine-tuning multilingual masked Large Language Models (LLMs), our approach demonstrates significant improvements in NER performance for Shahmukhi and Pashto. We further explore the capability of generative LLMs for NER and data augmentation using few-shot learning.
title Enhancing NER Performance in Low-Resource Pakistani Languages using Cross-Lingual Data Augmentation
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2504.08792