A Hybrid Method for Low-Resource Named Entity Recognition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Duc, Do Minh, Truong, Quan Xuan, Hong, Viet Tran, Anh, Le Hoang, Tra, Mac Thi Minh, Van Thuy, Nguyen, Ha, Le Hai, Van, Vinh Nguyen
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918485914812416
author Duc, Do Minh
Truong, Quan Xuan
Hong, Viet Tran
Anh, Le Hoang
Tra, Mac Thi Minh
Van Thuy, Nguyen
Ha, Le Hai
Van, Vinh Nguyen
author_facet Duc, Do Minh
Truong, Quan Xuan
Hong, Viet Tran
Anh, Le Hoang
Tra, Mac Thi Minh
Van Thuy, Nguyen
Ha, Le Hai
Van, Vinh Nguyen
contents Named Entity Recognition (NER) is a critical component of Natural Language Processing with diverse applications in information extraction and conversational AI. However, NER in specific domains for low-resource languages faces challenges such as limited annotated data and heterogeneous label sets. This study addresses these issues by proposing a hybrid neurosymbolic framework that integrates rule-based processing with deep learning models for Vietnamese NER. The core idea involves a two-stage pipeline: first, a rule-based component reduces label complexity by grouping relational and special categories; second, pre-trained language models are fine-tuned for high-precision extraction. A post-processing module is then utilized to restore fine-grained labels, preserving expressiveness for application-level usability. To mitigate data scarcity, a scalable data augmentation strategy leveraging Large Language Models (LLMs) is introduced to expand the label set without full re-annotation, which is a significant novelty of this work. The effectiveness of this method was evaluated across five specific-domain datasets, including logistics, wildlife, and healthcare. Experimental results demonstrate substantial improvements over strong RoBERTa-based baselines. Specifically, the proposed system achieved F1 scores of 90 percent in Customer Service, up from 83 percent; 84 percent in GAM, up from 73 percent; 83 percent in AI Fluent, up from 80 percent; 94 percent in PhoNER_Covid19, up from 91 percent; and 60 percent in Rare Wildlife, up from 36 percent. These findings confirm that the hybrid approach effectively captures the linguistic complexity of Vietnamese and contextual nuances in specialized domains, offering a robust contribution to low-resource NER research.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04489
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Hybrid Method for Low-Resource Named Entity Recognition
Duc, Do Minh
Truong, Quan Xuan
Hong, Viet Tran
Anh, Le Hoang
Tra, Mac Thi Minh
Van Thuy, Nguyen
Ha, Le Hai
Van, Vinh Nguyen
Computational Engineering, Finance, and Science
Artificial Intelligence
Computation and Language
Named Entity Recognition (NER) is a critical component of Natural Language Processing with diverse applications in information extraction and conversational AI. However, NER in specific domains for low-resource languages faces challenges such as limited annotated data and heterogeneous label sets. This study addresses these issues by proposing a hybrid neurosymbolic framework that integrates rule-based processing with deep learning models for Vietnamese NER. The core idea involves a two-stage pipeline: first, a rule-based component reduces label complexity by grouping relational and special categories; second, pre-trained language models are fine-tuned for high-precision extraction. A post-processing module is then utilized to restore fine-grained labels, preserving expressiveness for application-level usability. To mitigate data scarcity, a scalable data augmentation strategy leveraging Large Language Models (LLMs) is introduced to expand the label set without full re-annotation, which is a significant novelty of this work. The effectiveness of this method was evaluated across five specific-domain datasets, including logistics, wildlife, and healthcare. Experimental results demonstrate substantial improvements over strong RoBERTa-based baselines. Specifically, the proposed system achieved F1 scores of 90 percent in Customer Service, up from 83 percent; 84 percent in GAM, up from 73 percent; 83 percent in AI Fluent, up from 80 percent; 94 percent in PhoNER_Covid19, up from 91 percent; and 60 percent in Rare Wildlife, up from 36 percent. These findings confirm that the hybrid approach effectively captures the linguistic complexity of Vietnamese and contextual nuances in specialized domains, offering a robust contribution to low-resource NER research.
title A Hybrid Method for Low-Resource Named Entity Recognition
topic Computational Engineering, Finance, and Science
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.04489