An Experimental Study on Data Augmentation Techniques for Named Entity Recognition on Low-Resource Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Torres, Arthur Elwing, de Moura, Edleno Silva, da Silva, Altigran Soares, Nascimento, Mario A., Mesquita, Filipe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915901821943808
author Torres, Arthur Elwing
de Moura, Edleno Silva
da Silva, Altigran Soares
Nascimento, Mario A.
Mesquita, Filipe
author_facet Torres, Arthur Elwing
de Moura, Edleno Silva
da Silva, Altigran Soares
Nascimento, Mario A.
Mesquita, Filipe
contents Named Entity Recognition (NER) is a machine learning task that traditionally relies on supervised learning and annotated data. Acquiring such data is often a challenge, particularly in specialized fields like medical, legal, and financial sectors. Those are commonly referred to as low-resource domains, which comprise long-tail entities, due to the scarcity of available data. To address this, data augmentation techniques are increasingly being employed to generate additional training instances from the original dataset. In this study, we evaluate the effectiveness of two prominent text augmentation techniques, Mention Replacement and Contextual Word Replacement, on two widely-used NER models, Bi-LSTM+CRF and BERT. We conduct experiments on four datasets from low-resource domains, and we explore the impact of various combinations of training subset sizes and number of augmented examples. We not only confirm that data augmentation is particularly beneficial for smaller datasets, but we also demonstrate that there is no universally optimal number of augmented examples, i.e., NER practitioners must experiment with different quantities in order to fine-tune their projects.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14551
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Experimental Study on Data Augmentation Techniques for Named Entity Recognition on Low-Resource Domains
Torres, Arthur Elwing
de Moura, Edleno Silva
da Silva, Altigran Soares
Nascimento, Mario A.
Mesquita, Filipe
Computation and Language
Information Retrieval
Machine Learning
Named Entity Recognition (NER) is a machine learning task that traditionally relies on supervised learning and annotated data. Acquiring such data is often a challenge, particularly in specialized fields like medical, legal, and financial sectors. Those are commonly referred to as low-resource domains, which comprise long-tail entities, due to the scarcity of available data. To address this, data augmentation techniques are increasingly being employed to generate additional training instances from the original dataset. In this study, we evaluate the effectiveness of two prominent text augmentation techniques, Mention Replacement and Contextual Word Replacement, on two widely-used NER models, Bi-LSTM+CRF and BERT. We conduct experiments on four datasets from low-resource domains, and we explore the impact of various combinations of training subset sizes and number of augmented examples. We not only confirm that data augmentation is particularly beneficial for smaller datasets, but we also demonstrate that there is no universally optimal number of augmented examples, i.e., NER practitioners must experiment with different quantities in order to fine-tune their projects.
title An Experimental Study on Data Augmentation Techniques for Named Entity Recognition on Low-Resource Domains
topic Computation and Language
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2411.14551