Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Amalvy, Arthur, Labatut, Vincent, Dufour, Richard
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911829336260608
author Amalvy, Arthur
Labatut, Vincent
Dufour, Richard
author_facet Amalvy, Arthur
Labatut, Vincent
Dufour, Richard
contents While recent pre-trained transformer-based models can perform named entity recognition (NER) with great accuracy, their limited range remains an issue when applied to long documents such as whole novels. To alleviate this issue, a solution is to retrieve relevant context at the document level. Unfortunately, the lack of supervision for such a task means one has to settle for unsupervised approaches. Instead, we propose to generate a synthetic context retrieval training dataset using Alpaca, an instructiontuned large language model (LLM). Using this dataset, we train a neural context retriever based on a BERT model that is able to find relevant context for NER. We show that our method outperforms several retrieval baselines for the NER task on an English literary dataset composed of the first chapter of 40 books.
format Preprint
id arxiv_https___arxiv_org_abs_2310_10118
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset
Amalvy, Arthur
Labatut, Vincent
Dufour, Richard
Computation and Language
While recent pre-trained transformer-based models can perform named entity recognition (NER) with great accuracy, their limited range remains an issue when applied to long documents such as whole novels. To alleviate this issue, a solution is to retrieve relevant context at the document level. Unfortunately, the lack of supervision for such a task means one has to settle for unsupervised approaches. Instead, we propose to generate a synthetic context retrieval training dataset using Alpaca, an instructiontuned large language model (LLM). Using this dataset, we train a neural context retriever based on a BERT model that is able to find relevant context for NER. We show that our method outperforms several retrieval baselines for the NER task on an English literary dataset composed of the first chapter of 40 books.
title Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset
topic Computation and Language
url https://arxiv.org/abs/2310.10118