WikiNER-fr-gold: A Gold-Standard NER Corpus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Danrun, Béchet, Nicolas, Marteau, Pierre-François
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910919916781568
author Cao, Danrun
Béchet, Nicolas
Marteau, Pierre-François
author_facet Cao, Danrun
Béchet, Nicolas
Marteau, Pierre-François
contents We address in this article the the quality of the WikiNER corpus, a multilingual Named Entity Recognition corpus, and provide a consolidated version of it. The annotation of WikiNER was produced in a semi-supervised manner i.e. no manual verification has been carried out a posteriori. Such corpus is called silver-standard. In this paper we propose WikiNER-fr-gold which is a revised version of the French proportion of WikiNER. Our corpus consists of randomly sampled 20% of the original French sub-corpus (26,818 sentences with 700k tokens). We start by summarizing the entity types included in each category in order to define an annotation guideline, and then we proceed to revise the corpus. Finally we present an analysis of errors and inconsistency observed in the WikiNER-fr corpus, and we discuss potential future work directions.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00030
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WikiNER-fr-gold: A Gold-Standard NER Corpus
Cao, Danrun
Béchet, Nicolas
Marteau, Pierre-François
Computation and Language
Artificial Intelligence
Databases
We address in this article the the quality of the WikiNER corpus, a multilingual Named Entity Recognition corpus, and provide a consolidated version of it. The annotation of WikiNER was produced in a semi-supervised manner i.e. no manual verification has been carried out a posteriori. Such corpus is called silver-standard. In this paper we propose WikiNER-fr-gold which is a revised version of the French proportion of WikiNER. Our corpus consists of randomly sampled 20% of the original French sub-corpus (26,818 sentences with 700k tokens). We start by summarizing the entity types included in each category in order to define an annotation guideline, and then we proceed to revise the corpus. Finally we present an analysis of errors and inconsistency observed in the WikiNER-fr corpus, and we discuss potential future work directions.
title WikiNER-fr-gold: A Gold-Standard NER Corpus
topic Computation and Language
Artificial Intelligence
Databases
url https://arxiv.org/abs/2411.00030