FRASIMED: a Clinical French Annotated Resource Produced through Crosslingual BERT-Based Annotation Projection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zaghir, Jamil, Bjelogrlic, Mina, Goldman, Jean-Philippe, Aananou, Soukaïna, Gaudet-Blavignac, Christophe, Lovis, Christian
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912491000299520
author Zaghir, Jamil
Bjelogrlic, Mina
Goldman, Jean-Philippe
Aananou, Soukaïna
Gaudet-Blavignac, Christophe
Lovis, Christian
author_facet Zaghir, Jamil
Bjelogrlic, Mina
Goldman, Jean-Philippe
Aananou, Soukaïna
Gaudet-Blavignac, Christophe
Lovis, Christian
contents Natural language processing (NLP) applications such as named entity recognition (NER) for low-resource corpora do not benefit from recent advances in the development of large language models (LLMs) where there is still a need for larger annotated datasets. This research article introduces a methodology for generating translated versions of annotated datasets through crosslingual annotation projection. Leveraging a language agnostic BERT-based approach, it is an efficient solution to increase low-resource corpora with few human efforts and by only using already available open data resources. Quantitative and qualitative evaluations are often lacking when it comes to evaluating the quality and effectiveness of semi-automatic data generation strategies. The evaluation of our crosslingual annotation projection approach showed both effectiveness and high accuracy in the resulting dataset. As a practical application of this methodology, we present the creation of French Annotated Resource with Semantic Information for Medical Entities Detection (FRASIMED), an annotated corpus comprising 2'051 synthetic clinical cases in French. The corpus is now available for researchers and practitioners to develop and refine French natural language processing (NLP) applications in the clinical field (https://zenodo.org/record/8355629), making it the largest open annotated corpus with linked medical concepts in French.
format Preprint
id arxiv_https___arxiv_org_abs_2309_10770
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FRASIMED: a Clinical French Annotated Resource Produced through Crosslingual BERT-Based Annotation Projection
Zaghir, Jamil
Bjelogrlic, Mina
Goldman, Jean-Philippe
Aananou, Soukaïna
Gaudet-Blavignac, Christophe
Lovis, Christian
Computation and Language
Artificial Intelligence
Natural language processing (NLP) applications such as named entity recognition (NER) for low-resource corpora do not benefit from recent advances in the development of large language models (LLMs) where there is still a need for larger annotated datasets. This research article introduces a methodology for generating translated versions of annotated datasets through crosslingual annotation projection. Leveraging a language agnostic BERT-based approach, it is an efficient solution to increase low-resource corpora with few human efforts and by only using already available open data resources. Quantitative and qualitative evaluations are often lacking when it comes to evaluating the quality and effectiveness of semi-automatic data generation strategies. The evaluation of our crosslingual annotation projection approach showed both effectiveness and high accuracy in the resulting dataset. As a practical application of this methodology, we present the creation of French Annotated Resource with Semantic Information for Medical Entities Detection (FRASIMED), an annotated corpus comprising 2'051 synthetic clinical cases in French. The corpus is now available for researchers and practitioners to develop and refine French natural language processing (NLP) applications in the clinical field (https://zenodo.org/record/8355629), making it the largest open annotated corpus with linked medical concepts in French.
title FRASIMED: a Clinical French Annotated Resource Produced through Crosslingual BERT-Based Annotation Projection
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2309.10770