From LIMA to DeepLIMA: following a new path of interoperability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bocharov, Victor, Besançon, Romaric, de Chalendar, Gaël, Ferret, Olivier, Semmar, Nasredine
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929494675161088
author Bocharov, Victor
Besançon, Romaric
de Chalendar, Gaël
Ferret, Olivier
Semmar, Nasredine
author_facet Bocharov, Victor
Besançon, Romaric
de Chalendar, Gaël
Ferret, Olivier
Semmar, Nasredine
contents In this article, we describe the architecture of the LIMA (Libre Multilingual Analyzer) framework and its recent evolution with the addition of new text analysis modules based on deep neural networks. We extended the functionality of LIMA in terms of the number of supported languages while preserving existing configurable architecture and the availability of previously developed rule-based and statistical analysis components. Models were trained for more than 60 languages on the Universal Dependencies 2.5 corpora, WikiNer corpora, and CoNLL-03 dataset. Universal Dependencies allowed us to increase the number of supported languages and to generate models that could be integrated into other platforms. This integration of ubiquitous Deep Learning Natural Language Processing models and the use of standard annotated collections using Universal Dependencies can be viewed as a new path of interoperability, through the normalization of models and data, that are complementary to a more standard technical interoperability, implemented in LIMA through services available in Docker containers on Docker Hub.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06550
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From LIMA to DeepLIMA: following a new path of interoperability
Bocharov, Victor
Besançon, Romaric
de Chalendar, Gaël
Ferret, Olivier
Semmar, Nasredine
Computation and Language
In this article, we describe the architecture of the LIMA (Libre Multilingual Analyzer) framework and its recent evolution with the addition of new text analysis modules based on deep neural networks. We extended the functionality of LIMA in terms of the number of supported languages while preserving existing configurable architecture and the availability of previously developed rule-based and statistical analysis components. Models were trained for more than 60 languages on the Universal Dependencies 2.5 corpora, WikiNer corpora, and CoNLL-03 dataset. Universal Dependencies allowed us to increase the number of supported languages and to generate models that could be integrated into other platforms. This integration of ubiquitous Deep Learning Natural Language Processing models and the use of standard annotated collections using Universal Dependencies can be viewed as a new path of interoperability, through the normalization of models and data, that are complementary to a more standard technical interoperability, implemented in LIMA through services available in Docker containers on Docker Hub.
title From LIMA to DeepLIMA: following a new path of interoperability
topic Computation and Language
url https://arxiv.org/abs/2409.06550