LiMe: a Latin Corpus of Late Medieval Criminal Sentences

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bassani, Alessandra, Del Bo, Beatrice, Ferrara, Alfio, Mangini, Marta, Picascia, Sergio, Stefanello, Ambra
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916871606894592
author Bassani, Alessandra
Del Bo, Beatrice
Ferrara, Alfio
Mangini, Marta
Picascia, Sergio
Stefanello, Ambra
author_facet Bassani, Alessandra
Del Bo, Beatrice
Ferrara, Alfio
Mangini, Marta
Picascia, Sergio
Stefanello, Ambra
contents The Latin language has received attention from the computational linguistics research community, which has built, over the years, several valuable resources, ranging from detailed annotated corpora to sophisticated tools for linguistic analysis. With the recent advent of large language models, researchers have also started developing models capable of generating vector representations of Latin texts. The performances of such models remain behind the ones for modern languages, given the disparity in available data. In this paper, we present the LiMe dataset, a corpus of 325 documents extracted from a series of medieval manuscripts called Libri sententiarum potestatis Mediolani, and thoroughly annotated by experts, in order to be employed for masked language model, as well as supervised natural language processing tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2404_12829
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LiMe: a Latin Corpus of Late Medieval Criminal Sentences
Bassani, Alessandra
Del Bo, Beatrice
Ferrara, Alfio
Mangini, Marta
Picascia, Sergio
Stefanello, Ambra
Computation and Language
The Latin language has received attention from the computational linguistics research community, which has built, over the years, several valuable resources, ranging from detailed annotated corpora to sophisticated tools for linguistic analysis. With the recent advent of large language models, researchers have also started developing models capable of generating vector representations of Latin texts. The performances of such models remain behind the ones for modern languages, given the disparity in available data. In this paper, we present the LiMe dataset, a corpus of 325 documents extracted from a series of medieval manuscripts called Libri sententiarum potestatis Mediolani, and thoroughly annotated by experts, in order to be employed for masked language model, as well as supervised natural language processing tasks.
title LiMe: a Latin Corpus of Late Medieval Criminal Sentences
topic Computation and Language
url https://arxiv.org/abs/2404.12829