3CEL: A corpus of legal Spanish contract clauses

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: García, Nuria Aldama, Morales, Patricia Marsà, Sánchez, David Betancur, Jiménez, Álvaro Barbero, Nieto, Marta Guerrero, Coll, Pablo Haya, Chozas, Patricia Martín, Ponsoda, Elena Montiel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916585103425536
author García, Nuria Aldama
Morales, Patricia Marsà
Sánchez, David Betancur
Jiménez, Álvaro Barbero
Nieto, Marta Guerrero
Coll, Pablo Haya
Chozas, Patricia Martín
Ponsoda, Elena Montiel
author_facet García, Nuria Aldama
Morales, Patricia Marsà
Sánchez, David Betancur
Jiménez, Álvaro Barbero
Nieto, Marta Guerrero
Coll, Pablo Haya
Chozas, Patricia Martín
Ponsoda, Elena Montiel
contents Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Politécnica de Madrid (UPM) and developed by Instituto de Ingeniería del Conocimiento (IIC) to create a series of state-of-the-art NLP resources applied to the legal/administrative domain in Spanish. The goal of this paper is to present the Corpus of Legal Spanish Contract Clauses (3CEL), which is a contract information extraction corpus developed within the framework of INESData 2024. 3CEL contains 373 manually annotated tenders using 19 defined categories (4 782 total tags) that identify key information for contract understanding and reviewing.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15990
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3CEL: A corpus of legal Spanish contract clauses
García, Nuria Aldama
Morales, Patricia Marsà
Sánchez, David Betancur
Jiménez, Álvaro Barbero
Nieto, Marta Guerrero
Coll, Pablo Haya
Chozas, Patricia Martín
Ponsoda, Elena Montiel
Computation and Language
Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Politécnica de Madrid (UPM) and developed by Instituto de Ingeniería del Conocimiento (IIC) to create a series of state-of-the-art NLP resources applied to the legal/administrative domain in Spanish. The goal of this paper is to present the Corpus of Legal Spanish Contract Clauses (3CEL), which is a contract information extraction corpus developed within the framework of INESData 2024. 3CEL contains 373 manually annotated tenders using 19 defined categories (4 782 total tags) that identify key information for contract understanding and reviewing.
title 3CEL: A corpus of legal Spanish contract clauses
topic Computation and Language
url https://arxiv.org/abs/2501.15990