3CEL: A corpus of legal Spanish contract clauses
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916585103425536 |
|---|---|
| author | García, Nuria Aldama Morales, Patricia Marsà Sánchez, David Betancur Jiménez, Álvaro Barbero Nieto, Marta Guerrero Coll, Pablo Haya Chozas, Patricia Martín Ponsoda, Elena Montiel |
| author_facet | García, Nuria Aldama Morales, Patricia Marsà Sánchez, David Betancur Jiménez, Álvaro Barbero Nieto, Marta Guerrero Coll, Pablo Haya Chozas, Patricia Martín Ponsoda, Elena Montiel |
| contents | Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Politécnica de Madrid (UPM) and developed by Instituto de Ingeniería del Conocimiento (IIC) to create a series of state-of-the-art NLP resources applied to the legal/administrative domain in Spanish. The goal of this paper is to present the Corpus of Legal Spanish Contract Clauses (3CEL), which is a contract information extraction corpus developed within the framework of INESData 2024. 3CEL contains 373 manually annotated tenders using 19 defined categories (4 782 total tags) that identify key information for contract understanding and reviewing. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_15990 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | 3CEL: A corpus of legal Spanish contract clauses García, Nuria Aldama Morales, Patricia Marsà Sánchez, David Betancur Jiménez, Álvaro Barbero Nieto, Marta Guerrero Coll, Pablo Haya Chozas, Patricia Martín Ponsoda, Elena Montiel Computation and Language Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Politécnica de Madrid (UPM) and developed by Instituto de Ingeniería del Conocimiento (IIC) to create a series of state-of-the-art NLP resources applied to the legal/administrative domain in Spanish. The goal of this paper is to present the Corpus of Legal Spanish Contract Clauses (3CEL), which is a contract information extraction corpus developed within the framework of INESData 2024. 3CEL contains 373 manually annotated tenders using 19 defined categories (4 782 total tags) that identify key information for contract understanding and reviewing. |
| title | 3CEL: A corpus of legal Spanish contract clauses |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2501.15990 |