ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Portela, Johan R., Perez, Nicolás, Manrique, Rubén
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912270176485376
author Portela, Johan R.
Perez, Nicolás
Manrique, Rubén
author_facet Portela, Johan R.
Perez, Nicolás
Manrique, Rubén
contents Natural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), serves as a crucial area within the domain of Natural Language Processing (NLP). This area fundamentally empowers machines to discern semantic relationships between assorted sections of text. Even though considerable work has been executed for the English language, it has been observed that efforts for the Spanish language are relatively sparse. Keeping this in view, this paper focuses on generating a multi-genre Spanish dataset for NLI, ESNLIR, particularly accounting for causal Relationships. A preliminary baseline has been conceptualized and subjected to an evaluation, leveraging models drawn from the BERT family. The findings signify that the enrichment of genres essentially contributes to the enrichment of the model's capability to generalize. The code, notebooks and whole datasets for this experiments is available at: https://zenodo.org/records/15002575. If you are interested only in the dataset you can find it here: https://zenodo.org/records/15002371.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships
Portela, Johan R.
Perez, Nicolás
Manrique, Rubén
Computation and Language
Natural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), serves as a crucial area within the domain of Natural Language Processing (NLP). This area fundamentally empowers machines to discern semantic relationships between assorted sections of text. Even though considerable work has been executed for the English language, it has been observed that efforts for the Spanish language are relatively sparse. Keeping this in view, this paper focuses on generating a multi-genre Spanish dataset for NLI, ESNLIR, particularly accounting for causal Relationships. A preliminary baseline has been conceptualized and subjected to an evaluation, leveraging models drawn from the BERT family. The findings signify that the enrichment of genres essentially contributes to the enrichment of the model's capability to generalize. The code, notebooks and whole datasets for this experiments is available at: https://zenodo.org/records/15002575. If you are interested only in the dataset you can find it here: https://zenodo.org/records/15002371.
title ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships
topic Computation and Language
url https://arxiv.org/abs/2503.08803