| _version_ | 1866902210306113536 |
|---|---|
| author | Rodriguez Portela, Johan David Manrique Piramanrique, Rubén Francisco Perez Terán, Nicolás |
| author_facet | Rodriguez Portela, Johan David Manrique Piramanrique, Rubén Francisco Perez Terán, Nicolás |
| contents | <h1>ESNLIR: Expanding Spanish NLI Benchmarks with Multi-Genre and Causal Annotation</h1> <p>This is the complete code, model and datasets for the article <a href="https://link.springer.com/chapter/10.1007/978-3-032-07175-0_23">ESNLIR: Expanding Spanish NLI Benchmarks with Multi-genre and Causal Annotation</a></p> <p>In case you cannot access the article this preprint is available: <a href="https://arxiv.org/abs/2503.08803">ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships</a>.</p> <h2>How to cite:</h2> <p>Portela, J.R., Pérez-Terán, N., Manrique, R. (2026). ESNLIR: Expanding Spanish NLI Benchmarks with Multi-genre and Causal Annotation. In: Florez, H., Peluffo-Ordoñez, D. (eds) Applied Informatics. ICAI 2025. Communications in Computer and Information Science, vol 2667. Springer, Cham. https://doi.org/10.1007/978-3-032-07175-0_23</p> <h1>IMPORTANT UPDATE!!!</h1> <h2>It is strongly advised to work with the following links, instead of working directly from Zenodo:</h2> <ul> <li> <h2><a href="https://github.com/jd-rodriguezp1234/esnlir">CODE REPOSITORY:</a> This repository contains the code used for the article.</h2> </li> <li> <h2><a href="https://github.com/jd-rodriguezp1234/esnlir-training-example">SMALL EXAMPLE REPOSITORY</a>: This repository contains a small code example showing you how to train, and predict using a very small toy dataset, with the same structure.</h2> </li> <li> <h2><a href="https://huggingface.co/collections/Flaglab/esnlir">HUGGING FACE COLLECTION</a>: Huggingface collection containing the dataset and models.</h2> </li> </ul> <p> </p> <p>If you still want to use the Zenodo repository, follow the steps below. But once again, it is way easier to work with the links above.</p> <p>----------------------------------------------------------------------------------------------</p> <h1>Installation</h1> <p>This repository is a poetry project, which means that it can be installed easily by executing the following command from a shell in the repository folder:</p> <pre><code>poetry install</code></pre> <p>As this repository is script based, the README.md file contains all the commands executed to generate the dataset and train models.</p> <p>----------------------------------------------------------------------------------------------</p> <h1>Core code</h1> <p>The core code used for all the experiments is in the folder <strong>auto-nli </strong>and all the calls to the core code with the parameters requested are found in README.md</p> <p>----------------------------------------------------------------------------------------------</p> <h1>Parameters</h1> <p>All the parameters to create datasets and train models with the core code are found in the folder <strong>parameters.</strong></p> <p>----------------------------------------------------------------------------------------------</p> <h1>Models</h1> <h2>Model types</h2> <p>For BERT based models, all in pytorch, there are two types of models from huggingfaces that were used for training and also are required to load a dataset because of the tokenizer:</p> <ul> <li>RoBERTa (BERTIN): https://huggingface.co/<strong>bertin-project/bertin-roberta-base-spanish</strong></li> <li>XLMRoBERTa: https://huggingface.co/<strong>FacebookAI/xlm-roberta-base</strong></li> </ul> <h2>Model folder</h2> <p>The <strong>model </strong>folder contains all the trained models for the paper. There are three types of models:</p> <ul> <li>baseline: An XGBoost model that can be loaded with pickle.</li> <li>roberta: BERTIN based models in pytorch. You can load them with the model_path <strong><project_path>/model/roberta/model</strong></li> <li>xlmroberta: XLMRoBERTa based models in pytorch. You can load them with the model_path <strong><project_path>/model/xlmroberta/model</strong></li> </ul> <p>Models with the suffix <strong>_annot </strong> are models trained with the premise (first sentence) only. Apart from the pytorch <strong>model</strong> folder, each model result folder (ex: <strong><project_path>/model/xlmroberta/</strong>) contains the test results for the test set and the stress test sets (ex: <strong><project_path>/model/xlmroberta/test</strong>)</p> <h2>Load model</h2> <p>Models are found in the folder <strong>model </strong>and all of them are pytorch models which can be loaded with the huggingface interface:</p> <pre><code>from transformers import AutoModel model = AutoModel.from_pretrained('<model_path>',local_files_only=True)</code><br><br></pre> <p>----------------------------------------------------------------------------------------------</p> <h1>Dataset</h1> <h2>labeled_final_dataset.jsonl</h2> <p>This file is included outside the ZIP containing all other files, and it contains the final test dataset with 974 examples selected by human majority label matching the original linking phrase label.</p> <h2>Other datasets:</h2> <p>The datasets can be found in the folder <strong>data</strong> that is divided in the following folders:</p> <h3>base_dataset</h3> <p>The splits to train, validate and test the models.</p> <h3>splits_data</h3> <p>Splits of train-val-test extracted for each corpora. They are used to generate base_dataset.</p> <h3>sentence_data</h3> <p>Pairs of sentences found in each corpus. They are used to generate splits_data.</p> <h2>Dataset dictionary</h2> <p>This repository contains the splits that resulted from the research project "ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships". All the splits are in JSONL format and have the same fields per example:</p> <ul> <li>sentence_1: First sentence of the pair.</li> <li>sentence_2: Second sentence of the pair.</li> <li>connector: Linking phrase used to extract pair.</li> <li>connector_type: NLI label, between "contrasting", "entailment", "reasoning" or "neutral"</li> <li>extraction_strategy: "linking_phrase" for "contrasting", "entailment", "reasoning" and "none" for neutral.</li> <li>distance: How many sentences before the connector is the sentence_1</li> <li>sentence_1_position: Number of sentence for sentence_1 in the source document</li> <li>sentence_1_paragraph: Number of paragraph for sentence_1 in the source document</li> <li>sentence_2_position: Number of sentence for sentence_2 in the source document</li> <li>sentence_2_paragraph: Number of paragraph for sentence_2 in the source document</li> <li>id: Unique identifier for the example</li> <li>dataset: Source corpus of the pair. Metadata of corpus, including source can be found in dataset_metadata.xlsx.</li> <li>genre: Writing genre of the dataset.</li> <li>domain: Domain genre of the dataset. </li> </ul> <p>Example: </p> <p><code>{"sentence_1":"sefior Bcajavides no es moderado, tampoco lo convertirse e\u00f1 declarada divergencia de miras polileido en griego","sentence_2":"era mayor claricomentarios, as\u00ed de los peri\u00f3dicos como de los homes dado \u00e1 la voluntad de los hombres, sin que sobreticas","connector":"por consiguiente,","connector_type":"reasoning","extraction_strategy":"linking_phrase","distance":1.0,"sentence_1_paragraph":4,"sentence_1_position":86,"sentence_2_paragraph":4,"sentence_2_position":87,"id":"esnews__spanish_pd_news__531537","dataset":"esnews__spanish_pd_news","genre":"news","domain":"spanish_public_domain_news"}</code></p> <h2>Dataset load</h2> <p>To load a dataset/split as a pytorch object used to train-validate-test models you must use the custom class dataset </p> <div> <div><code>from auto_nli.model.bert_based.dataset import BERTDataset<br></code></div> <div> </div> <div><code>dataset = BERTDataset(</code></div> <div><code>os.path.join(dataset_folder, <Path to jsonl>),</code></div> <div><code>max_len=<max length of sentences>,</code></div> <div> <div> <div><code>model_type=</code></div> </div> <code><type of model to use for tokenizer>,</code></div> <div><code>only_premise=<True to load a dataset with only the first sentences>,</code></div> <div><code>max_samples=<Maximum number of examples in the dataset>)</code></div> <div> </div> <p>----------------------------------------------------------------------------------------------</p> <h1>Notebooks</h1> <div>The folder notebooks contains a collection of jupyter notebooks used to preprocess datasets and visualize results.</div> </div> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_15002575 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Complete code and datasets for "ESNLIR: Expanding Spanish NLI Benchmarks with Multi-Genre and Causal Annotation" Rodriguez Portela, Johan David Manrique Piramanrique, Rubén Francisco Perez Terán, Nicolás <h1>ESNLIR: Expanding Spanish NLI Benchmarks with Multi-Genre and Causal Annotation</h1> <p>This is the complete code, model and datasets for the article <a href="https://link.springer.com/chapter/10.1007/978-3-032-07175-0_23">ESNLIR: Expanding Spanish NLI Benchmarks with Multi-genre and Causal Annotation</a></p> <p>In case you cannot access the article this preprint is available: <a href="https://arxiv.org/abs/2503.08803">ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships</a>.</p> <h2>How to cite:</h2> <p>Portela, J.R., Pérez-Terán, N., Manrique, R. (2026). ESNLIR: Expanding Spanish NLI Benchmarks with Multi-genre and Causal Annotation. In: Florez, H., Peluffo-Ordoñez, D. (eds) Applied Informatics. ICAI 2025. Communications in Computer and Information Science, vol 2667. Springer, Cham. https://doi.org/10.1007/978-3-032-07175-0_23</p> <h1>IMPORTANT UPDATE!!!</h1> <h2>It is strongly advised to work with the following links, instead of working directly from Zenodo:</h2> <ul> <li> <h2><a href="https://github.com/jd-rodriguezp1234/esnlir">CODE REPOSITORY:</a> This repository contains the code used for the article.</h2> </li> <li> <h2><a href="https://github.com/jd-rodriguezp1234/esnlir-training-example">SMALL EXAMPLE REPOSITORY</a>: This repository contains a small code example showing you how to train, and predict using a very small toy dataset, with the same structure.</h2> </li> <li> <h2><a href="https://huggingface.co/collections/Flaglab/esnlir">HUGGING FACE COLLECTION</a>: Huggingface collection containing the dataset and models.</h2> </li> </ul> <p> </p> <p>If you still want to use the Zenodo repository, follow the steps below. But once again, it is way easier to work with the links above.</p> <p>----------------------------------------------------------------------------------------------</p> <h1>Installation</h1> <p>This repository is a poetry project, which means that it can be installed easily by executing the following command from a shell in the repository folder:</p> <pre><code>poetry install</code></pre> <p>As this repository is script based, the README.md file contains all the commands executed to generate the dataset and train models.</p> <p>----------------------------------------------------------------------------------------------</p> <h1>Core code</h1> <p>The core code used for all the experiments is in the folder <strong>auto-nli </strong>and all the calls to the core code with the parameters requested are found in README.md</p> <p>----------------------------------------------------------------------------------------------</p> <h1>Parameters</h1> <p>All the parameters to create datasets and train models with the core code are found in the folder <strong>parameters.</strong></p> <p>----------------------------------------------------------------------------------------------</p> <h1>Models</h1> <h2>Model types</h2> <p>For BERT based models, all in pytorch, there are two types of models from huggingfaces that were used for training and also are required to load a dataset because of the tokenizer:</p> <ul> <li>RoBERTa (BERTIN): https://huggingface.co/<strong>bertin-project/bertin-roberta-base-spanish</strong></li> <li>XLMRoBERTa: https://huggingface.co/<strong>FacebookAI/xlm-roberta-base</strong></li> </ul> <h2>Model folder</h2> <p>The <strong>model </strong>folder contains all the trained models for the paper. There are three types of models:</p> <ul> <li>baseline: An XGBoost model that can be loaded with pickle.</li> <li>roberta: BERTIN based models in pytorch. You can load them with the model_path <strong><project_path>/model/roberta/model</strong></li> <li>xlmroberta: XLMRoBERTa based models in pytorch. You can load them with the model_path <strong><project_path>/model/xlmroberta/model</strong></li> </ul> <p>Models with the suffix <strong>_annot </strong> are models trained with the premise (first sentence) only. Apart from the pytorch <strong>model</strong> folder, each model result folder (ex: <strong><project_path>/model/xlmroberta/</strong>) contains the test results for the test set and the stress test sets (ex: <strong><project_path>/model/xlmroberta/test</strong>)</p> <h2>Load model</h2> <p>Models are found in the folder <strong>model </strong>and all of them are pytorch models which can be loaded with the huggingface interface:</p> <pre><code>from transformers import AutoModel model = AutoModel.from_pretrained('<model_path>',local_files_only=True)</code><br><br></pre> <p>----------------------------------------------------------------------------------------------</p> <h1>Dataset</h1> <h2>labeled_final_dataset.jsonl</h2> <p>This file is included outside the ZIP containing all other files, and it contains the final test dataset with 974 examples selected by human majority label matching the original linking phrase label.</p> <h2>Other datasets:</h2> <p>The datasets can be found in the folder <strong>data</strong> that is divided in the following folders:</p> <h3>base_dataset</h3> <p>The splits to train, validate and test the models.</p> <h3>splits_data</h3> <p>Splits of train-val-test extracted for each corpora. They are used to generate base_dataset.</p> <h3>sentence_data</h3> <p>Pairs of sentences found in each corpus. They are used to generate splits_data.</p> <h2>Dataset dictionary</h2> <p>This repository contains the splits that resulted from the research project "ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships". All the splits are in JSONL format and have the same fields per example:</p> <ul> <li>sentence_1: First sentence of the pair.</li> <li>sentence_2: Second sentence of the pair.</li> <li>connector: Linking phrase used to extract pair.</li> <li>connector_type: NLI label, between "contrasting", "entailment", "reasoning" or "neutral"</li> <li>extraction_strategy: "linking_phrase" for "contrasting", "entailment", "reasoning" and "none" for neutral.</li> <li>distance: How many sentences before the connector is the sentence_1</li> <li>sentence_1_position: Number of sentence for sentence_1 in the source document</li> <li>sentence_1_paragraph: Number of paragraph for sentence_1 in the source document</li> <li>sentence_2_position: Number of sentence for sentence_2 in the source document</li> <li>sentence_2_paragraph: Number of paragraph for sentence_2 in the source document</li> <li>id: Unique identifier for the example</li> <li>dataset: Source corpus of the pair. Metadata of corpus, including source can be found in dataset_metadata.xlsx.</li> <li>genre: Writing genre of the dataset.</li> <li>domain: Domain genre of the dataset. </li> </ul> <p>Example: </p> <p><code>{"sentence_1":"sefior Bcajavides no es moderado, tampoco lo convertirse e\u00f1 declarada divergencia de miras polileido en griego","sentence_2":"era mayor claricomentarios, as\u00ed de los peri\u00f3dicos como de los homes dado \u00e1 la voluntad de los hombres, sin que sobreticas","connector":"por consiguiente,","connector_type":"reasoning","extraction_strategy":"linking_phrase","distance":1.0,"sentence_1_paragraph":4,"sentence_1_position":86,"sentence_2_paragraph":4,"sentence_2_position":87,"id":"esnews__spanish_pd_news__531537","dataset":"esnews__spanish_pd_news","genre":"news","domain":"spanish_public_domain_news"}</code></p> <h2>Dataset load</h2> <p>To load a dataset/split as a pytorch object used to train-validate-test models you must use the custom class dataset </p> <div> <div><code>from auto_nli.model.bert_based.dataset import BERTDataset<br></code></div> <div> </div> <div><code>dataset = BERTDataset(</code></div> <div><code>os.path.join(dataset_folder, <Path to jsonl>),</code></div> <div><code>max_len=<max length of sentences>,</code></div> <div> <div> <div><code>model_type=</code></div> </div> <code><type of model to use for tokenizer>,</code></div> <div><code>only_premise=<True to load a dataset with only the first sentences>,</code></div> <div><code>max_samples=<Maximum number of examples in the dataset>)</code></div> <div> </div> <p>----------------------------------------------------------------------------------------------</p> <h1>Notebooks</h1> <div>The folder notebooks contains a collection of jupyter notebooks used to preprocess datasets and visualize results.</div> </div> |
| title | Complete code and datasets for "ESNLIR: Expanding Spanish NLI Benchmarks with Multi-Genre and Causal Annotation" |
| url | https://doi.org/10.5281/zenodo.15002575 |