Sequence-to-Sequence Spanish Pre-trained Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Araujo, Vladimir, Trusca, Maria Mihaela, Tufiño, Rodrigo, Moens, Marie-Francine
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917618700517376
author Araujo, Vladimir
Trusca, Maria Mihaela
Tufiño, Rodrigo
Moens, Marie-Francine
author_facet Araujo, Vladimir
Trusca, Maria Mihaela
Tufiño, Rodrigo
Moens, Marie-Francine
contents In recent years, significant advancements in pre-trained language models have driven the creation of numerous non-English language variants, with a particular emphasis on encoder-only and decoder-only architectures. While Spanish language models based on BERT and GPT have demonstrated proficiency in natural language understanding and generation, there remains a noticeable scarcity of encoder-decoder models explicitly designed for sequence-to-sequence tasks, which aim to map input sequences to generate output sequences conditionally. This paper breaks new ground by introducing the implementation and evaluation of renowned encoder-decoder architectures exclusively pre-trained on Spanish corpora. Specifically, we present Spanish versions of BART, T5, and BERT2BERT-style models and subject them to a comprehensive assessment across various sequence-to-sequence tasks, including summarization, question answering, split-and-rephrase, dialogue, and translation. Our findings underscore the competitive performance of all models, with the BART- and T5-based models emerging as top performers across all tasks. We have made all models publicly available to the research community to foster future explorations and advancements in Spanish NLP: https://github.com/vgaraujov/Seq2Seq-Spanish-PLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2309_11259
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Sequence-to-Sequence Spanish Pre-trained Language Models
Araujo, Vladimir
Trusca, Maria Mihaela
Tufiño, Rodrigo
Moens, Marie-Francine
Computation and Language
Artificial Intelligence
Machine Learning
In recent years, significant advancements in pre-trained language models have driven the creation of numerous non-English language variants, with a particular emphasis on encoder-only and decoder-only architectures. While Spanish language models based on BERT and GPT have demonstrated proficiency in natural language understanding and generation, there remains a noticeable scarcity of encoder-decoder models explicitly designed for sequence-to-sequence tasks, which aim to map input sequences to generate output sequences conditionally. This paper breaks new ground by introducing the implementation and evaluation of renowned encoder-decoder architectures exclusively pre-trained on Spanish corpora. Specifically, we present Spanish versions of BART, T5, and BERT2BERT-style models and subject them to a comprehensive assessment across various sequence-to-sequence tasks, including summarization, question answering, split-and-rephrase, dialogue, and translation. Our findings underscore the competitive performance of all models, with the BART- and T5-based models emerging as top performers across all tasks. We have made all models publicly available to the research community to foster future explorations and advancements in Spanish NLP: https://github.com/vgaraujov/Seq2Seq-Spanish-PLMs.
title Sequence-to-Sequence Spanish Pre-trained Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2309.11259