ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Piau, Marcos, Lotufo, Roberto, Nogueira, Rodrigo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917839707832320
author Piau, Marcos
Lotufo, Roberto
Nogueira, Rodrigo
author_facet Piau, Marcos
Lotufo, Roberto
Nogueira, Rodrigo
contents Despite advancements in Natural Language Processing (NLP) and the growing availability of pretrained models, the English language remains the primary focus of model development. Continued pretraining on language-specific corpora provides a practical solution for adapting models to other languages. However, the impact of different pretraining settings on downstream tasks remains underexplored. This work introduces $\texttt{ptt5-v2}$, investigating the continued pretraining of T5 models for Portuguese. We first develop a baseline set of settings and pretrain models with sizes up to 3B parameters. Finetuning on three Portuguese downstream tasks (assin2 STS, assin2 RTE, and TweetSentBR) yields SOTA results on the latter two. We then explore the effects of different pretraining configurations, including pretraining data quality, optimization strategies, and multi-epoch pretraining. Perhaps surprisingly, their impact remains subtle compared to our baseline. We release $\texttt{ptt5-v2}$ pretrained checkpoints and their MonoT5-based finetuned $\texttt{MonoPTT5}$ rerankers on HuggingFace in their respective collections at \url{https://huggingface.co/unicamp-dl}.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10806
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language
Piau, Marcos
Lotufo, Roberto
Nogueira, Rodrigo
Computation and Language
Artificial Intelligence
Information Retrieval
Despite advancements in Natural Language Processing (NLP) and the growing availability of pretrained models, the English language remains the primary focus of model development. Continued pretraining on language-specific corpora provides a practical solution for adapting models to other languages. However, the impact of different pretraining settings on downstream tasks remains underexplored. This work introduces $\texttt{ptt5-v2}$, investigating the continued pretraining of T5 models for Portuguese. We first develop a baseline set of settings and pretrain models with sizes up to 3B parameters. Finetuning on three Portuguese downstream tasks (assin2 STS, assin2 RTE, and TweetSentBR) yields SOTA results on the latter two. We then explore the effects of different pretraining configurations, including pretraining data quality, optimization strategies, and multi-epoch pretraining. Perhaps surprisingly, their impact remains subtle compared to our baseline. We release $\texttt{ptt5-v2}$ pretrained checkpoints and their MonoT5-based finetuned $\texttt{MonoPTT5}$ rerankers on HuggingFace in their respective collections at \url{https://huggingface.co/unicamp-dl}.
title ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2406.10806