PeLLE: Encoder-based language models for Brazilian Portuguese based on open data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: de Mello, Guilherme Lamartine, Finger, Marcelo, Serras, and Felipe, Carpi, Miguel de Mello, Jose, Marcos Menon, Domingues, Pedro Henrique, Cavalim, Paulo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916142171291648
author de Mello, Guilherme Lamartine
Finger, Marcelo
Serras, and Felipe
Carpi, Miguel de Mello
Jose, Marcos Menon
Domingues, Pedro Henrique
Cavalim, Paulo
author_facet de Mello, Guilherme Lamartine
Finger, Marcelo
Serras, and Felipe
Carpi, Miguel de Mello
Jose, Marcos Menon
Domingues, Pedro Henrique
Cavalim, Paulo
contents In this paper we present PeLLE, a family of large language models based on the RoBERTa architecture, for Brazilian Portuguese, trained on curated, open data from the Carolina corpus. Aiming at reproducible results, we describe details of the pretraining of the models. We also evaluate PeLLE models against a set of existing multilingual and PT-BR refined pretrained Transformer-based LLM encoders, contrasting performance of large versus smaller-but-curated pretrained models in several downstream tasks. We conclude that several tasks perform better with larger models, but some tasks benefit from smaller-but-curated data in its pretraining.
format Preprint
id arxiv_https___arxiv_org_abs_2402_19204
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PeLLE: Encoder-based language models for Brazilian Portuguese based on open data
de Mello, Guilherme Lamartine
Finger, Marcelo
Serras, and Felipe
Carpi, Miguel de Mello
Jose, Marcos Menon
Domingues, Pedro Henrique
Cavalim, Paulo
Computation and Language
I.2.7
In this paper we present PeLLE, a family of large language models based on the RoBERTa architecture, for Brazilian Portuguese, trained on curated, open data from the Carolina corpus. Aiming at reproducible results, we describe details of the pretraining of the models. We also evaluate PeLLE models against a set of existing multilingual and PT-BR refined pretrained Transformer-based LLM encoders, contrasting performance of large versus smaller-but-curated pretrained models in several downstream tasks. We conclude that several tasks perform better with larger models, but some tasks benefit from smaller-but-curated data in its pretraining.
title PeLLE: Encoder-based language models for Brazilian Portuguese based on open data
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2402.19204