AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Simplício, Afonso, Vinagre, Gonçalo, Ramos, Miguel Moura, Tavares, Diogo, Ferreira, Rafael, Attanasio, Giuseppe, Alves, Duarte M., Calvo, Inês, Vieira, Inês, Guerra, Rui, Furtado, James, Canaverde, Beatriz, Paulo, Iago, Ramos, Vasco, Glória-Silva, Diogo, Faria, Miguel, Treviso, Marcos, Gomes, Daniel, Gomes, Pedro, Semedo, David, Martins, André, Magalhães, João
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915895323918336
author Simplício, Afonso
Vinagre, Gonçalo
Ramos, Miguel Moura
Tavares, Diogo
Ferreira, Rafael
Attanasio, Giuseppe
Alves, Duarte M.
Calvo, Inês
Vieira, Inês
Guerra, Rui
Furtado, James
Canaverde, Beatriz
Paulo, Iago
Ramos, Vasco
Glória-Silva, Diogo
Faria, Miguel
Treviso, Marcos
Gomes, Daniel
Gomes, Pedro
Semedo, David
Martins, André
Magalhães, João
author_facet Simplício, Afonso
Vinagre, Gonçalo
Ramos, Miguel Moura
Tavares, Diogo
Ferreira, Rafael
Attanasio, Giuseppe
Alves, Duarte M.
Calvo, Inês
Vieira, Inês
Guerra, Rui
Furtado, James
Canaverde, Beatriz
Paulo, Iago
Ramos, Vasco
Glória-Silva, Diogo
Faria, Miguel
Treviso, Marcos
Gomes, Daniel
Gomes, Pedro
Semedo, David
Martins, André
Magalhães, João
contents Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26511
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
Simplício, Afonso
Vinagre, Gonçalo
Ramos, Miguel Moura
Tavares, Diogo
Ferreira, Rafael
Attanasio, Giuseppe
Alves, Duarte M.
Calvo, Inês
Vieira, Inês
Guerra, Rui
Furtado, James
Canaverde, Beatriz
Paulo, Iago
Ramos, Vasco
Glória-Silva, Diogo
Faria, Miguel
Treviso, Marcos
Gomes, Daniel
Gomes, Pedro
Semedo, David
Martins, André
Magalhães, João
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese.
title AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
url https://arxiv.org/abs/2603.26511