AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915895323918336 |
|---|---|
| author | Simplício, Afonso Vinagre, Gonçalo Ramos, Miguel Moura Tavares, Diogo Ferreira, Rafael Attanasio, Giuseppe Alves, Duarte M. Calvo, Inês Vieira, Inês Guerra, Rui Furtado, James Canaverde, Beatriz Paulo, Iago Ramos, Vasco Glória-Silva, Diogo Faria, Miguel Treviso, Marcos Gomes, Daniel Gomes, Pedro Semedo, David Martins, André Magalhães, João |
| author_facet | Simplício, Afonso Vinagre, Gonçalo Ramos, Miguel Moura Tavares, Diogo Ferreira, Rafael Attanasio, Giuseppe Alves, Duarte M. Calvo, Inês Vieira, Inês Guerra, Rui Furtado, James Canaverde, Beatriz Paulo, Iago Ramos, Vasco Glória-Silva, Diogo Faria, Miguel Treviso, Marcos Gomes, Daniel Gomes, Pedro Semedo, David Martins, André Magalhães, João |
| contents | Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_26511 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese Simplício, Afonso Vinagre, Gonçalo Ramos, Miguel Moura Tavares, Diogo Ferreira, Rafael Attanasio, Giuseppe Alves, Duarte M. Calvo, Inês Vieira, Inês Guerra, Rui Furtado, James Canaverde, Beatriz Paulo, Iago Ramos, Vasco Glória-Silva, Diogo Faria, Miguel Treviso, Marcos Gomes, Daniel Gomes, Pedro Semedo, David Martins, André Magalhães, João Computation and Language Artificial Intelligence Machine Learning I.2.7 Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese. |
| title | AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese |
| topic | Computation and Language Artificial Intelligence Machine Learning I.2.7 |
| url | https://arxiv.org/abs/2603.26511 |