Sabiá-4 Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911504208494592 |
|---|---|
| author | Laitz, Thiago Almeida, Thales Sales Abonizio, Hugo Junior, Roseval Malaquias Bonás, Giovana Kerche Piau, Marcos Larcher, Celio Pires, Ramon Nogueira, Rodrigo |
| author_facet | Laitz, Thiago Almeida, Thales Sales Abonizio, Hugo Junior, Roseval Malaquias Bonás, Giovana Kerche Piau, Marcos Larcher, Celio Pires, Ramon Nogueira, Rodrigo |
| contents | This technical report presents Sabiá-4 and Sabiazinho-4, a new generation of Portuguese language models with a focus on Brazilian Portuguese language. The models were developed through a four-stage training pipeline: continued pre-training on Portuguese and Brazilian legal corpora, long-context extension to 128K tokens, supervised fine-tuning on instruction data spanning chat, code, legal tasks, and function calling, and preference alignment. We evaluate the models on six benchmark categories: conversational capabilities in Brazilian Portuguese, knowledge of Brazilian legislation, long-context understanding, instruction following, standardized exams, and agentic capabilities including tool use and web navigation. Results show that Sabiá-4 and Sabiazinho-4 achieve a favorable cost-performance trade-off compared to other models, positioning them in the upper-left region of the pricing-accuracy chart. The models show improvements over previous generations in legal document drafting, multi-turn dialogue quality, and agentic task completion. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_10213 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Sabiá-4 Technical Report Laitz, Thiago Almeida, Thales Sales Abonizio, Hugo Junior, Roseval Malaquias Bonás, Giovana Kerche Piau, Marcos Larcher, Celio Pires, Ramon Nogueira, Rodrigo Computation and Language This technical report presents Sabiá-4 and Sabiazinho-4, a new generation of Portuguese language models with a focus on Brazilian Portuguese language. The models were developed through a four-stage training pipeline: continued pre-training on Portuguese and Brazilian legal corpora, long-context extension to 128K tokens, supervised fine-tuning on instruction data spanning chat, code, legal tasks, and function calling, and preference alignment. We evaluate the models on six benchmark categories: conversational capabilities in Brazilian Portuguese, knowledge of Brazilian legislation, long-context understanding, instruction following, standardized exams, and agentic capabilities including tool use and web navigation. Results show that Sabiá-4 and Sabiazinho-4 achieve a favorable cost-performance trade-off compared to other models, positioning them in the upper-left region of the pricing-accuracy chart. The models show improvements over previous generations in legal document drafting, multi-turn dialogue quality, and agentic task completion. |
| title | Sabiá-4 Technical Report |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2603.10213 |