Sabiá-3 Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | Abonizio, Hugo, Almeida, Thales Sales, Laitz, Thiago, Junior, Roseval Malaquias, Bonás, Giovana Kerche, Nogueira, Rodrigo, Pires, Ramon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026)
by: Laitz, Thiago, et al.
Published: (2026)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026)
by: Junior, Roseval Malaquias, et al.
Published: (2026)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024)
by: Almeida, Thales Sales, et al.
Published: (2024)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Automatic Legal Writing Evaluation of LLMs
by: Pires, Ramon, et al.
Published: (2025)
by: Pires, Ramon, et al.
Published: (2025)
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
by: Santos, João Guilherme Alves, et al.
Published: (2025)
by: Santos, João Guilherme Alves, et al.
Published: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
The interplay between domain specialization and model size
by: Junior, Roseval Malaquias, et al.
Published: (2025)
by: Junior, Roseval Malaquias, et al.
Published: (2025)
Juru: Legal Brazilian Large Language Model from Reputable Sources
by: Junior, Roseval Malaquias, et al.
Published: (2024)
by: Junior, Roseval Malaquias, et al.
Published: (2024)
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
ExaRanker-Open: Synthetic Explanation for IR using Open-Source LLMs
by: Ferraretto, Fernando, et al.
Published: (2024)
by: Ferraretto, Fernando, et al.
Published: (2024)
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
by: Abonizio, Hugo, et al.
Published: (2025)
by: Abonizio, Hugo, et al.
Published: (2025)
SurveySum: A Dataset for Summarizing Multiple Scientific Articles into a Survey Section
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Gemma 3 Technical Report
by: Gemma Team, et al.
Published: (2025)
by: Gemma Team, et al.
Published: (2025)
DeepSeek-V3 Technical Report
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
GPT-4 Technical Report
by: OpenAI, et al.
Published: (2023)
by: OpenAI, et al.
Published: (2023)
Ruyi2 Technical Report
by: Song, Huan, et al.
Published: (2026)
by: Song, Huan, et al.
Published: (2026)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
Jupiter-N Technical Report
by: Drayson, George
Published: (2026)
by: Drayson, George
Published: (2026)
QZhou-Embedding Technical Report
by: Yu, Peng, et al.
Published: (2025)
by: Yu, Peng, et al.
Published: (2025)
Xmodel-LM Technical Report
by: Wang, Yichuan, et al.
Published: (2024)
by: Wang, Yichuan, et al.
Published: (2024)
Phi-4 Technical Report
by: Abdin, Marah, et al.
Published: (2024)
by: Abdin, Marah, et al.
Published: (2024)
Qwen2 Technical Report
by: Yang, An, et al.
Published: (2024)
by: Yang, An, et al.
Published: (2024)
Similar Items
-
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026) -
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026) -
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026) -
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024) -
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)