PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
Fuente:
arXiv
Saved in:
| Main Authors: | Almeida, Thales Sales, Pires, Ramon, Abonizio, Hugo, Nogueira, Rodrigo, Pedrini, Hélio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024)
by: Almeida, Thales Sales, et al.
Published: (2024)
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026)
by: Junior, Roseval Malaquias, et al.
Published: (2026)
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
by: Abonizio, Hugo, et al.
Published: (2025)
by: Abonizio, Hugo, et al.
Published: (2025)
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Sabiá-3 Technical Report
by: Abonizio, Hugo, et al.
Published: (2024)
by: Abonizio, Hugo, et al.
Published: (2024)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026)
by: Laitz, Thiago, et al.
Published: (2026)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
The interplay between domain specialization and model size
by: Junior, Roseval Malaquias, et al.
Published: (2025)
by: Junior, Roseval Malaquias, et al.
Published: (2025)
Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines
by: Abonizio, Hugo, et al.
Published: (2026)
by: Abonizio, Hugo, et al.
Published: (2026)
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language
by: Piau, Marcos, et al.
Published: (2024)
by: Piau, Marcos, et al.
Published: (2024)
Measuring Cross-lingual Transfer in Bytes
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
Juru: Legal Brazilian Large Language Model from Reputable Sources
by: Junior, Roseval Malaquias, et al.
Published: (2024)
by: Junior, Roseval Malaquias, et al.
Published: (2024)
Automatic Legal Writing Evaluation of LLMs
by: Pires, Ramon, et al.
Published: (2025)
by: Pires, Ramon, et al.
Published: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
SurveySum: A Dataset for Summarizing Multiple Scientific Articles into a Survey Section
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
Clustering Discourses: Racial Biases in Short Stories about Women Generated by Large Language Models
by: Bonil, Gustavo, et al.
Published: (2025)
by: Bonil, Gustavo, et al.
Published: (2025)
Yet another algorithmic bias: A Discursive Analysis of Large Language Models Reinforcing Dominant Discourses on Gender and Race
by: Bonil, Gustavo, et al.
Published: (2025)
by: Bonil, Gustavo, et al.
Published: (2025)
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
by: Chang, Emily, et al.
Published: (2025)
by: Chang, Emily, et al.
Published: (2025)
Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study
by: Silva, Jhessica, et al.
Published: (2025)
by: Silva, Jhessica, et al.
Published: (2025)
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
by: Santos, João Guilherme Alves, et al.
Published: (2025)
by: Santos, João Guilherme Alves, et al.
Published: (2025)
Towards High-Fidelity Synthetic Multi-platform Social Media Datasets via Large Language Models
by: Tari, Henry, et al.
Published: (2025)
by: Tari, Henry, et al.
Published: (2025)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
by: Zhu, Kaijie, et al.
Published: (2023)
by: Zhu, Kaijie, et al.
Published: (2023)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
by: Kytöniemi, Joona, et al.
Published: (2025)
by: Kytöniemi, Joona, et al.
Published: (2025)
Predictive Authoring for Brazilian Portuguese Augmentative and Alternative Communication
by: Pereira, Jayr, et al.
Published: (2023)
by: Pereira, Jayr, et al.
Published: (2023)
PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding
by: Sanchez-Bayona, Elisa, et al.
Published: (2025)
by: Sanchez-Bayona, Elisa, et al.
Published: (2025)
Power-of-Two Quantization-Aware-Training (PoT-QAT) in Large Language Models (LLMs)
by: Elgenedy, Mahmoud
Published: (2026)
by: Elgenedy, Mahmoud
Published: (2026)
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion
by: Zhang, Zekai, et al.
Published: (2024)
by: Zhang, Zekai, et al.
Published: (2024)
Evaluating the Retrieval Robustness of Large Language Models
by: Cao, Shuyang, et al.
Published: (2025)
by: Cao, Shuyang, et al.
Published: (2025)
Enhancing Portuguese Variety Identification with Cross-Domain Approaches
by: Sousa, Hugo, et al.
Published: (2025)
by: Sousa, Hugo, et al.
Published: (2025)
On the Cross-lingual Transferability of Pre-trained wav2vec2-based Models
by: Grosman, Jonatas, et al.
Published: (2025)
by: Grosman, Jonatas, et al.
Published: (2025)
Similar Items
-
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024) -
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2026) -
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
by: Almeida, Thales Sales, et al.
Published: (2025) -
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025) -
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)