BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Almeida, Thales Sales, Bonás, Giovana Kerche, Santos, João Guilherme Alves |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
von: Santos, João Guilherme Alves, et al.
Veröffentlicht: (2025)
von: Santos, João Guilherme Alves, et al.
Veröffentlicht: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2026)
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2026)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
von: Bonás, Giovana Kerche, et al.
Veröffentlicht: (2026)
von: Bonás, Giovana Kerche, et al.
Veröffentlicht: (2026)
Sabiá-3 Technical Report
von: Abonizio, Hugo, et al.
Veröffentlicht: (2024)
von: Abonizio, Hugo, et al.
Veröffentlicht: (2024)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
von: Nogueira, Rodrigo, et al.
Veröffentlicht: (2026)
von: Nogueira, Rodrigo, et al.
Veröffentlicht: (2026)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
von: Nogueira, Rodrigo, et al.
Veröffentlicht: (2026)
von: Nogueira, Rodrigo, et al.
Veröffentlicht: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2026)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2026)
Sabiá-4 Technical Report
von: Laitz, Thiago, et al.
Veröffentlicht: (2026)
von: Laitz, Thiago, et al.
Veröffentlicht: (2026)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2026)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2026)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
von: Pires, Ramon, et al.
Veröffentlicht: (2026)
von: Pires, Ramon, et al.
Veröffentlicht: (2026)
Sabiá-2: A New Generation of Portuguese Large Language Models
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2024)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2024)
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Measuring Cross-lingual Transfer in Bytes
von: de Souza, Leandro Rodrigues, et al.
Veröffentlicht: (2024)
von: de Souza, Leandro Rodrigues, et al.
Veröffentlicht: (2024)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
von: Abonizio, Hugo, et al.
Veröffentlicht: (2025)
von: Abonizio, Hugo, et al.
Veröffentlicht: (2025)
ELEPHANT: Measuring and understanding social sycophancy in LLMs
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
Advancing Generative AI for Portuguese with Open Decoder Gervásio PT*
von: Santos, Rodrigo, et al.
Veröffentlicht: (2024)
von: Santos, Rodrigo, et al.
Veröffentlicht: (2024)
Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family
von: Gomes, Luís, et al.
Veröffentlicht: (2024)
von: Gomes, Luís, et al.
Veröffentlicht: (2024)
LinGO: A Linguistic Graph Optimization Framework with LLMs for Interpreting Intents of Online Uncivil Discourse
von: Zhang, Yuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuan, et al.
Veröffentlicht: (2026)
Advancing Neural Encoding of Portuguese with Transformer Albertina PT-*
von: Rodrigues, João, et al.
Veröffentlicht: (2023)
von: Rodrigues, João, et al.
Veröffentlicht: (2023)
Adapting LLMs for the Medical Domain in Portuguese: A Study on Fine-Tuning and Model Evaluation
von: Paiola, Pedro Henrique, et al.
Veröffentlicht: (2024)
von: Paiola, Pedro Henrique, et al.
Veröffentlicht: (2024)
so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs
von: Bhyravajjula, Sriharsh, et al.
Veröffentlicht: (2025)
von: Bhyravajjula, Sriharsh, et al.
Veröffentlicht: (2025)
Fostering the Ecosystem of Open Neural Encoders for Portuguese with Albertina PT* Family
von: Santos, Rodrigo, et al.
Veröffentlicht: (2024)
von: Santos, Rodrigo, et al.
Veröffentlicht: (2024)
From Brazilian Portuguese to European Portuguese
von: Sanches, João, et al.
Veröffentlicht: (2024)
von: Sanches, João, et al.
Veröffentlicht: (2024)
SurveySum: A Dataset for Summarizing Multiple Scientific Articles into a Survey Section
von: Fernandes, Leandro Carísio, et al.
Veröffentlicht: (2024)
von: Fernandes, Leandro Carísio, et al.
Veröffentlicht: (2024)
The interplay between domain specialization and model size
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2025)
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2025)
Using Shapley interactions to understand how models use structure
von: Singhvi, Divyansh, et al.
Veröffentlicht: (2024)
von: Singhvi, Divyansh, et al.
Veröffentlicht: (2024)
Clinical named entity recognition in the Portuguese language: a benchmark of modern BERT models and LLMs
von: de Almeida, Vinicius Anjos, et al.
Veröffentlicht: (2026)
von: de Almeida, Vinicius Anjos, et al.
Veröffentlicht: (2026)
CMMLU: Measuring massive multitask language understanding in Chinese
von: Li, Haonan, et al.
Veröffentlicht: (2023)
von: Li, Haonan, et al.
Veröffentlicht: (2023)
Progressing beyond Art Masterpieces or Touristic Clichés: how to assess your LLMs for cultural alignment?
von: Branco, António, et al.
Veröffentlicht: (2026)
von: Branco, António, et al.
Veröffentlicht: (2026)
Performance in a dialectal profiling task of LLMs for varieties of Brazilian Portuguese
von: Freitag, Raquel Meister Ko, et al.
Veröffentlicht: (2024)
von: Freitag, Raquel Meister Ko, et al.
Veröffentlicht: (2024)
How much do language models memorize?
von: Morris, John X., et al.
Veröffentlicht: (2025)
von: Morris, John X., et al.
Veröffentlicht: (2025)
PORTULAN ExtraGLUE Datasets and Models: Kick-starting a Benchmark for the Neural Processing of Portuguese
von: Osório, Tomás, et al.
Veröffentlicht: (2024)
von: Osório, Tomás, et al.
Veröffentlicht: (2024)
Enhancing Portuguese Variety Identification with Cross-Domain Approaches
von: Sousa, Hugo, et al.
Veröffentlicht: (2025)
von: Sousa, Hugo, et al.
Veröffentlicht: (2025)
CLARIN-PT-LDB: An Open LLM Leaderboard for Portuguese to assess Language, Culture and Civility
von: Silva, João, et al.
Veröffentlicht: (2026)
von: Silva, João, et al.
Veröffentlicht: (2026)
Effectiveness of Yoruba proverbs in acquiring Yoruba language and culture
von: Adebimpe Adegbite
Veröffentlicht: (2025)
von: Adebimpe Adegbite
Veröffentlicht: (2025)
Tucano 2 Cool: Better Open Source LLMs for Portuguese
von: Corrêa, Nicholas Kluge, et al.
Veröffentlicht: (2026)
von: Corrêa, Nicholas Kluge, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
von: Santos, João Guilherme Alves, et al.
Veröffentlicht: (2025) -
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025) -
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025) -
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2026) -
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
von: Bonás, Giovana Kerche, et al.
Veröffentlicht: (2026)