TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
Fuente:
arXiv
Saved in:
| Main Authors: | Almeida, Thales Sales, Bonás, Giovana Kerche, Santos, João Guilherme Alves, Abonizio, Hugo, Nogueira, Rodrigo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
by: Santos, João Guilherme Alves, et al.
Published: (2025)
by: Santos, João Guilherme Alves, et al.
Published: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Sabiá-3 Technical Report
by: Abonizio, Hugo, et al.
Published: (2024)
by: Abonizio, Hugo, et al.
Published: (2024)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026)
by: Junior, Roseval Malaquias, et al.
Published: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026)
by: Laitz, Thiago, et al.
Published: (2026)
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024)
by: Almeida, Thales Sales, et al.
Published: (2024)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
by: Abonizio, Hugo, et al.
Published: (2025)
by: Abonizio, Hugo, et al.
Published: (2025)
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines
by: Abonizio, Hugo, et al.
Published: (2026)
by: Abonizio, Hugo, et al.
Published: (2026)
Measuring Cross-lingual Transfer in Bytes
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
Estudo do uso de plantas medicinais pela comunidade quilombola Senhor do Bonfim - Areia-PB
by: Giovana Patrícia dos Santos Sales
Published: (2009)
by: Giovana Patrícia dos Santos Sales
Published: (2009)
CONFLITOS À MESA: Vegetarianos, consumo e identidade
by: Juliana Abonizio
Published: (2016)
by: Juliana Abonizio
Published: (2016)
Por uma quiromancia da vida urbana
by: Juliana Abonizio
Published: (2011)
by: Juliana Abonizio
Published: (2011)
Consumo alimentar e anticonsumismo: veganos e freeganos
by: Juliana Abonizio
Published: (2013)
by: Juliana Abonizio
Published: (2013)
INDEPENDÊNCIA, PODER JUDICIÁRIO E MINISTÉRIO PÚBLICO
by: Fábio Kerche
Published: (2018)
by: Fábio Kerche
Published: (2018)
MINISTÉRIO PÚBLICO, LAVA JATO E MÃOS LIMPAS: UMA ABORDAGEM INSTITUCIONAL
by: Fábio Kerche
Published: (2018)
by: Fábio Kerche
Published: (2018)
Os Conselhos Nacionais de Justiça e do Ministério Público no Brasil: instrumentos de accountability?
by: Fábio Kerche
Published: (2020)
by: Fábio Kerche
Published: (2020)
Autonomia e Discricionariedade do Ministério Público no Brasil
by: Fábio Kerche
Published: (2007)
by: Fábio Kerche
Published: (2007)
The interplay between domain specialization and model size
by: Junior, Roseval Malaquias, et al.
Published: (2025)
by: Junior, Roseval Malaquias, et al.
Published: (2025)
DE SEM-TERRA A SEM-TERRA: MEMÓRIAS E IDENTIDADES
by: Natália Kerche Alvaides
Published: (2013)
by: Natália Kerche Alvaides
Published: (2013)
SurveySum: A Dataset for Summarizing Multiple Scientific Articles into a Survey Section
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
Principais Padrões de Verão da Pressão ao Nível do Mar sobre a Região da América do Sul no Clima Presente e em Projeções Futuras
by: Thales Alves Teodoro
Published: (2022)
by: Thales Alves Teodoro
Published: (2022)
Fungicide application technology for controlling the sugarcane orange rust
by: Thales Cassemiro Alves
Published: (2019)
by: Thales Cassemiro Alves
Published: (2019)
Forecasting Events in Soccer Matches Through Language
by: Mendes-Neves, Tiago, et al.
Published: (2024)
by: Mendes-Neves, Tiago, et al.
Published: (2024)
The Best, the Notable and the Recommended.
Published: (1994)
Published: (1994)
Notable Books of 1968
Published: (1969)
Published: (1969)
The Best, the Notable & the Recommended.
Published: (1995)
Published: (1995)
Notable Documents 1986.
Published: (1987)
Published: (1987)
Event Segmentation Applications in Large Language Model Enabled Automated Recall Assessments
by: Panela, Ryan A., et al.
Published: (2025)
by: Panela, Ryan A., et al.
Published: (2025)
Skip That Beat: Augmenting Meter Tracking Models for Underrepresented Time Signatures
by: Morais, Giovana, et al.
Published: (2025)
by: Morais, Giovana, et al.
Published: (2025)
Do Large Language Models Understand Data Visualization Rules?
by: Sinnona, Martin, et al.
Published: (2026)
by: Sinnona, Martin, et al.
Published: (2026)
Similar Items
-
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025) -
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
by: Santos, João Guilherme Alves, et al.
Published: (2025) -
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025) -
Sabiá-3 Technical Report
by: Abonizio, Hugo, et al.
Published: (2024) -
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)