CLARIN-PT-LDB: An Open LLM Leaderboard for Portuguese to assess Language, Culture and Civility
Fuente:
arXiv
Saved in:
| Main Authors: | Silva, João, Gomes, Luís, Branco, António |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Generative AI for Portuguese with Open Decoder Gervásio PT*
by: Santos, Rodrigo, et al.
Published: (2024)
by: Santos, Rodrigo, et al.
Published: (2024)
Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family
by: Gomes, Luís, et al.
Published: (2024)
by: Gomes, Luís, et al.
Published: (2024)
Fostering the Ecosystem of Open Neural Encoders for Portuguese with Albertina PT* Family
by: Santos, Rodrigo, et al.
Published: (2024)
by: Santos, Rodrigo, et al.
Published: (2024)
Advancing Neural Encoding of Portuguese with Transformer Albertina PT-*
by: Rodrigues, João, et al.
Published: (2023)
by: Rodrigues, João, et al.
Published: (2023)
ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
by: Campos, Ricardo, et al.
Published: (2026)
by: Campos, Ricardo, et al.
Published: (2026)
PORTULAN ExtraGLUE Datasets and Models: Kick-starting a Benchmark for the Neural Processing of Portuguese
by: Osório, Tomás, et al.
Published: (2024)
by: Osório, Tomás, et al.
Published: (2024)
Open Universal Arabic ASR Leaderboard
by: Wang, Yingzhi, et al.
Published: (2024)
by: Wang, Yingzhi, et al.
Published: (2024)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
by: Hong, Giwon, et al.
Published: (2024)
by: Hong, Giwon, et al.
Published: (2024)
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
by: Grandury, María, et al.
Published: (2025)
by: Grandury, María, et al.
Published: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
ACE-2005-PT: Corpus for Event Extraction in Portuguese
by: Cunha, Luís Filipe, et al.
Published: (2024)
by: Cunha, Luís Filipe, et al.
Published: (2024)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
by: Myrzakhan, Aidar, et al.
Published: (2024)
by: Myrzakhan, Aidar, et al.
Published: (2024)
Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training
by: Santos, Rodrigo, et al.
Published: (2025)
by: Santos, Rodrigo, et al.
Published: (2025)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Improving LLM Leaderboards with Psychometrical Methodology
by: Federiakin, Denis
Published: (2025)
by: Federiakin, Denis
Published: (2025)
LegalBench.PT: A Benchmark for Portuguese Law
by: Canaverde, Beatriz, et al.
Published: (2025)
by: Canaverde, Beatriz, et al.
Published: (2025)
GlórIA -- A Generative and Open Large Language Model for Portuguese
by: Lopes, Ricardo, et al.
Published: (2024)
by: Lopes, Ricardo, et al.
Published: (2024)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
Meta-prompting Optimized Retrieval-augmented Generation
by: Rodrigues, João, et al.
Published: (2024)
by: Rodrigues, João, et al.
Published: (2024)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
by: Rao, Varun, et al.
Published: (2025)
by: Rao, Varun, et al.
Published: (2025)
The Trust Paradox: How CS Researchers Engage LLM Leaderboards
by: Sadeghi, Pouya, et al.
Published: (2026)
by: Sadeghi, Pouya, et al.
Published: (2026)
Prompt-to-Leaderboard
by: Frick, Evan, et al.
Published: (2025)
by: Frick, Evan, et al.
Published: (2025)
Effective Context Selection in LLM-based Leaderboard Generation: An Empirical Study
by: Kabongo, Salomon, et al.
Published: (2024)
by: Kabongo, Salomon, et al.
Published: (2024)
Sovereign AI-based Public Services are Viable and Affordable
by: Branco, António, et al.
Published: (2026)
by: Branco, António, et al.
Published: (2026)
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
by: Li, Haonan, et al.
Published: (2024)
by: Li, Haonan, et al.
Published: (2024)
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Progressing beyond Art Masterpieces or Touristic Clichés: how to assess your LLMs for cultural alignment?
by: Branco, António, et al.
Published: (2026)
by: Branco, António, et al.
Published: (2026)
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification
by: North, Kai, et al.
Published: (2022)
by: North, Kai, et al.
Published: (2022)
Leveraging LLMs for On-the-Fly Instruction Guided Image Editing
by: Santos, Rodrigo, et al.
Published: (2024)
by: Santos, Rodrigo, et al.
Published: (2024)
LEGOBench: Scientific Leaderboard Generation Benchmark
by: Singh, Shruti, et al.
Published: (2024)
by: Singh, Shruti, et al.
Published: (2024)
MedPT: A Massive Medical Question Answering Dataset for Brazilian-Portuguese Speakers
by: Färber, Fernanda Bufon, et al.
Published: (2025)
by: Färber, Fernanda Bufon, et al.
Published: (2025)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
League: Leaderboard Generation on Demand
by: Wu, Jian, et al.
Published: (2025)
by: Wu, Jian, et al.
Published: (2025)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
by: Topsakal, Oguzhan, et al.
Published: (2024)
by: Topsakal, Oguzhan, et al.
Published: (2024)
SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
by: Du, Jiayu, et al.
Published: (2024)
by: Du, Jiayu, et al.
Published: (2024)
CLARIN
Published: (2023)
Published: (2023)
Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
by: Şahinuç, Furkan, et al.
Published: (2024)
by: Şahinuç, Furkan, et al.
Published: (2024)
Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing
by: Boughorbel, Sabri, et al.
Published: (2025)
by: Boughorbel, Sabri, et al.
Published: (2025)
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
by: Cheng, Aileen, et al.
Published: (2025)
by: Cheng, Aileen, et al.
Published: (2025)
Similar Items
-
Advancing Generative AI for Portuguese with Open Decoder Gervásio PT*
by: Santos, Rodrigo, et al.
Published: (2024) -
Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family
by: Gomes, Luís, et al.
Published: (2024) -
Fostering the Ecosystem of Open Neural Encoders for Portuguese with Albertina PT* Family
by: Santos, Rodrigo, et al.
Published: (2024) -
Advancing Neural Encoding of Portuguese with Transformer Albertina PT-*
by: Rodrigues, João, et al.
Published: (2023) -
ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
by: Campos, Ricardo, et al.
Published: (2026)