Guardado en:
| Autores principales: | Grandury, María, Aula-Blasco, Javier, Falcão, Júlia, Fourrier, Clémentine, González, Miguel, Martínez, Gonzalo, Santamaría, Gonzalo, Agerri, Rodrigo, Aldama, Nuria, Chiruzzo, Luis, Conde, Javier, Gómez, Helena, Guerrero, Marta, Ivetta, Guido, López, Natalia, Plaza-del-Arco, Flor Miriam, Martín-Valdivia, María Teresa, Montoro, Helena, Muñoz, Carmen, Reviriego, Pedro, Rosado, Leire, Vaca, Alejandro, Vallecillo-Rodríguez, María Estrella, Vallego, Jorge, Zubiaga, Irune |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2507.00999 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
A LLM-Based Ranking Method for the Evaluation of Automatic Counter-Narrative Generation
por: Zubiaga, Irune, et al.
Publicado: (2024)
por: Zubiaga, Irune, et al.
Publicado: (2024)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
por: Hong, Giwon, et al.
Publicado: (2024)
por: Hong, Giwon, et al.
Publicado: (2024)
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
por: Mayor-Rocher, Marina, et al.
Publicado: (2024)
por: Mayor-Rocher, Marina, et al.
Publicado: (2024)
It's the same but not the same: Do LLMs distinguish Spanish varieties?
por: Mayor-Rocher, Marina, et al.
Publicado: (2025)
por: Mayor-Rocher, Marina, et al.
Publicado: (2025)
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
por: Ferrando, Raquel, et al.
Publicado: (2025)
por: Ferrando, Raquel, et al.
Publicado: (2025)
Pluralistic Leaderboards
por: Haghtalab, Nika, et al.
Publicado: (2026)
por: Haghtalab, Nika, et al.
Publicado: (2026)
Prompt-to-Leaderboard
por: Frick, Evan, et al.
Publicado: (2025)
por: Frick, Evan, et al.
Publicado: (2025)
The Leaderboard Illusion
por: Singh, Shivalika, et al.
Publicado: (2025)
por: Singh, Shivalika, et al.
Publicado: (2025)
On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
por: Zhao, Zhimin, et al.
Publicado: (2024)
por: Zhao, Zhimin, et al.
Publicado: (2024)
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
por: Li, Haonan, et al.
Publicado: (2024)
por: Li, Haonan, et al.
Publicado: (2024)
PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing Games
por: Góngora, Santiago, et al.
Publicado: (2025)
por: Góngora, Santiago, et al.
Publicado: (2025)
League: Leaderboard Generation on Demand
por: Wu, Jian, et al.
Publicado: (2025)
por: Wu, Jian, et al.
Publicado: (2025)
The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations
por: Arriaga, Carlos, et al.
Publicado: (2025)
por: Arriaga, Carlos, et al.
Publicado: (2025)
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
por: Chen, Wenting, et al.
Publicado: (2025)
por: Chen, Wenting, et al.
Publicado: (2025)
Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times
por: Clark, Thomas Hikaru, et al.
Publicado: (2026)
por: Clark, Thomas Hikaru, et al.
Publicado: (2026)
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
por: Cheng, Aileen, et al.
Publicado: (2025)
por: Cheng, Aileen, et al.
Publicado: (2025)
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
por: Alzahrani, Norah, et al.
Publicado: (2024)
por: Alzahrani, Norah, et al.
Publicado: (2024)
Open Universal Arabic ASR Leaderboard
por: Wang, Yingzhi, et al.
Publicado: (2024)
por: Wang, Yingzhi, et al.
Publicado: (2024)
Exploring the Latest LLMs for Leaderboard Extraction
por: Kabongo, Salomon, et al.
Publicado: (2024)
por: Kabongo, Salomon, et al.
Publicado: (2024)
Improving LLM Leaderboards with Psychometrical Methodology
por: Federiakin, Denis
Publicado: (2025)
por: Federiakin, Denis
Publicado: (2025)
LEGOBench: Scientific Leaderboard Generation Benchmark
por: Singh, Shruti, et al.
Publicado: (2024)
por: Singh, Shruti, et al.
Publicado: (2024)
The #Somos600M Project: Generating NLP resources that represent the diversity of the languages from LATAM, the Caribbean, and Spain
por: Grandury, María
Publicado: (2024)
por: Grandury, María
Publicado: (2024)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
por: Plaza, Irene, et al.
Publicado: (2024)
por: Plaza, Irene, et al.
Publicado: (2024)
Reliable, Reproducible, and Really Fast Leaderboards with Evalica
por: Ustalov, Dmitry
Publicado: (2024)
por: Ustalov, Dmitry
Publicado: (2024)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
por: Tamber, Manveer Singh, et al.
Publicado: (2025)
por: Tamber, Manveer Singh, et al.
Publicado: (2025)
Identifying Models Behind Text-to-Image Leaderboards
por: Naseh, Ali, et al.
Publicado: (2026)
por: Naseh, Ali, et al.
Publicado: (2026)
MULTI: Multimodal Understanding Leaderboard with Text and Images
por: Zhu, Zichen, et al.
Publicado: (2024)
por: Zhu, Zichen, et al.
Publicado: (2024)
Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?
por: Fu, Tarian, et al.
Publicado: (2025)
por: Fu, Tarian, et al.
Publicado: (2025)
Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism
por: Fu, Tairan, et al.
Publicado: (2026)
por: Fu, Tairan, et al.
Publicado: (2026)
Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
por: Şahinuç, Furkan, et al.
Publicado: (2024)
por: Şahinuç, Furkan, et al.
Publicado: (2024)
Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing
por: Boughorbel, Sabri, et al.
Publicado: (2025)
por: Boughorbel, Sabri, et al.
Publicado: (2025)
CLARIN-PT-LDB: An Open LLM Leaderboard for Portuguese to assess Language, Culture and Civility
por: Silva, João, et al.
Publicado: (2026)
por: Silva, João, et al.
Publicado: (2026)
RepairBench: Leaderboard of Frontier Models for Program Repair
por: Silva, André, et al.
Publicado: (2024)
por: Silva, André, et al.
Publicado: (2024)
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
por: Mazaheri, Parsa, et al.
Publicado: (2026)
por: Mazaheri, Parsa, et al.
Publicado: (2026)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
por: Huang, Yangsibo, et al.
Publicado: (2025)
por: Huang, Yangsibo, et al.
Publicado: (2025)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
por: Suri, Anshuman, et al.
Publicado: (2025)
por: Suri, Anshuman, et al.
Publicado: (2025)
LLM Robustness Leaderboard v1 --Technical report
por: Lefebvre, Pierre Peigné -, et al.
Publicado: (2025)
por: Lefebvre, Pierre Peigné -, et al.
Publicado: (2025)
Ejemplares similares
-
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025) -
A LLM-Based Ranking Method for the Evaluation of Automatic Counter-Narrative Generation
por: Zubiaga, Irune, et al.
Publicado: (2024) -
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
por: Hong, Giwon, et al.
Publicado: (2024) -
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
por: Conde, Javier, et al.
Publicado: (2025) -
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
por: Mayor-Rocher, Marina, et al.
Publicado: (2024)