Spanish and LLM Benchmarks: is MMLU Lost in Translation?
Fuente:
arXiv
Guardado en:
| Autores principales: | Plaza, Irene, Melero, Nina, del Pozo, Cristina, Conde, Javier, Reviriego, Pedro, Mayor-Rocher, Marina, Grandury, María |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
It's the same but not the same: Do LLMs distinguish Spanish varieties?
por: Mayor-Rocher, Marina, et al.
Publicado: (2025)
por: Mayor-Rocher, Marina, et al.
Publicado: (2025)
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
por: Mayor-Rocher, Marina, et al.
Publicado: (2024)
por: Mayor-Rocher, Marina, et al.
Publicado: (2024)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
por: Awad, Samer, et al.
Publicado: (2026)
por: Awad, Samer, et al.
Publicado: (2026)
The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations
por: Arriaga, Carlos, et al.
Publicado: (2025)
por: Arriaga, Carlos, et al.
Publicado: (2025)
How does fine-tuning improve sensorimotor representations in large language models?
por: Wu, Minghua, et al.
Publicado: (2026)
por: Wu, Minghua, et al.
Publicado: (2026)
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
por: Ferrando, Raquel, et al.
Publicado: (2025)
por: Ferrando, Raquel, et al.
Publicado: (2025)
Are We Done with MMLU?
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
por: Bsharat, Sondos Mahmoud, et al.
Publicado: (2025)
por: Bsharat, Sondos Mahmoud, et al.
Publicado: (2025)
Can ChatGPT Learn to Count Letters?
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Playing with words: Comparing the vocabulary and lexical diversity of ChatGPT and humans
por: Reviriego, Pedro, et al.
Publicado: (2023)
por: Reviriego, Pedro, et al.
Publicado: (2023)
Speed and Conversational Large Language Models: Not All Is About Tokens per Second
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
por: Altakrori, Malik H., et al.
Publicado: (2025)
por: Altakrori, Malik H., et al.
Publicado: (2025)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
por: Ghahroodi, Omid, et al.
Publicado: (2024)
por: Ghahroodi, Omid, et al.
Publicado: (2024)
Concurrent Linguistic Error Detection (CLED): a New Methodology for Error Detection in Large Language Models
por: Zhu, Jinhua, et al.
Publicado: (2024)
por: Zhu, Jinhua, et al.
Publicado: (2024)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
por: KJ, Sankalp, et al.
Publicado: (2025)
por: KJ, Sankalp, et al.
Publicado: (2025)
Open Conversational LLMs do not know most Spanish words
por: Conde, Javier, et al.
Publicado: (2024)
por: Conde, Javier, et al.
Publicado: (2024)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
por: Wang, Wentian, et al.
Publicado: (2024)
por: Wang, Wentian, et al.
Publicado: (2024)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
por: Zhao, Qihao, et al.
Publicado: (2024)
por: Zhao, Qihao, et al.
Publicado: (2024)
Large Language Models and Book Summarization: Reading or Remembering, Which Is Better?
por: Fu, Tairan, et al.
Publicado: (2026)
por: Fu, Tairan, et al.
Publicado: (2026)
Reactor Mk.1 performances: MMLU, HumanEval and BBH test results
por: Dunham, TJ, et al.
Publicado: (2024)
por: Dunham, TJ, et al.
Publicado: (2024)
Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan
por: Chatterjee, Ahan, et al.
Publicado: (2026)
por: Chatterjee, Ahan, et al.
Publicado: (2026)
Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models
por: Zhao, Juntu, et al.
Publicado: (2024)
por: Zhao, Juntu, et al.
Publicado: (2024)
Stochastic Streets: A Walk Through Random LLM Address Generation in four European Cities
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
por: Shukla, Vaibhav, et al.
Publicado: (2026)
por: Shukla, Vaibhav, et al.
Publicado: (2026)
Lost in Translation: The Algorithmic Gap Between LMs and the Brain
por: Tosato, Tommaso, et al.
Publicado: (2024)
por: Tosato, Tommaso, et al.
Publicado: (2024)
Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation
por: Huang, Xu, et al.
Publicado: (2024)
por: Huang, Xu, et al.
Publicado: (2024)
Assessing Latency in ASR Systems: A Methodological Perspective for Real-Time Use
por: Arriaga, Carlos, et al.
Publicado: (2024)
por: Arriaga, Carlos, et al.
Publicado: (2024)
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks
por: Zhao, Justin, et al.
Publicado: (2024)
por: Zhao, Justin, et al.
Publicado: (2024)
Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News
por: López, Alejandro Buitrago, et al.
Publicado: (2026)
por: López, Alejandro Buitrago, et al.
Publicado: (2026)
Improving LLM Abilities in Idiomatic Translation
por: Donthi, Sundesh, et al.
Publicado: (2024)
por: Donthi, Sundesh, et al.
Publicado: (2024)
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
por: Gupta, Aman, et al.
Publicado: (2025)
por: Gupta, Aman, et al.
Publicado: (2025)
What Has Been Lost with Synthetic Evaluation?
por: Gill, Alexander, et al.
Publicado: (2025)
por: Gill, Alexander, et al.
Publicado: (2025)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
por: Kim, Dongjun, et al.
Publicado: (2025)
por: Kim, Dongjun, et al.
Publicado: (2025)
Real-time Spatial Retrieval Augmented Generation for Urban Environments
por: Campo, David Nazareno, et al.
Publicado: (2025)
por: Campo, David Nazareno, et al.
Publicado: (2025)
Understanding the Impact of Artificial Intelligence in Academic Writing: Metadata to the Rescue
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Improving Low-Resource Translation with Dictionary-Guided Fine-Tuning and RL: A Spanish-to-Wayuunaiki Study
por: Mosquera, Manuel, et al.
Publicado: (2025)
por: Mosquera, Manuel, et al.
Publicado: (2025)
LLM-Supported Natural Language to Bash Translation
por: Westenfelder, Finnian, et al.
Publicado: (2025)
por: Westenfelder, Finnian, et al.
Publicado: (2025)
Training language models to be warm and empathetic makes them less reliable and more sycophantic
por: Ibrahim, Lujain, et al.
Publicado: (2025)
por: Ibrahim, Lujain, et al.
Publicado: (2025)
Ejemplares similares
-
It's the same but not the same: Do LLMs distinguish Spanish varieties?
por: Mayor-Rocher, Marina, et al.
Publicado: (2025) -
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
por: Mayor-Rocher, Marina, et al.
Publicado: (2024) -
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025) -
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
por: Conde, Javier, et al.
Publicado: (2025) -
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
por: Awad, Samer, et al.
Publicado: (2026)