Guardado en:
| Autores principales: | Mayor-Rocher, Marina, Pozo, Cristina, Melero, Nina, Martínez, Gonzalo, Grandury, María, Reviriego, Pedro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2504.20049 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
por: Plaza, Irene, et al.
Publicado: (2024)
por: Plaza, Irene, et al.
Publicado: (2024)
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
por: Mayor-Rocher, Marina, et al.
Publicado: (2024)
por: Mayor-Rocher, Marina, et al.
Publicado: (2024)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
Open Conversational LLMs do not know most Spanish words
por: Conde, Javier, et al.
Publicado: (2024)
por: Conde, Javier, et al.
Publicado: (2024)
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Do LLMs exhibit the same commonsense capabilities across languages?
por: Martínez-Murillo, Ivan, et al.
Publicado: (2025)
por: Martínez-Murillo, Ivan, et al.
Publicado: (2025)
Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Beware of Words: Evaluating the Lexical Diversity of Conversational LLMs using ChatGPT as Case Study
por: Martínez, Gonzalo, et al.
Publicado: (2024)
por: Martínez, Gonzalo, et al.
Publicado: (2024)
Large Language Models and Book Summarization: Reading or Remembering, Which Is Better?
por: Fu, Tairan, et al.
Publicado: (2026)
por: Fu, Tairan, et al.
Publicado: (2026)
The #Somos600M Project: Generating NLP resources that represent the diversity of the languages from LATAM, the Caribbean, and Spain
por: Grandury, María
Publicado: (2024)
por: Grandury, María
Publicado: (2024)
Why Do Large Language Models (LLMs) Struggle to Count Letters?
por: Fu, Tairan, et al.
Publicado: (2024)
por: Fu, Tairan, et al.
Publicado: (2024)
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
por: Ferrando, Raquel, et al.
Publicado: (2025)
por: Ferrando, Raquel, et al.
Publicado: (2025)
Text Difficulty Study: Do machines behave the same as humans regarding text difficulty?
por: Chen, Bowen, et al.
Publicado: (2022)
por: Chen, Bowen, et al.
Publicado: (2022)
The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations
por: Arriaga, Carlos, et al.
Publicado: (2025)
por: Arriaga, Carlos, et al.
Publicado: (2025)
To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times
por: Clark, Thomas Hikaru, et al.
Publicado: (2026)
por: Clark, Thomas Hikaru, et al.
Publicado: (2026)
LLMs can hide text in other text of the same length
por: Norelli, Antonio, et al.
Publicado: (2025)
por: Norelli, Antonio, et al.
Publicado: (2025)
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
por: Grandury, María, et al.
Publicado: (2025)
por: Grandury, María, et al.
Publicado: (2025)
Using large language models to estimate features of multi-word expressions: Concreteness, valence, arousal
por: Martínez, Gonzalo, et al.
Publicado: (2024)
por: Martínez, Gonzalo, et al.
Publicado: (2024)
Can ChatGPT Learn to Count Letters?
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Playing with words: Comparing the vocabulary and lexical diversity of ChatGPT and humans
por: Reviriego, Pedro, et al.
Publicado: (2023)
por: Reviriego, Pedro, et al.
Publicado: (2023)
Does Burrows' Delta really confirm that Rowling and Galbraith are the same author?
por: Orekhov, Boris
Publicado: (2024)
por: Orekhov, Boris
Publicado: (2024)
Different types of syntactic agreement recruit the same units within large language models
por: Kryvosheieva, Daria, et al.
Publicado: (2025)
por: Kryvosheieva, Daria, et al.
Publicado: (2025)
Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models
por: Martínez, Gonzalo, et al.
Publicado: (2023)
por: Martínez, Gonzalo, et al.
Publicado: (2023)
Verifying Graph Algorithms in Separation Logic: A Case for an Algebraic Approach (Extended Version)
por: Grandury, Marcos, et al.
Publicado: (2025)
por: Grandury, Marcos, et al.
Publicado: (2025)
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
por: Awad, Samer, et al.
Publicado: (2026)
por: Awad, Samer, et al.
Publicado: (2026)
How does fine-tuning improve sensorimotor representations in large language models?
por: Wu, Minghua, et al.
Publicado: (2026)
por: Wu, Minghua, et al.
Publicado: (2026)
Are we describing the same sound? An analysis of word embedding spaces of expressive piano performance
por: Peter, Silvan David, et al.
Publicado: (2023)
por: Peter, Silvan David, et al.
Publicado: (2023)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
por: Stewart, Ian, et al.
Publicado: (2024)
por: Stewart, Ian, et al.
Publicado: (2024)
Do LLMs Know What Luxembourgish Borrows? Probing Lexical Neology in Low-Resource Multilingual Models
por: Hosseini-Kivanani, Nina
Publicado: (2026)
por: Hosseini-Kivanani, Nina
Publicado: (2026)
The power of Prompts: Evaluating and Mitigating Gender Bias in MT with LLMs
por: Sant, Aleix, et al.
Publicado: (2024)
por: Sant, Aleix, et al.
Publicado: (2024)
Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
Speed and Conversational Large Language Models: Not All Is About Tokens per Second
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
On convergence empirics: same evidence for Spanish regions
por: Ana Lamo
Publicado: (2000)
por: Ana Lamo
Publicado: (2000)
Concurrent Linguistic Error Detection (CLED): a New Methodology for Error Detection in Large Language Models
por: Zhu, Jinhua, et al.
Publicado: (2024)
por: Zhu, Jinhua, et al.
Publicado: (2024)
Training language models to be warm and empathetic makes them less reliable and more sycophantic
por: Ibrahim, Lujain, et al.
Publicado: (2025)
por: Ibrahim, Lujain, et al.
Publicado: (2025)
Into the crossfire: evaluating the use of a language model to crowdsource gun violence reports
por: Belisario, Adriano, et al.
Publicado: (2024)
por: Belisario, Adriano, et al.
Publicado: (2024)
Overview of ADoBo at IberLEF 2025: Automatic Detection of Anglicisms in Spanish
por: Alvarez-Mellado, Elena, et al.
Publicado: (2025)
por: Alvarez-Mellado, Elena, et al.
Publicado: (2025)
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
por: Hafner, Franziska Sofia, et al.
Publicado: (2025)
por: Hafner, Franziska Sofia, et al.
Publicado: (2025)
Alignment Drift in CEFR-prompted LLMs for Interactive Spanish Tutoring
por: Almasi, Mina, et al.
Publicado: (2025)
por: Almasi, Mina, et al.
Publicado: (2025)
Digital Linguistic Bias in Spanish: Evidence from Lexical Variation in LLMs
por: Kawasaki, Yoshifumi
Publicado: (2026)
por: Kawasaki, Yoshifumi
Publicado: (2026)
Ejemplares similares
-
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
por: Plaza, Irene, et al.
Publicado: (2024) -
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
por: Mayor-Rocher, Marina, et al.
Publicado: (2024) -
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025) -
Open Conversational LLMs do not know most Spanish words
por: Conde, Javier, et al.
Publicado: (2024) -
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
por: Conde, Javier, et al.
Publicado: (2025)