Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
Fuente:
arXiv
Guardado en:
| Autores principales: | Mi, Maggie, Villavicencio, Aline, Moosavi, Nafise Sadat |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
por: Mi, Maggie, et al.
Publicado: (2025)
por: Mi, Maggie, et al.
Publicado: (2025)
How to Leverage Digit Embeddings to Represent Numbers?
por: Sivakumar, Jasivan Alex, et al.
Publicado: (2024)
por: Sivakumar, Jasivan Alex, et al.
Publicado: (2024)
Stands to Reason: Investigating the Effect of Reasoning on Idiomaticity Detection
por: Phelps, Dylan, et al.
Publicado: (2025)
por: Phelps, Dylan, et al.
Publicado: (2025)
Sign of the Times: Evaluating the use of Large Language Models for Idiomaticity Detection
por: Phelps, Dylan, et al.
Publicado: (2024)
por: Phelps, Dylan, et al.
Publicado: (2024)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
por: Liu, Yiqi, et al.
Publicado: (2023)
por: Liu, Yiqi, et al.
Publicado: (2023)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
por: Kennedy, Ian W., et al.
Publicado: (2026)
por: Kennedy, Ian W., et al.
Publicado: (2026)
SemEval-2025 Task 1: AdMIRe -- Advancing Multimodal Idiomaticity Representation
por: Pickard, Thomas, et al.
Publicado: (2025)
por: Pickard, Thomas, et al.
Publicado: (2025)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
por: Pandya, Mugdha, et al.
Publicado: (2024)
por: Pandya, Mugdha, et al.
Publicado: (2024)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
por: Shafiei, Mohammadamin, et al.
Publicado: (2025)
por: Shafiei, Mohammadamin, et al.
Publicado: (2025)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
por: Shafiei, Mohammadamin, et al.
Publicado: (2025)
por: Shafiei, Mohammadamin, et al.
Publicado: (2025)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
por: Aghaebe, Favour Yahdii, et al.
Publicado: (2025)
por: Aghaebe, Favour Yahdii, et al.
Publicado: (2025)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
por: Xue, Huiyin, et al.
Publicado: (2025)
por: Xue, Huiyin, et al.
Publicado: (2025)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
por: Pastorino, Valeria, et al.
Publicado: (2024)
por: Pastorino, Valeria, et al.
Publicado: (2024)
Investigating Idiomaticity in Word Representations
por: He, Wei, et al.
Publicado: (2024)
por: He, Wei, et al.
Publicado: (2024)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
por: Permadi, Vynska Amalia, et al.
Publicado: (2026)
por: Permadi, Vynska Amalia, et al.
Publicado: (2026)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
por: Aghaebe, Favour Yahdii, et al.
Publicado: (2026)
por: Aghaebe, Favour Yahdii, et al.
Publicado: (2026)
Enhancing Idiomatic Representation in Multiple Languages via an Adaptive Contrastive Triplet Loss
por: He, Wei, et al.
Publicado: (2024)
por: He, Wei, et al.
Publicado: (2024)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
por: Saffari, Hamidreza, et al.
Publicado: (2024)
por: Saffari, Hamidreza, et al.
Publicado: (2024)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
por: James, Joseph, et al.
Publicado: (2026)
por: James, Joseph, et al.
Publicado: (2026)
Exploring Gender Disparities in Automatic Speech Recognition Technology
por: ElGhazaly, Hend, et al.
Publicado: (2025)
por: ElGhazaly, Hend, et al.
Publicado: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
por: Stacey, Joe, et al.
Publicado: (2026)
por: Stacey, Joe, et al.
Publicado: (2026)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
por: Wang, Xiao, et al.
Publicado: (2025)
por: Wang, Xiao, et al.
Publicado: (2025)
How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
por: Yamaguchi, Atsuki, et al.
Publicado: (2024)
por: Yamaguchi, Atsuki, et al.
Publicado: (2024)
How LLMs Fail to Support Fact-Checking
por: Proma, Adiba Mahbub, et al.
Publicado: (2025)
por: Proma, Adiba Mahbub, et al.
Publicado: (2025)
CLIX: Cross-Lingual Explanations of Idiomatic Expressions
por: Gluck, Aaron, et al.
Publicado: (2025)
por: Gluck, Aaron, et al.
Publicado: (2025)
A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
por: Torunoğlu-Selamet, Dilara, et al.
Publicado: (2026)
por: Torunoğlu-Selamet, Dilara, et al.
Publicado: (2026)
Improving LLM Abilities in Idiomatic Translation
por: Donthi, Sundesh, et al.
Publicado: (2024)
por: Donthi, Sundesh, et al.
Publicado: (2024)
Graph-Assisted Culturally Adaptable Idiomatic Translation for Indic Languages
por: Singh, Pratik Rakesh, et al.
Publicado: (2025)
por: Singh, Pratik Rakesh, et al.
Publicado: (2025)
Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
por: Yamaguchi, Atsuki, et al.
Publicado: (2025)
por: Yamaguchi, Atsuki, et al.
Publicado: (2025)
DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
por: Li, Yiqi, et al.
Publicado: (2025)
por: Li, Yiqi, et al.
Publicado: (2025)
Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers
por: Zeng, Fanqin, et al.
Publicado: (2026)
por: Zeng, Fanqin, et al.
Publicado: (2026)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
por: Roy, Amartya, et al.
Publicado: (2026)
por: Roy, Amartya, et al.
Publicado: (2026)
Span Modeling for Idiomaticity and Figurative Language Detection with Span Contrastive Loss
por: Matheny, Blake, et al.
Publicado: (2026)
por: Matheny, Blake, et al.
Publicado: (2026)
Claim Check-Worthiness Detection: How Well do LLMs Grasp Annotation Guidelines?
por: Majer, Laura, et al.
Publicado: (2024)
por: Majer, Laura, et al.
Publicado: (2024)
Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
por: Yamaguchi, Atsuki, et al.
Publicado: (2026)
por: Yamaguchi, Atsuki, et al.
Publicado: (2026)
Biasless Language Models Learn Unnaturally: How LLMs Fail to Distinguish the Possible from the Impossible
por: Ziv, Imry, et al.
Publicado: (2025)
por: Ziv, Imry, et al.
Publicado: (2025)
When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities
por: Das, Sarmistha, et al.
Publicado: (2026)
por: Das, Sarmistha, et al.
Publicado: (2026)
A Data-Driven Approach to Idiomaticity Based on Experts' Criteria in Theoretical Linguistics
por: Mikhalkova, Elena, et al.
Publicado: (2026)
por: Mikhalkova, Elena, et al.
Publicado: (2026)
IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions
por: Hashiloni, Kai Golan, et al.
Publicado: (2026)
por: Hashiloni, Kai Golan, et al.
Publicado: (2026)
Preferences for Idiomatic Language are Acquired Slowly -- and Forgotten Quickly: A Case Study on Swedish
por: Kunz, Jenny
Publicado: (2026)
por: Kunz, Jenny
Publicado: (2026)
Ejemplares similares
-
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
por: Mi, Maggie, et al.
Publicado: (2025) -
How to Leverage Digit Embeddings to Represent Numbers?
por: Sivakumar, Jasivan Alex, et al.
Publicado: (2024) -
Stands to Reason: Investigating the Effect of Reasoning on Idiomaticity Detection
por: Phelps, Dylan, et al.
Publicado: (2025) -
Sign of the Times: Evaluating the use of Large Language Models for Idiomaticity Detection
por: Phelps, Dylan, et al.
Publicado: (2024) -
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
por: Liu, Yiqi, et al.
Publicado: (2023)