When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sakhawat, Adib, Parveen, Shamim Ara, Amin, Md Ruhul, Mahmud, Shamim Al, Islam, Md Saiful, Khatun, Tahera
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908833448722432
author Sakhawat, Adib
Parveen, Shamim Ara
Amin, Md Ruhul
Mahmud, Shamim Al
Islam, Md Saiful
Khatun, Tahera
author_facet Sakhawat, Adib
Parveen, Shamim Ara
Amin, Md Ruhul
Mahmud, Shamim Al
Islam, Md Saiful
Khatun, Tahera
contents Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce a new idiom dataset, a large-scale, culturally-grounded corpus of 10,361 Bengali idioms. Each idiom is annotated under a comprehensive 19-field schema, established and refined through a deliberative expert consensus process, that captures its semantic, syntactic, cultural, and religious dimensions, providing a rich, structured resource for computational linguistics. To establish a robust benchmark for Bangla figurative language understanding, we evaluate 30 state-of-the-art multilingual and instruction-tuned LLMs on the task of inferring figurative meaning. Our results reveal a critical performance gap, with no model surpassing 50% accuracy, a stark contrast to significantly higher human performance (83.4%). This underscores the limitations of existing models in cross-linguistic and cultural reasoning. By releasing the new idiom dataset and benchmark, we provide foundational infrastructure for advancing figurative language understanding and cultural grounding in LLMs for Bengali and other low-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2602_12921
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms
Sakhawat, Adib
Parveen, Shamim Ara
Amin, Md Ruhul
Mahmud, Shamim Al
Islam, Md Saiful
Khatun, Tahera
Computation and Language
Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce a new idiom dataset, a large-scale, culturally-grounded corpus of 10,361 Bengali idioms. Each idiom is annotated under a comprehensive 19-field schema, established and refined through a deliberative expert consensus process, that captures its semantic, syntactic, cultural, and religious dimensions, providing a rich, structured resource for computational linguistics. To establish a robust benchmark for Bangla figurative language understanding, we evaluate 30 state-of-the-art multilingual and instruction-tuned LLMs on the task of inferring figurative meaning. Our results reveal a critical performance gap, with no model surpassing 50% accuracy, a stark contrast to significantly higher human performance (83.4%). This underscores the limitations of existing models in cross-linguistic and cultural reasoning. By releasing the new idiom dataset and benchmark, we provide foundational infrastructure for advancing figurative language understanding and cultural grounding in LLMs for Bengali and other low-resource languages.
title When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms
topic Computation and Language
url https://arxiv.org/abs/2602.12921