Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alghallabi, Wafa, Thawkar, Ritesh, Ghaboura, Sara, More, Ketan, Thawakar, Omkar, Cholakkal, Hisham, Khan, Salman, Anwer, Rao Muhammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909623002333184
author Alghallabi, Wafa
Thawkar, Ritesh
Ghaboura, Sara
More, Ketan
Thawakar, Omkar
Cholakkal, Hisham
Khan, Salman
Anwer, Rao Muhammad
author_facet Alghallabi, Wafa
Thawkar, Ritesh
Ghaboura, Sara
More, Ketan
Thawakar, Omkar
Cholakkal, Hisham
Khan, Salman
Anwer, Rao Muhammad
contents Arabic poetry is one of the richest and most culturally rooted forms of expression in the Arabic language, known for its layered meanings, stylistic diversity, and deep historical continuity. Although large language models (LLMs) have demonstrated strong performance across languages and tasks, their ability to understand Arabic poetry remains largely unexplored. In this work, we introduce \emph{Fann or Flop}, the first benchmark designed to assess the comprehension of Arabic poetry by LLMs in 12 historical eras, covering 14 core poetic genres and a variety of metrical forms, from classical structures to contemporary free verse. The benchmark comprises a curated corpus of poems with explanations that assess semantic understanding, metaphor interpretation, prosodic awareness, and cultural context. We argue that poetic comprehension offers a strong indicator for testing how good the LLM understands classical Arabic through Arabic poetry. Unlike surface-level tasks, this domain demands deeper interpretive reasoning and cultural sensitivity. Our evaluation of state-of-the-art LLMs shows that most models struggle with poetic understanding despite strong results on standard Arabic benchmarks. We release "Fann or Flop" along with the evaluation suite as an open-source resource to enable rigorous evaluation and advancement for Arabic language models. Code is available at: https://github.com/mbzuai-oryx/FannOrFlop.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18152
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
Alghallabi, Wafa
Thawkar, Ritesh
Ghaboura, Sara
More, Ketan
Thawakar, Omkar
Cholakkal, Hisham
Khan, Salman
Anwer, Rao Muhammad
Computation and Language
Arabic poetry is one of the richest and most culturally rooted forms of expression in the Arabic language, known for its layered meanings, stylistic diversity, and deep historical continuity. Although large language models (LLMs) have demonstrated strong performance across languages and tasks, their ability to understand Arabic poetry remains largely unexplored. In this work, we introduce \emph{Fann or Flop}, the first benchmark designed to assess the comprehension of Arabic poetry by LLMs in 12 historical eras, covering 14 core poetic genres and a variety of metrical forms, from classical structures to contemporary free verse. The benchmark comprises a curated corpus of poems with explanations that assess semantic understanding, metaphor interpretation, prosodic awareness, and cultural context. We argue that poetic comprehension offers a strong indicator for testing how good the LLM understands classical Arabic through Arabic poetry. Unlike surface-level tasks, this domain demands deeper interpretive reasoning and cultural sensitivity. Our evaluation of state-of-the-art LLMs shows that most models struggle with poetic understanding despite strong results on standard Arabic benchmarks. We release "Fann or Flop" along with the evaluation suite as an open-source resource to enable rigorous evaluation and advancement for Arabic language models. Code is available at: https://github.com/mbzuai-oryx/FannOrFlop.
title Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
topic Computation and Language
url https://arxiv.org/abs/2505.18152