Evaluating the Creativity of LLMs in Persian Literary Text Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tourajmehr, Armin, Modarres, Mohammad Reza, Yaghoobzadeh, Yadollah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911444163887104
author Tourajmehr, Armin
Modarres, Mohammad Reza
Yaghoobzadeh, Yadollah
author_facet Tourajmehr, Armin
Modarres, Mohammad Reza
Yaghoobzadeh, Yadollah
contents Large language models (LLMs) have demonstrated notable creative abilities in generating literary texts, including poetry and short stories. However, prior research has primarily centered on English, with limited exploration of non-English literary traditions and without standardized methods for assessing creativity. In this paper, we evaluate the capacity of LLMs to generate Persian literary text enriched with culturally relevant expressions. We build a dataset of user-generated Persian literary spanning 20 diverse topics and assess model outputs along four creativity dimensions-originality, fluency, flexibility, and elaboration-by adapting the Torrance Tests of Creative Thinking. To reduce evaluation costs, we adopt an LLM as a judge for automated scoring and validate its reliability against human judgments using intraclass correlation coefficients, observing strong agreement. In addition, we analyze the models' ability to understand and employ four core literary devices: simile, metaphor, hyperbole, and antithesis. Our results highlight both the strengths and limitations of LLMs in Persian literary text generation, underscoring the need for further refinement.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18401
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating the Creativity of LLMs in Persian Literary Text Generation
Tourajmehr, Armin
Modarres, Mohammad Reza
Yaghoobzadeh, Yadollah
Computation and Language
Large language models (LLMs) have demonstrated notable creative abilities in generating literary texts, including poetry and short stories. However, prior research has primarily centered on English, with limited exploration of non-English literary traditions and without standardized methods for assessing creativity. In this paper, we evaluate the capacity of LLMs to generate Persian literary text enriched with culturally relevant expressions. We build a dataset of user-generated Persian literary spanning 20 diverse topics and assess model outputs along four creativity dimensions-originality, fluency, flexibility, and elaboration-by adapting the Torrance Tests of Creative Thinking. To reduce evaluation costs, we adopt an LLM as a judge for automated scoring and validate its reliability against human judgments using intraclass correlation coefficients, observing strong agreement. In addition, we analyze the models' ability to understand and employ four core literary devices: simile, metaphor, hyperbole, and antithesis. Our results highlight both the strengths and limitations of LLMs in Persian literary text generation, underscoring the need for further refinement.
title Evaluating the Creativity of LLMs in Persian Literary Text Generation
topic Computation and Language
url https://arxiv.org/abs/2509.18401