LLMs Behind the Scenes: Enabling Narrative Scene Illustration

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Roemmele, Melissa, Chung, John Joon Young, Kim, Taewook, Sun, Yuqian, Calderwood, Alex, Kreminski, Max
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908562145411072
author Roemmele, Melissa
Chung, John Joon Young
Kim, Taewook
Sun, Yuqian
Calderwood, Alex
Kreminski, Max
author_facet Roemmele, Melissa
Chung, John Joon Young
Kim, Taewook
Sun, Yuqian
Calderwood, Alex
Kreminski, Max
contents Generative AI has established the opportunity to readily transform content from one medium to another. This capability is especially powerful for storytelling, where visual illustrations can illuminate a story originally expressed in text. In this paper, we focus on the task of narrative scene illustration, which involves automatically generating an image depicting a scene in a story. Motivated by recent progress on text-to-image models, we consider a pipeline that uses LLMs as an interface for prompting text-to-image models to generate scene illustrations given raw story text. We apply variations of this pipeline to a prominent story corpus in order to synthesize illustrations for scenes in these stories. We conduct a human annotation task to obtain pairwise quality judgments for these illustrations. The outcome of this process is the SceneIllustrations dataset, which we release as a new resource for future work on cross-modal narrative transformation. Through our analysis of this dataset and experiments modeling illustration quality, we demonstrate that LLMs can effectively verbalize scene knowledge implicitly evoked by story text. Moreover, this capability is impactful for generating and evaluating illustrations.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22940
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMs Behind the Scenes: Enabling Narrative Scene Illustration
Roemmele, Melissa
Chung, John Joon Young
Kim, Taewook
Sun, Yuqian
Calderwood, Alex
Kreminski, Max
Computation and Language
Computer Vision and Pattern Recognition
Generative AI has established the opportunity to readily transform content from one medium to another. This capability is especially powerful for storytelling, where visual illustrations can illuminate a story originally expressed in text. In this paper, we focus on the task of narrative scene illustration, which involves automatically generating an image depicting a scene in a story. Motivated by recent progress on text-to-image models, we consider a pipeline that uses LLMs as an interface for prompting text-to-image models to generate scene illustrations given raw story text. We apply variations of this pipeline to a prominent story corpus in order to synthesize illustrations for scenes in these stories. We conduct a human annotation task to obtain pairwise quality judgments for these illustrations. The outcome of this process is the SceneIllustrations dataset, which we release as a new resource for future work on cross-modal narrative transformation. Through our analysis of this dataset and experiments modeling illustration quality, we demonstrate that LLMs can effectively verbalize scene knowledge implicitly evoked by story text. Moreover, this capability is impactful for generating and evaluating illustrations.
title LLMs Behind the Scenes: Enabling Narrative Scene Illustration
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.22940