VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gado, Mohamed, Taliee, Towhid, Memon, Muhammad, Ignatov, Dmitry, Timofte, Radu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912421961007104
author Gado, Mohamed
Taliee, Towhid
Memon, Muhammad
Ignatov, Dmitry
Timofte, Radu
author_facet Gado, Mohamed
Taliee, Towhid
Memon, Muhammad
Ignatov, Dmitry
Timofte, Radu
contents Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in multimodal models, specifically adapting transformer-based architectures and large multimodal models, for the visual storytelling task. Leveraging the large-scale Visual Storytelling (VIST) dataset, our VIST-GPT model produces visually grounded, contextually appropriate narratives. We address the limitations of traditional evaluation metrics, such as BLEU, METEOR, ROUGE, and CIDEr, which are not suitable for this task. Instead, we utilize RoViST and GROOVIST, novel reference-free metrics designed to assess visual storytelling, focusing on visual grounding, coherence, and non-redundancy. These metrics provide a more nuanced evaluation of narrative quality, aligning closely with human judgment.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19267
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
Gado, Mohamed
Taliee, Towhid
Memon, Muhammad
Ignatov, Dmitry
Timofte, Radu
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in multimodal models, specifically adapting transformer-based architectures and large multimodal models, for the visual storytelling task. Leveraging the large-scale Visual Storytelling (VIST) dataset, our VIST-GPT model produces visually grounded, contextually appropriate narratives. We address the limitations of traditional evaluation metrics, such as BLEU, METEOR, ROUGE, and CIDEr, which are not suitable for this task. Instead, we utilize RoViST and GROOVIST, novel reference-free metrics designed to assess visual storytelling, focusing on visual grounding, coherence, and non-redundancy. These metrics provide a more nuanced evaluation of narrative quality, aligning closely with human judgment.
title VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2504.19267