From Image Captioning to Visual Storytelling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Passadakis, Admitos, Song, Yingjin, Gatt, Albert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915452023734272
author Passadakis, Admitos
Song, Yingjin
Gatt, Albert
author_facet Passadakis, Admitos
Song, Yingjin
Gatt, Albert
contents Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence but also narrative and coherent. The aim of this work is to balance between these aspects, by treating Visual Storytelling as a superset of Image Captioning, an approach quite different compared to most of prior relevant studies. This means that we firstly employ a vision-to-language model for obtaining captions of the input images, and then, these captions are transformed into coherent narratives using language-to-language methods. Our multifarious evaluation shows that integrating captioning and storytelling under a unified framework, has a positive impact on the quality of the produced stories. In addition, compared to numerous previous studies, this approach accelerates training time and makes our framework readily reusable and reproducible by anyone interested. Lastly, we propose a new metric/tool, named ideality, that can be used to simulate how far some results are from an oracle model, and we apply it to emulate human-likeness in visual storytelling.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14045
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Image Captioning to Visual Storytelling
Passadakis, Admitos
Song, Yingjin
Gatt, Albert
Computation and Language
Computer Vision and Pattern Recognition
Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence but also narrative and coherent. The aim of this work is to balance between these aspects, by treating Visual Storytelling as a superset of Image Captioning, an approach quite different compared to most of prior relevant studies. This means that we firstly employ a vision-to-language model for obtaining captions of the input images, and then, these captions are transformed into coherent narratives using language-to-language methods. Our multifarious evaluation shows that integrating captioning and storytelling under a unified framework, has a positive impact on the quality of the produced stories. In addition, compared to numerous previous studies, this approach accelerates training time and makes our framework readily reusable and reproducible by anyone interested. Lastly, we propose a new metric/tool, named ideality, that can be used to simulate how far some results are from an oracle model, and we apply it to emulate human-likeness in visual storytelling.
title From Image Captioning to Visual Storytelling
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.14045