Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Song, Yingjin, Paperno, Denis, Gatt, Albert
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929456236462080
author Song, Yingjin
Paperno, Denis
Gatt, Albert
author_facet Song, Yingjin
Paperno, Denis
Gatt, Albert
contents Visual storytelling systems generate multi-sentence stories from image sequences. In this task, capturing contextual information and bridging visual variation bring additional challenges. We propose a simple yet effective framework that leverages the generalization capabilities of pretrained foundation models, only training a lightweight vision-language mapping network to connect modalities, while incorporating context to enhance coherence. We introduce a multimodal contrastive objective that also improves visual relevance and story informativeness. Extensive experimental results, across both automatic metrics and human evaluations, demonstrate that the stories generated by our framework are diverse, coherent, informative, and interesting.
format Preprint
id arxiv_https___arxiv_org_abs_2408_06259
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
Song, Yingjin
Paperno, Denis
Gatt, Albert
Computation and Language
Computer Vision and Pattern Recognition
Visual storytelling systems generate multi-sentence stories from image sequences. In this task, capturing contextual information and bridging visual variation bring additional challenges. We propose a simple yet effective framework that leverages the generalization capabilities of pretrained foundation models, only training a lightweight vision-language mapping network to connect modalities, while incorporating context to enhance coherence. We introduce a multimodal contrastive objective that also improves visual relevance and story informativeness. Extensive experimental results, across both automatic metrics and human evaluations, demonstrate that the stories generated by our framework are diverse, coherent, informative, and interesting.
title Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.06259