The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shin, Andrew, Mori, Yusuke, Kaneko, Kunitake
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909202017943552
author Shin, Andrew
Mori, Yusuke
Kaneko, Kunitake
author_facet Shin, Andrew
Mori, Yusuke
Kaneko, Kunitake
contents Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models are invariably focused on conveying the visual elements of a single scene, and have so far been indifferent to another important potential of the medium, namely a storytelling. In this paper, we examine text-to-video generation from a storytelling perspective, which has been hardly investigated, and make empirical remarks that spotlight the limitations of current text-to-video generation scheme. We also propose an evaluation framework for storytelling aspects of videos, and discuss the potential future directions.
format Preprint
id arxiv_https___arxiv_org_abs_2405_08720
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective
Shin, Andrew
Mori, Yusuke
Kaneko, Kunitake
Computer Vision and Pattern Recognition
Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models are invariably focused on conveying the visual elements of a single scene, and have so far been indifferent to another important potential of the medium, namely a storytelling. In this paper, we examine text-to-video generation from a storytelling perspective, which has been hardly investigated, and make empirical remarks that spotlight the limitations of current text-to-video generation scheme. We also propose an evaluation framework for storytelling aspects of videos, and discuss the potential future directions.
title The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.08720