The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909202017943552 |
|---|---|
| author | Shin, Andrew Mori, Yusuke Kaneko, Kunitake |
| author_facet | Shin, Andrew Mori, Yusuke Kaneko, Kunitake |
| contents | Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models are invariably focused on conveying the visual elements of a single scene, and have so far been indifferent to another important potential of the medium, namely a storytelling. In this paper, we examine text-to-video generation from a storytelling perspective, which has been hardly investigated, and make empirical remarks that spotlight the limitations of current text-to-video generation scheme. We also propose an evaluation framework for storytelling aspects of videos, and discuss the potential future directions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_08720 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective Shin, Andrew Mori, Yusuke Kaneko, Kunitake Computer Vision and Pattern Recognition Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models are invariably focused on conveying the visual elements of a single scene, and have so far been indifferent to another important potential of the medium, namely a storytelling. In this paper, we examine text-to-video generation from a storytelling perspective, which has been hardly investigated, and make empirical remarks that spotlight the limitations of current text-to-video generation scheme. We also propose an evaluation framework for storytelling aspects of videos, and discuss the potential future directions. |
| title | The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.08720 |