From Sora What We Can See: A Survey of Text-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Rui, Zhang, Yumin, Shah, Tejal, Sun, Jiahao, Zhang, Shuoying, Li, Wenqi, Duan, Haoran, Wei, Bo, Ranjan, Rajiv
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910450670632960
author Sun, Rui
Zhang, Yumin
Shah, Tejal
Sun, Jiahao
Zhang, Shuoying
Li, Wenqi
Duan, Haoran
Wei, Bo
Ranjan, Rajiv
author_facet Sun, Rui
Zhang, Yumin
Shah, Tejal
Sun, Jiahao
Zhang, Shuoying
Li, Wenqi
Duan, Haoran
Wei, Bo
Ranjan, Rajiv
contents With impressive achievements made, artificial intelligence is on the path forward to artificial general intelligence. Sora, developed by OpenAI, which is capable of minute-level world-simulative abilities can be considered as a milestone on this developmental path. However, despite its notable successes, Sora still encounters various obstacles that need to be resolved. In this survey, we embark from the perspective of disassembling Sora in text-to-video generation, and conducting a comprehensive review of literature, trying to answer the question, \textit{From Sora What We Can See}. Specifically, after basic preliminaries regarding the general algorithms are introduced, the literature is categorized from three mutually perpendicular dimensions: evolutionary generators, excellent pursuit, and realistic panorama. Subsequently, the widely used datasets and metrics are organized in detail. Last but more importantly, we identify several challenges and open problems in this domain and propose potential future directions for research and development.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10674
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Sora What We Can See: A Survey of Text-to-Video Generation
Sun, Rui
Zhang, Yumin
Shah, Tejal
Sun, Jiahao
Zhang, Shuoying
Li, Wenqi
Duan, Haoran
Wei, Bo
Ranjan, Rajiv
Computer Vision and Pattern Recognition
Artificial Intelligence
With impressive achievements made, artificial intelligence is on the path forward to artificial general intelligence. Sora, developed by OpenAI, which is capable of minute-level world-simulative abilities can be considered as a milestone on this developmental path. However, despite its notable successes, Sora still encounters various obstacles that need to be resolved. In this survey, we embark from the perspective of disassembling Sora in text-to-video generation, and conducting a comprehensive review of literature, trying to answer the question, \textit{From Sora What We Can See}. Specifically, after basic preliminaries regarding the general algorithms are introduced, the literature is categorized from three mutually perpendicular dimensions: evolutionary generators, excellent pursuit, and realistic panorama. Subsequently, the widely used datasets and metrics are organized in detail. Last but more importantly, we identify several challenges and open problems in this domain and propose potential future directions for research and development.
title From Sora What We Can See: A Survey of Text-to-Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.10674