Sora as a World Model? A Complete Survey on Text-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Puspitasari, Fachrina Dewi, Zhang, Chaoning, Cho, Joseph, Haider, Adnan, Eman, Noor Ul, Amin, Omer, Mankowski, Alexis, Umair, Muhammad, Zheng, Jingyao, Zheng, Sheng, Lee, Lik-Hang, Qin, Caiyan, Kim, Tae-Ho, Hong, Choong Seon, Yang, Yang, Shen, Heng Tao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917215451742208
author Puspitasari, Fachrina Dewi
Zhang, Chaoning
Cho, Joseph
Haider, Adnan
Eman, Noor Ul
Amin, Omer
Mankowski, Alexis
Umair, Muhammad
Zheng, Jingyao
Zheng, Sheng
Lee, Lik-Hang
Qin, Caiyan
Kim, Tae-Ho
Hong, Choong Seon
Yang, Yang
Shen, Heng Tao
author_facet Puspitasari, Fachrina Dewi
Zhang, Chaoning
Cho, Joseph
Haider, Adnan
Eman, Noor Ul
Amin, Omer
Mankowski, Alexis
Umair, Muhammad
Zheng, Jingyao
Zheng, Sheng
Lee, Lik-Hang
Qin, Caiyan
Kim, Tae-Ho
Hong, Choong Seon
Yang, Yang
Shen, Heng Tao
contents The evolution of video generation from text, from animating MNIST to simulating the world with Sora, has progressed at a breakneck speed. Here, we systematically discuss how far text-to-video generation technology supports essential requirements in world modeling. We curate 250+ studies on text-based video synthesis and world modeling. We then observe that recent models increasingly support spatial, action, and strategic intelligences in world modeling through adherence to completeness, consistency, invention, as well as human interaction and control. We conclude that text-to-video generation is adept at world modeling, although homework in several aspects, such as the diversity-consistency trade-offs, remains to be addressed.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05131
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sora as a World Model? A Complete Survey on Text-to-Video Generation
Puspitasari, Fachrina Dewi
Zhang, Chaoning
Cho, Joseph
Haider, Adnan
Eman, Noor Ul
Amin, Omer
Mankowski, Alexis
Umair, Muhammad
Zheng, Jingyao
Zheng, Sheng
Lee, Lik-Hang
Qin, Caiyan
Kim, Tae-Ho
Hong, Choong Seon
Yang, Yang
Shen, Heng Tao
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.10
The evolution of video generation from text, from animating MNIST to simulating the world with Sora, has progressed at a breakneck speed. Here, we systematically discuss how far text-to-video generation technology supports essential requirements in world modeling. We curate 250+ studies on text-based video synthesis and world modeling. We then observe that recent models increasingly support spatial, action, and strategic intelligences in world modeling through adherence to completeness, consistency, invention, as well as human interaction and control. We conclude that text-to-video generation is adept at world modeling, although homework in several aspects, such as the diversity-consistency trade-offs, remains to be addressed.
title Sora as a World Model? A Complete Survey on Text-to-Video Generation
topic Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.10
url https://arxiv.org/abs/2403.05131