WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Deshun, Hu, Luhui, Tian, Yu, Li, Zihao, Kelly, Chris, Yang, Bang, Yang, Cindy, Zou, Yuexian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
WorldGPT: Empowering LLM as Multimodal World Model
by: Ge, Zhiqi, et al.
Published: (2024)
by: Ge, Zhiqi, et al.
Published: (2024)
Sora as a World Model? A Complete Survey on Text-to-Video Generation
by: Puspitasari, Fachrina Dewi, et al.
Published: (2024)
by: Puspitasari, Fachrina Dewi, et al.
Published: (2024)
RF-GPT: Teaching AI to See the Wireless World
by: Zou, Hang, et al.
Published: (2026)
by: Zou, Hang, et al.
Published: (2026)
Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
by: Zhu, Zheng, et al.
Published: (2024)
by: Zhu, Zheng, et al.
Published: (2024)
AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era
by: Jiang, Yudong, et al.
Published: (2024)
by: Jiang, Yudong, et al.
Published: (2024)
Open-Sora: Democratizing Efficient Video Production for All
by: Zheng, Zangwei, et al.
Published: (2024)
by: Zheng, Zangwei, et al.
Published: (2024)
Sora OpenAI's Prelude: Social Media Perspectives on Sora OpenAI and the Future of AI Video Generation
by: Mogavi, Reza Hadi, et al.
Published: (2024)
by: Mogavi, Reza Hadi, et al.
Published: (2024)
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
by: Dai, Josef, et al.
Published: (2024)
by: Dai, Josef, et al.
Published: (2024)
VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation
by: Sun, Yongxu, et al.
Published: (2026)
by: Sun, Yongxu, et al.
Published: (2026)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
by: Xin, Yifei, et al.
Published: (2023)
by: Xin, Yifei, et al.
Published: (2023)
MultiWorld: Scalable Multi-Agent Multi-View Video World Models
by: Wu, Haoyu, et al.
Published: (2026)
by: Wu, Haoyu, et al.
Published: (2026)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
"Sora is Incredible and Scary": Emerging Governance Challenges of Text-to-Video Generative AI Models
by: Zhou, Kyrie Zhixuan, et al.
Published: (2024)
by: Zhou, Kyrie Zhixuan, et al.
Published: (2024)
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
by: Su, Zihan, et al.
Published: (2025)
by: Su, Zihan, et al.
Published: (2025)
Sora-generated video: "A video of a person with a disability"
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
World Models as an Intermediary between Agents and the Real World
by: Yang, Sherry
Published: (2026)
by: Yang, Sherry
Published: (2026)
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
by: Yang, Bang, et al.
Published: (2024)
by: Yang, Bang, et al.
Published: (2024)
World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge
by: Son, Moo Hyun, et al.
Published: (2025)
by: Son, Moo Hyun, et al.
Published: (2025)
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
by: Hu, Xiaotao, et al.
Published: (2024)
by: Hu, Xiaotao, et al.
Published: (2024)
OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
by: Wang, Lening, et al.
Published: (2024)
by: Wang, Lening, et al.
Published: (2024)
Analysing the Public Discourse around OpenAI's Text-To-Video Model 'Sora' using Topic Modeling
by: Parikh, Vatsal Vinay
Published: (2024)
by: Parikh, Vatsal Vinay
Published: (2024)
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
by: Zheng, Minghang, et al.
Published: (2026)
by: Zheng, Minghang, et al.
Published: (2026)
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
by: Yang, Bang, et al.
Published: (2023)
by: Yang, Bang, et al.
Published: (2023)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
by: Chu, Zhixuan, et al.
Published: (2024)
by: Chu, Zhixuan, et al.
Published: (2024)
Gender Bias in Text-to-Video Generation Models: A case study of Sora
by: Nadeem, Mohammad, et al.
Published: (2024)
by: Nadeem, Mohammad, et al.
Published: (2024)
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
Open-Sora Plan: Open-Source Large Video Generation Model
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Syntax-Aware Complex-Valued Neural Machine Translation
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
HPIC: The Habitable Worlds Observatory Preliminary Input Catalog
by: Tuchow, Noah, et al.
Published: (2024)
by: Tuchow, Noah, et al.
Published: (2024)
Técnica para dibujar crisantemo / Wang Deshun
by: Wang Deshun
by: Wang Deshun
Técnica para dibujar peonía / Wang Deshun
by: Wang Deshun
by: Wang Deshun
MAPF-World: Action World Model for Multi-Agent Path Finding
by: Yang, Zhanjiang, et al.
Published: (2025)
by: Yang, Zhanjiang, et al.
Published: (2025)
Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
by: Qiao, Xiangshuo, et al.
Published: (2024)
by: Qiao, Xiangshuo, et al.
Published: (2024)
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
APISR: Anime Production Inspired Real-World Anime Super-Resolution
by: Wang, Boyang, et al.
Published: (2024)
by: Wang, Boyang, et al.
Published: (2024)
Similar Items
-
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
by: Kelly, Chris, et al.
Published: (2024) -
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024) -
WorldGPT: Empowering LLM as Multimodal World Model
by: Ge, Zhiqi, et al.
Published: (2024) -
Sora as a World Model? A Complete Survey on Text-to-Video Generation
by: Puspitasari, Fachrina Dewi, et al.
Published: (2024) -
RF-GPT: Teaching AI to See the Wireless World
by: Zou, Hang, et al.
Published: (2026)