WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Deshun, Hu, Luhui, Tian, Yu, Li, Zihao, Kelly, Chris, Yang, Bang, Yang, Cindy, Zou, Yuexian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
WorldGPT: Empowering LLM as Multimodal World Model
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
Sora as a World Model? A Complete Survey on Text-to-Video Generation
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024)
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024)
RF-GPT: Teaching AI to See the Wireless World
von: Zou, Hang, et al.
Veröffentlicht: (2026)
von: Zou, Hang, et al.
Veröffentlicht: (2026)
Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
von: Zhu, Zheng, et al.
Veröffentlicht: (2024)
von: Zhu, Zheng, et al.
Veröffentlicht: (2024)
AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era
von: Jiang, Yudong, et al.
Veröffentlicht: (2024)
von: Jiang, Yudong, et al.
Veröffentlicht: (2024)
Open-Sora: Democratizing Efficient Video Production for All
von: Zheng, Zangwei, et al.
Veröffentlicht: (2024)
von: Zheng, Zangwei, et al.
Veröffentlicht: (2024)
Sora OpenAI's Prelude: Social Media Perspectives on Sora OpenAI and the Future of AI Video Generation
von: Mogavi, Reza Hadi, et al.
Veröffentlicht: (2024)
von: Mogavi, Reza Hadi, et al.
Veröffentlicht: (2024)
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
von: Dai, Josef, et al.
Veröffentlicht: (2024)
von: Dai, Josef, et al.
Veröffentlicht: (2024)
VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation
von: Sun, Yongxu, et al.
Veröffentlicht: (2026)
von: Sun, Yongxu, et al.
Veröffentlicht: (2026)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
MultiWorld: Scalable Multi-Agent Multi-View Video World Models
von: Wu, Haoyu, et al.
Veröffentlicht: (2026)
von: Wu, Haoyu, et al.
Veröffentlicht: (2026)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
von: Liu, Zhiheng, et al.
Veröffentlicht: (2025)
von: Liu, Zhiheng, et al.
Veröffentlicht: (2025)
"Sora is Incredible and Scary": Emerging Governance Challenges of Text-to-Video Generative AI Models
von: Zhou, Kyrie Zhixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Kyrie Zhixuan, et al.
Veröffentlicht: (2024)
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
von: Su, Zihan, et al.
Veröffentlicht: (2025)
von: Su, Zihan, et al.
Veröffentlicht: (2025)
Sora-generated video: "A video of a person with a disability"
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
World Models as an Intermediary between Agents and the Real World
von: Yang, Sherry
Veröffentlicht: (2026)
von: Yang, Sherry
Veröffentlicht: (2026)
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
von: Yang, Bang, et al.
Veröffentlicht: (2024)
von: Yang, Bang, et al.
Veröffentlicht: (2024)
World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge
von: Son, Moo Hyun, et al.
Veröffentlicht: (2025)
von: Son, Moo Hyun, et al.
Veröffentlicht: (2025)
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
von: Hu, Xiaotao, et al.
Veröffentlicht: (2024)
von: Hu, Xiaotao, et al.
Veröffentlicht: (2024)
OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
von: Wang, Lening, et al.
Veröffentlicht: (2024)
von: Wang, Lening, et al.
Veröffentlicht: (2024)
Analysing the Public Discourse around OpenAI's Text-To-Video Model 'Sora' using Topic Modeling
von: Parikh, Vatsal Vinay
Veröffentlicht: (2024)
von: Parikh, Vatsal Vinay
Veröffentlicht: (2024)
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2026)
von: Zheng, Minghang, et al.
Veröffentlicht: (2026)
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
von: Yang, Bang, et al.
Veröffentlicht: (2023)
von: Yang, Bang, et al.
Veröffentlicht: (2023)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
Gender Bias in Text-to-Video Generation Models: A case study of Sora
von: Nadeem, Mohammad, et al.
Veröffentlicht: (2024)
von: Nadeem, Mohammad, et al.
Veröffentlicht: (2024)
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
von: Wang, Weijie, et al.
Veröffentlicht: (2026)
von: Wang, Weijie, et al.
Veröffentlicht: (2026)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
von: Chen, Tong, et al.
Veröffentlicht: (2024)
von: Chen, Tong, et al.
Veröffentlicht: (2024)
Open-Sora Plan: Open-Source Large Video Generation Model
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
Syntax-Aware Complex-Valued Neural Machine Translation
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
HPIC: The Habitable Worlds Observatory Preliminary Input Catalog
von: Tuchow, Noah, et al.
Veröffentlicht: (2024)
von: Tuchow, Noah, et al.
Veröffentlicht: (2024)
Técnica para dibujar crisantemo / Wang Deshun
von: Wang Deshun
von: Wang Deshun
Técnica para dibujar peonía / Wang Deshun
von: Wang Deshun
von: Wang Deshun
MAPF-World: Action World Model for Multi-Agent Path Finding
von: Yang, Zhanjiang, et al.
Veröffentlicht: (2025)
von: Yang, Zhanjiang, et al.
Veröffentlicht: (2025)
Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
APISR: Anime Production Inspired Real-World Anime Super-Resolution
von: Wang, Boyang, et al.
Veröffentlicht: (2024)
von: Wang, Boyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
von: Kelly, Chris, et al.
Veröffentlicht: (2024) -
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
von: Kelly, Chris, et al.
Veröffentlicht: (2024) -
WorldGPT: Empowering LLM as Multimodal World Model
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024) -
Sora as a World Model? A Complete Survey on Text-to-Video Generation
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024) -
RF-GPT: Teaching AI to See the Wireless World
von: Zou, Hang, et al.
Veröffentlicht: (2026)