MultiWorld: Scalable Multi-Agent Multi-View Video World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Haoyu, Yu, Jiwen, Zou, Yingtian, Liu, Xihui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
by: Huang, Kaiyi, et al.
Published: (2024)
by: Huang, Kaiyi, et al.
Published: (2024)
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
by: Liu, Fangfu, et al.
Published: (2026)
by: Liu, Fangfu, et al.
Published: (2026)
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
by: Xiang, Chendong, et al.
Published: (2026)
by: Xiang, Chendong, et al.
Published: (2026)
4Diffusion: Multi-view Video Diffusion Model for 4D Generation
by: Zhang, Haiyu, et al.
Published: (2024)
by: Zhang, Haiyu, et al.
Published: (2024)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
by: Eskandar, George, et al.
Published: (2026)
by: Eskandar, George, et al.
Published: (2026)
Boosting Multi-View Stereo with Depth Foundation Model in the Absence of Real-World Labels
by: Zhu, Jie, et al.
Published: (2025)
by: Zhu, Jie, et al.
Published: (2025)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
by: Zheng, Chaoda, et al.
Published: (2026)
by: Zheng, Chaoda, et al.
Published: (2026)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
GameFactory: Creating New Games with Generative Interactive Videos
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
by: Zou, Jialv, et al.
Published: (2024)
by: Zou, Jialv, et al.
Published: (2024)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
by: Huang, Jiehui, et al.
Published: (2025)
by: Huang, Jiehui, et al.
Published: (2025)
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
by: Yang, Yunhan, et al.
Published: (2025)
by: Yang, Yunhan, et al.
Published: (2025)
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
by: Qin, Yiran, et al.
Published: (2026)
by: Qin, Yiran, et al.
Published: (2026)
ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models
by: Peng, Baorui, et al.
Published: (2026)
by: Peng, Baorui, et al.
Published: (2026)
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
by: Zhu, Shangwen, et al.
Published: (2026)
by: Zhu, Shangwen, et al.
Published: (2026)
WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models
by: Inbasekar, Karthik, et al.
Published: (2026)
by: Inbasekar, Karthik, et al.
Published: (2026)
MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving
by: Wang, Haiguang, et al.
Published: (2025)
by: Wang, Haiguang, et al.
Published: (2025)
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
MV-VTON: Multi-View Virtual Try-On with Diffusion Models
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
by: Ke, Yan, et al.
Published: (2025)
by: Ke, Yan, et al.
Published: (2025)
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
by: HY-World, Team, et al.
Published: (2026)
by: HY-World, Team, et al.
Published: (2026)
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
by: Qiu, Lu, et al.
Published: (2025)
by: Qiu, Lu, et al.
Published: (2025)
Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
coDrawAgents: A Multi-Agent Dialogue Framework for Compositional Image Generation
by: Li, Chunhan, et al.
Published: (2026)
by: Li, Chunhan, et al.
Published: (2026)
Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
by: Xie, Haoyu, et al.
Published: (2026)
by: Xie, Haoyu, et al.
Published: (2026)
Position: Interactive Generative Video as Next-Generation Game Engine
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
GenWorld: Towards Detecting AI-generated Real-world Simulation Videos
by: Chen, Weiliang, et al.
Published: (2025)
by: Chen, Weiliang, et al.
Published: (2025)
Multi-View Attentive Contextualization for Multi-View 3D Object Detection
by: Liu, Xianpeng, et al.
Published: (2024)
by: Liu, Xianpeng, et al.
Published: (2024)
Vivid-ZOO: Multi-View Video Generation with Diffusion Model
by: Li, Bing, et al.
Published: (2024)
by: Li, Bing, et al.
Published: (2024)
SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain
by: Zhou, Jiawei, et al.
Published: (2025)
by: Zhou, Jiawei, et al.
Published: (2025)
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
by: Meng, Fanqing, et al.
Published: (2026)
by: Meng, Fanqing, et al.
Published: (2026)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
World Guidance: World Modeling in Condition Space for Action Generation
by: Su, Yue, et al.
Published: (2026)
by: Su, Yue, et al.
Published: (2026)
AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models
by: Pourkeshavatz, Mozhgan, et al.
Published: (2026)
by: Pourkeshavatz, Mozhgan, et al.
Published: (2026)
Similar Items
-
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024) -
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
by: Huang, Kaiyi, et al.
Published: (2024) -
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
by: Zhou, Yang, et al.
Published: (2025) -
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
by: Liu, Fangfu, et al.
Published: (2026) -
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
by: Xiang, Chendong, et al.
Published: (2026)