Saved in:
Bibliographic Details
Main Authors: Savva, Georgy, Michel, Oscar, Lu, Daohan, Waiwitlikhit, Suppakit, Meehan, Timothy, Mishra, Dhairya, Poddar, Srivats, Lu, Jack, Xie, Saining
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.22208
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918357081522176
author Savva, Georgy
Michel, Oscar
Lu, Daohan
Waiwitlikhit, Suppakit
Meehan, Timothy
Mishra, Dhairya
Poddar, Srivats
Lu, Jack
Xie, Saining
author_facet Savva, Georgy
Michel, Oscar
Lu, Daohan
Waiwitlikhit, Suppakit
Meehan, Timothy
Mishra, Dhairya
Poddar, Srivats
Lu, Jack
Xie, Saining
contents Existing action-conditioned video generation models (video world models) are limited to single-agent perspectives, failing to capture the multi-agent interactions of real-world environments. We introduce Solaris, a multiplayer video world model that simulates consistent multi-view observations. To enable this, we develop a multiplayer data system designed for robust, continuous, and automated data collection on video games such as Minecraft. Unlike prior platforms built for single-player settings, our system supports coordinated multi-agent interaction and synchronized videos + actions capture. Using this system, we collect 12.64 million multiplayer frames and propose an evaluation framework for multiplayer movement, memory, grounding, building, and view consistency. We train Solaris using a staged pipeline that progressively transitions from single-player to multiplayer modeling, combining bidirectional, causal, and Self Forcing training. In the final stage, we introduce Checkpointed Self Forcing, a memory-efficient Self Forcing variant that enables a longer-horizon teacher. Results show our architecture and training design outperform existing baselines. Through open-sourcing our system and models, we hope to lay the groundwork for a new generation of multi-agent world models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_22208
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Solaris: Building a Multiplayer Video World Model in Minecraft
Savva, Georgy
Michel, Oscar
Lu, Daohan
Waiwitlikhit, Suppakit
Meehan, Timothy
Mishra, Dhairya
Poddar, Srivats
Lu, Jack
Xie, Saining
Computer Vision and Pattern Recognition
Existing action-conditioned video generation models (video world models) are limited to single-agent perspectives, failing to capture the multi-agent interactions of real-world environments. We introduce Solaris, a multiplayer video world model that simulates consistent multi-view observations. To enable this, we develop a multiplayer data system designed for robust, continuous, and automated data collection on video games such as Minecraft. Unlike prior platforms built for single-player settings, our system supports coordinated multi-agent interaction and synchronized videos + actions capture. Using this system, we collect 12.64 million multiplayer frames and propose an evaluation framework for multiplayer movement, memory, grounding, building, and view consistency. We train Solaris using a staged pipeline that progressively transitions from single-player to multiplayer modeling, combining bidirectional, causal, and Self Forcing training. In the final stage, we introduce Checkpointed Self Forcing, a memory-efficient Self Forcing variant that enables a longer-horizon teacher. Results show our architecture and training design outperform existing baselines. Through open-sourcing our system and models, we hope to lay the groundwork for a new generation of multi-agent world models.
title Solaris: Building a Multiplayer Video World Model in Minecraft
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.22208