Solaris: Building a Multiplayer Video World Model in Minecraft
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Savva, Georgy, Michel, Oscar, Lu, Daohan, Waiwitlikhit, Suppakit, Meehan, Timothy, Mishra, Dhairya, Poddar, Srivats, Lu, Jack, Xie, Saining |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Scaling Up 3D Gaussian Splatting Training
von: Zhao, Hexu, et al.
Veröffentlicht: (2024)
von: Zhao, Hexu, et al.
Veröffentlicht: (2024)
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
von: Li, Chenyu, et al.
Veröffentlicht: (2025)
von: Li, Chenyu, et al.
Veröffentlicht: (2025)
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
von: Zhang, Lechao, et al.
Veröffentlicht: (2026)
von: Zhang, Lechao, et al.
Veröffentlicht: (2026)
iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
von: Peng, Weikun, et al.
Veröffentlicht: (2025)
von: Peng, Weikun, et al.
Veröffentlicht: (2025)
Cambrian-S: Towards Spatial Supersensing in Video
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)
Fast Encoding and Decoding for Implicit Video Representation
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Reinforcement Learning Friendly Vision-Language Model for Minecraft
von: Jiang, Haobin, et al.
Veröffentlicht: (2023)
von: Jiang, Haobin, et al.
Veröffentlicht: (2023)
EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates
von: Peng, Weikun, et al.
Veröffentlicht: (2026)
von: Peng, Weikun, et al.
Veröffentlicht: (2026)
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
von: Guo, Junliang, et al.
Veröffentlicht: (2025)
von: Guo, Junliang, et al.
Veröffentlicht: (2025)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines
von: Po, Ryan, et al.
Veröffentlicht: (2026)
von: Po, Ryan, et al.
Veröffentlicht: (2026)
Self-Refining Video Sampling
von: Jang, Sangwon, et al.
Veröffentlicht: (2026)
von: Jang, Sangwon, et al.
Veröffentlicht: (2026)
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
von: Qin, Yiran, et al.
Veröffentlicht: (2023)
von: Qin, Yiran, et al.
Veröffentlicht: (2023)
Cambrian-P: Pose-Grounded Video Understanding
von: Yang, Jihan, et al.
Veröffentlicht: (2026)
von: Yang, Jihan, et al.
Veröffentlicht: (2026)
Trustless Audits without Revealing Data or Models
von: Waiwitlikhit, Suppakit, et al.
Veröffentlicht: (2024)
von: Waiwitlikhit, Suppakit, et al.
Veröffentlicht: (2024)
Transition Matching Distillation for Fast Video Generation
von: Nie, Weili, et al.
Veröffentlicht: (2026)
von: Nie, Weili, et al.
Veröffentlicht: (2026)
WildSmoke: Ready-to-Use Dynamic 3D Smoke Assets from a Single Video in the Wild
von: Liu, Yuqiu, et al.
Veröffentlicht: (2025)
von: Liu, Yuqiu, et al.
Veröffentlicht: (2025)
STEVE Series: Step-by-Step Construction of Agent Systems in Minecraft
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
Hawk: Learning to Understand Open-World Video Anomalies
von: Tang, Jiaqi, et al.
Veröffentlicht: (2024)
von: Tang, Jiaqi, et al.
Veröffentlicht: (2024)
Flow Map Distillation Without Data
von: Tong, Shangyuan, et al.
Veröffentlicht: (2025)
von: Tong, Shangyuan, et al.
Veröffentlicht: (2025)
Diffusion Transformers with Representation Autoencoders
von: Zheng, Boyang, et al.
Veröffentlicht: (2025)
von: Zheng, Boyang, et al.
Veröffentlicht: (2025)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
von: Chen, Xinlei, et al.
Veröffentlicht: (2024)
von: Chen, Xinlei, et al.
Veröffentlicht: (2024)
Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes
von: Merino, Tim, et al.
Veröffentlicht: (2026)
von: Merino, Tim, et al.
Veröffentlicht: (2026)
Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
von: Huang, Junchao, et al.
Veröffentlicht: (2025)
von: Huang, Junchao, et al.
Veröffentlicht: (2025)
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
von: Tang, Bingda, et al.
Veröffentlicht: (2025)
von: Tang, Bingda, et al.
Veröffentlicht: (2025)
Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
von: Lu, Jiacheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2026)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024)
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024)
OptiWorld: Optimal Control for Video World Generation under Physical Constraints
von: Yuan, Yu, et al.
Veröffentlicht: (2026)
von: Yuan, Yu, et al.
Veröffentlicht: (2026)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
Survey on Modeling of Human-made Articulated Objects
von: Liu, Jiayi, et al.
Veröffentlicht: (2024)
von: Liu, Jiayi, et al.
Veröffentlicht: (2024)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Image Sculpting: Precise Object Editing with 3D Geometry Control
von: Yenphraphai, Jiraphon, et al.
Veröffentlicht: (2024)
von: Yenphraphai, Jiraphon, et al.
Veröffentlicht: (2024)
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
von: Chen, Jiacheng, et al.
Veröffentlicht: (2025)
von: Chen, Jiacheng, et al.
Veröffentlicht: (2025)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
von: Chai, Wenhao, et al.
Veröffentlicht: (2024)
von: Chai, Wenhao, et al.
Veröffentlicht: (2024)
WorldSimBench: Towards Video Generation Models as World Simulators
von: Qin, Yiran, et al.
Veröffentlicht: (2024)
von: Qin, Yiran, et al.
Veröffentlicht: (2024)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Text-to-3D Shape Generation
von: Lee, Han-Hung, et al.
Veröffentlicht: (2024)
von: Lee, Han-Hung, et al.
Veröffentlicht: (2024)
ALIVE: Animate Your World with Lifelike Audio-Video Generation
von: Guo, Ying, et al.
Veröffentlicht: (2026)
von: Guo, Ying, et al.
Veröffentlicht: (2026)
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
von: Ni, Jingcheng, et al.
Veröffentlicht: (2025)
von: Ni, Jingcheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On Scaling Up 3D Gaussian Splatting Training
von: Zhao, Hexu, et al.
Veröffentlicht: (2024) -
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
von: Li, Chenyu, et al.
Veröffentlicht: (2025) -
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
von: Zhang, Lechao, et al.
Veröffentlicht: (2026) -
iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
von: Peng, Weikun, et al.
Veröffentlicht: (2025) -
Cambrian-S: Towards Spatial Supersensing in Video
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)