Saved in:
| Main Authors: | Savva, Georgy, Michel, Oscar, Lu, Daohan, Waiwitlikhit, Suppakit, Meehan, Timothy, Mishra, Dhairya, Poddar, Srivats, Lu, Jack, Xie, Saining |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.22208 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trustless Audits without Revealing Data or Models
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
On Scaling Up 3D Gaussian Splatting Training
by: Zhao, Hexu, et al.
Published: (2024)
by: Zhao, Hexu, et al.
Published: (2024)
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
by: Li, Chenyu, et al.
Published: (2025)
by: Li, Chenyu, et al.
Published: (2025)
Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?
by: Hu, Chuxuan, et al.
Published: (2025)
by: Hu, Chuxuan, et al.
Published: (2025)
iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
by: Peng, Weikun, et al.
Published: (2025)
by: Peng, Weikun, et al.
Published: (2025)
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
by: Zhang, Lechao, et al.
Published: (2026)
by: Zhang, Lechao, et al.
Published: (2026)
Reinforcement Learning Friendly Vision-Language Model for Minecraft
by: Jiang, Haobin, et al.
Published: (2023)
by: Jiang, Haobin, et al.
Published: (2023)
Cambrian-S: Towards Spatial Supersensing in Video
by: Yang, Shusheng, et al.
Published: (2025)
by: Yang, Shusheng, et al.
Published: (2025)
Fast Encoding and Decoding for Implicit Video Representation
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates
by: Peng, Weikun, et al.
Published: (2026)
by: Peng, Weikun, et al.
Published: (2026)
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
by: Guo, Junliang, et al.
Published: (2025)
by: Guo, Junliang, et al.
Published: (2025)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
by: Brown, Ellis, et al.
Published: (2025)
by: Brown, Ellis, et al.
Published: (2025)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
by: Chen, Xinlei, et al.
Published: (2024)
by: Chen, Xinlei, et al.
Published: (2024)
Self-Refining Video Sampling
by: Jang, Sangwon, et al.
Published: (2026)
by: Jang, Sangwon, et al.
Published: (2026)
MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines
by: Po, Ryan, et al.
Published: (2026)
by: Po, Ryan, et al.
Published: (2026)
Fisher Information Is the Squared Lorentz Factor: Why the Qubit Is Special
by: Srivats, Bharath G
Published: (2026)
by: Srivats, Bharath G
Published: (2026)
A Modular Speed Limit on the Qubit: Bloch Precession, Saturation, Channel Data-Processing, and the Vanishing of Modular Flow at Pure States
by: Srivats, Bharath G
Published: (2026)
by: Srivats, Bharath G
Published: (2026)
Transition Matching Distillation for Fast Video Generation
by: Nie, Weili, et al.
Published: (2026)
by: Nie, Weili, et al.
Published: (2026)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025)
by: Tang, Bingda, et al.
Published: (2025)
Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
by: Lu, Jiacheng, et al.
Published: (2026)
by: Lu, Jiacheng, et al.
Published: (2026)
WildSmoke: Ready-to-Use Dynamic 3D Smoke Assets from a Single Video in the Wild
by: Liu, Yuqiu, et al.
Published: (2025)
by: Liu, Yuqiu, et al.
Published: (2025)
Survey on Modeling of Human-made Articulated Objects
by: Liu, Jiayi, et al.
Published: (2024)
by: Liu, Jiayi, et al.
Published: (2024)
Cambrian-P: Pose-Grounded Video Understanding
by: Yang, Jihan, et al.
Published: (2026)
by: Yang, Jihan, et al.
Published: (2026)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
by: Qin, Yiran, et al.
Published: (2023)
by: Qin, Yiran, et al.
Published: (2023)
Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes
by: Merino, Tim, et al.
Published: (2026)
by: Merino, Tim, et al.
Published: (2026)
Flow Map Distillation Without Data
by: Tong, Shangyuan, et al.
Published: (2025)
by: Tong, Shangyuan, et al.
Published: (2025)
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025)
by: Zheng, Boyang, et al.
Published: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
by: Brown, Ellis, et al.
Published: (2025)
by: Brown, Ellis, et al.
Published: (2025)
STEVE Series: Step-by-Step Construction of Agent Systems in Minecraft
by: Zhao, Zhonghan, et al.
Published: (2024)
by: Zhao, Zhonghan, et al.
Published: (2024)
WorldModelBench: Judging Video Generation Models As World Models
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
Image Sculpting: Precise Object Editing with 3D Geometry Control
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
by: Chen, Jiacheng, et al.
Published: (2025)
by: Chen, Jiacheng, et al.
Published: (2025)
Multiplayer Information Asymmetric Contextual Bandits
by: Chang, William, et al.
Published: (2025)
by: Chang, William, et al.
Published: (2025)
V-IRL: Grounding Virtual Intelligence in Real Life
by: Yang, Jihan, et al.
Published: (2024)
by: Yang, Jihan, et al.
Published: (2024)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
by: Yang, Jihan, et al.
Published: (2024)
by: Yang, Jihan, et al.
Published: (2024)
Similar Items
-
Trustless Audits without Revealing Data or Models
by: Waiwitlikhit, Suppakit, et al.
Published: (2024) -
On Scaling Up 3D Gaussian Splatting Training
by: Zhao, Hexu, et al.
Published: (2024) -
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
by: Li, Chenyu, et al.
Published: (2025) -
Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?
by: Hu, Chuxuan, et al.
Published: (2025) -
iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
by: Peng, Weikun, et al.
Published: (2025)