WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Bohai, Wu, Taiyi, Yuan, Yueyang, Liu, Jian, Lu, Xiaocheng, Du, Dazhao, Zhang, Jie, Lai, Jinxiang, Yang, Shuai, Zhao, Xiaotong, Zhao, Alan, Guo, Song |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation
by: Wu, Jiafeng, et al.
Published: (2026)
by: Wu, Jiafeng, et al.
Published: (2026)
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
by: Yang, Songlin, et al.
Published: (2026)
by: Yang, Songlin, et al.
Published: (2026)
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
by: Du, Dazhao, et al.
Published: (2026)
by: Du, Dazhao, et al.
Published: (2026)
Building Explicit World Model for Zero-Shot Open-World Object Manipulation
by: Li, Xiaotong, et al.
Published: (2026)
by: Li, Xiaotong, et al.
Published: (2026)
Spider: Any-to-Many Multimodal LLM
by: Lai, Jinxiang, et al.
Published: (2024)
by: Lai, Jinxiang, et al.
Published: (2024)
Vid2World: Crafting Video Diffusion Models to Interactive World Models
by: Huang, Siqiao, et al.
Published: (2025)
by: Huang, Siqiao, et al.
Published: (2025)
Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
by: Tang, Junshu, et al.
Published: (2025)
by: Tang, Junshu, et al.
Published: (2025)
DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos
by: Li, Can, et al.
Published: (2026)
by: Li, Can, et al.
Published: (2026)
Towards High-Consistency Embodied World Model with Multi-View Trajectory Videos
by: Su, Taiyi, et al.
Published: (2025)
by: Su, Taiyi, et al.
Published: (2025)
Flow-Guided Diffusion for Video Inpainting
by: Gu, Bohai, et al.
Published: (2023)
by: Gu, Bohai, et al.
Published: (2023)
Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion
by: Gu, Bohai, et al.
Published: (2024)
by: Gu, Bohai, et al.
Published: (2024)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
by: Fu, Hongming, et al.
Published: (2026)
by: Fu, Hongming, et al.
Published: (2026)
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
Predicting the Future by Retrieving the Past
by: Du, Dazhao, et al.
Published: (2025)
by: Du, Dazhao, et al.
Published: (2025)
PhysGen3D: Crafting a Miniature Interactive World from a Single Image
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
AgentWorld: An Interactive Simulation Platform for Scene Construction and Mobile Robotic Manipulation
by: Zhang, Yizheng, et al.
Published: (2025)
by: Zhang, Yizheng, et al.
Published: (2025)
World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
by: Jiang, Zhennan, et al.
Published: (2025)
by: Jiang, Zhennan, et al.
Published: (2025)
Vision-based Manipulation from Single Human Video with Open-World Object Graphs
by: Zhu, Yifeng, et al.
Published: (2024)
by: Zhu, Yifeng, et al.
Published: (2024)
Navigating in High-Dimensional Search Space: A Hierarchical Bayesian Optimization Approach
by: Li, Wenxuan, et al.
Published: (2024)
by: Li, Wenxuan, et al.
Published: (2024)
Cooperative-Competitive Team Play of Real-World Craft Robots
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
PartCraft: Crafting Creative Objects by Parts
by: Ng, Kam Woh, et al.
Published: (2024)
by: Ng, Kam Woh, et al.
Published: (2024)
Disruptions as Opportunities
by: Sun, Taiyi
Published: (2023)
by: Sun, Taiyi
Published: (2023)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
World Craft: Agentic Framework to Create Visualizable Worlds via Text
by: Sun, Jianwen, et al.
Published: (2026)
by: Sun, Jianwen, et al.
Published: (2026)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
Object-Centric World Model for Language-Guided Manipulation
by: Jeong, Youngjoon, et al.
Published: (2025)
by: Jeong, Youngjoon, et al.
Published: (2025)
Adaptive Mobile Manipulation for Articulated Objects In the Open World
by: Xiong, Haoyu, et al.
Published: (2024)
by: Xiong, Haoyu, et al.
Published: (2024)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
Stereo World Model: Camera-Guided Stereo Video Generation
by: Sun, Yang-Tian, et al.
Published: (2026)
by: Sun, Yang-Tian, et al.
Published: (2026)
Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control
by: Xie, Linxi, et al.
Published: (2026)
by: Xie, Linxi, et al.
Published: (2026)
Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation
by: Gu, Songen, et al.
Published: (2026)
by: Gu, Songen, et al.
Published: (2026)
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
by: Song, Zijian, et al.
Published: (2026)
by: Song, Zijian, et al.
Published: (2026)
MONA: Moving Object Detection from Videos Shot by Dynamic Camera
by: Hu, Boxun, et al.
Published: (2025)
by: Hu, Boxun, et al.
Published: (2025)
Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation
by: Xue, Han, et al.
Published: (2026)
by: Xue, Han, et al.
Published: (2026)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
by: Dharmarajan, Karthik, et al.
Published: (2025)
by: Dharmarajan, Karthik, et al.
Published: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
OmniNWM: Omniscient Driving Navigation World Models
by: Li, Bohan, et al.
Published: (2025)
by: Li, Bohan, et al.
Published: (2025)
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
by: Chen, Yuzhi, et al.
Published: (2026)
by: Chen, Yuzhi, et al.
Published: (2026)
Similar Items
-
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
by: Gu, Bohai, et al.
Published: (2026) -
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
by: Liu, Xinhang, et al.
Published: (2025) -
From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation
by: Wu, Jiafeng, et al.
Published: (2026) -
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
by: Yang, Songlin, et al.
Published: (2026) -
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
by: Du, Dazhao, et al.
Published: (2026)