VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Sixiao, Yin, Minghao, Hu, Wenbo, Li, Xiaoyu, Shan, Ying, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
by: YU, Mark, et al.
Published: (2025)
by: YU, Mark, et al.
Published: (2025)
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
by: Zheng, Sixiao, et al.
Published: (2024)
by: Zheng, Sixiao, et al.
Published: (2024)
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
by: Xu, Tian-Xing, et al.
Published: (2025)
by: Xu, Tian-Xing, et al.
Published: (2025)
ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
by: Yu, Wangbo, et al.
Published: (2024)
by: Yu, Wangbo, et al.
Published: (2024)
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
by: Hu, Wenbo, et al.
Published: (2024)
by: Hu, Wenbo, et al.
Published: (2024)
Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers
by: Yin, Minghao, et al.
Published: (2026)
by: Yin, Minghao, et al.
Published: (2026)
MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE
by: Zhu, Ruijie, et al.
Published: (2026)
by: Zhu, Ruijie, et al.
Published: (2026)
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
by: Zheng, Sixiao, et al.
Published: (2024)
by: Zheng, Sixiao, et al.
Published: (2024)
DeepVerse: 4D Autoregressive Video Generation as a World Model
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
by: Dhiman, Ankit, et al.
Published: (2025)
by: Dhiman, Ankit, et al.
Published: (2025)
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
by: Yang, Yuxue, et al.
Published: (2026)
by: Yang, Yuxue, et al.
Published: (2026)
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
by: Chen, Haoxin, et al.
Published: (2024)
by: Chen, Haoxin, et al.
Published: (2024)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
by: Liu, Yaofang, et al.
Published: (2023)
by: Liu, Yaofang, et al.
Published: (2023)
NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
by: Bin, Yanrui, et al.
Published: (2025)
by: Bin, Yanrui, et al.
Published: (2025)
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
by: Wen, Kairun, et al.
Published: (2025)
by: Wen, Kairun, et al.
Published: (2025)
AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
by: Niu, Muyao, et al.
Published: (2025)
by: Niu, Muyao, et al.
Published: (2025)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
by: Zheng, Sixiao, et al.
Published: (2025)
by: Zheng, Sixiao, et al.
Published: (2025)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
by: Yan, Yunzhi, et al.
Published: (2024)
by: Yan, Yunzhi, et al.
Published: (2024)
Mind-of-Director: Multi-modal Agent-Driven Film Previsualization via Collaborative Decision-Making
by: Nan, Shufeng, et al.
Published: (2026)
by: Nan, Shufeng, et al.
Published: (2026)
A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
ToonCrafter: Generative Cartoon Interpolation
by: Xing, Jinbo, et al.
Published: (2024)
by: Xing, Jinbo, et al.
Published: (2024)
StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter
by: Liu, Gongye, et al.
Published: (2023)
by: Liu, Gongye, et al.
Published: (2023)
WonderVerse: Extendable 3D Scene Generation with Video Generative Models
by: Feng, Hao, et al.
Published: (2025)
by: Feng, Hao, et al.
Published: (2025)
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
by: Lu, Jiahao, et al.
Published: (2026)
by: Lu, Jiahao, et al.
Published: (2026)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion
by: Wang, Chaoyang, et al.
Published: (2024)
by: Wang, Chaoyang, et al.
Published: (2024)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation
by: Wang, Rongsheng, et al.
Published: (2026)
by: Wang, Rongsheng, et al.
Published: (2026)
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
by: Gao, Jianxiong, et al.
Published: (2025)
by: Gao, Jianxiong, et al.
Published: (2025)
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
by: Liu, Kunhao, et al.
Published: (2025)
by: Liu, Kunhao, et al.
Published: (2025)
PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
by: Mao, Qing, et al.
Published: (2025)
by: Mao, Qing, et al.
Published: (2025)
Measuring 3D Spatial Geometric Consistency in Dynamic Generated Videos
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control
by: Zhong, Yong, et al.
Published: (2024)
by: Zhong, Yong, et al.
Published: (2024)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
by: Wang, Qilin, et al.
Published: (2024)
by: Wang, Qilin, et al.
Published: (2024)
HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model
by: Nguyen, Hieu T., et al.
Published: (2024)
by: Nguyen, Hieu T., et al.
Published: (2024)
CityCraft: A Real Crafter for 3D City Generation
by: Deng, Jie, et al.
Published: (2024)
by: Deng, Jie, et al.
Published: (2024)
Similar Items
-
TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
by: YU, Mark, et al.
Published: (2025) -
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
by: Zheng, Sixiao, et al.
Published: (2024) -
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
by: Xu, Tian-Xing, et al.
Published: (2025) -
ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
by: Yu, Wangbo, et al.
Published: (2024) -
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
by: Hu, Wenbo, et al.
Published: (2024)