MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Xuehai, Zhou, Shijie, Venkateswaran, Thivyanth, Zheng, Kaizhi, Wan, Ziyu, Kadambi, Achuta, Wang, Xin Eric
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914076340256768
author He, Xuehai
Zhou, Shijie
Venkateswaran, Thivyanth
Zheng, Kaizhi
Wan, Ziyu
Kadambi, Achuta
Wang, Xin Eric
author_facet He, Xuehai
Zhou, Shijie
Venkateswaran, Thivyanth
Zheng, Kaizhi
Wan, Ziyu
Kadambi, Achuta
Wang, Xin Eric
contents World models that support controllable and editable spatiotemporal environments are valuable for robotics, enabling scalable training data, repro ducible evaluation, and flexible task design. While recent text-to-video models generate realistic dynam ics, they are constrained to 2D views and offer limited interaction. We introduce MorphoSim, a language guided framework that generates 4D scenes with multi-view consistency and object-level controls. From natural language instructions, MorphoSim produces dynamic environments where objects can be directed, recolored, or removed, and scenes can be observed from arbitrary viewpoints. The framework integrates trajectory-guided generation with feature field dis tillation, allowing edits to be applied interactively without full re-generation. Experiments show that Mor phoSim maintains high scene fidelity while enabling controllability and editability. The code is available at https://github.com/eric-ai-lab/Morph4D.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
He, Xuehai
Zhou, Shijie
Venkateswaran, Thivyanth
Zheng, Kaizhi
Wan, Ziyu
Kadambi, Achuta
Wang, Xin Eric
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
World models that support controllable and editable spatiotemporal environments are valuable for robotics, enabling scalable training data, repro ducible evaluation, and flexible task design. While recent text-to-video models generate realistic dynam ics, they are constrained to 2D views and offer limited interaction. We introduce MorphoSim, a language guided framework that generates 4D scenes with multi-view consistency and object-level controls. From natural language instructions, MorphoSim produces dynamic environments where objects can be directed, recolored, or removed, and scenes can be observed from arbitrary viewpoints. The framework integrates trajectory-guided generation with feature field dis tillation, allowing edits to be applied interactively without full re-generation. Experiments show that Mor phoSim maintains high scene fidelity while enabling controllability and editability. The code is available at https://github.com/eric-ai-lab/Morph4D.
title MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.04390