MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xuehai, Zhou, Shijie, Venkateswaran, Thivyanth, Zheng, Kaizhi, Wan, Ziyu, Kadambi, Achuta, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
by: Zheng, Kaizhi, et al.
Published: (2023)
by: Zheng, Kaizhi, et al.
Published: (2023)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
by: Jeon, Changwoo, et al.
Published: (2026)
by: Jeon, Changwoo, et al.
Published: (2026)
Self-Evolving 3D Scene Generation from a Single Image
by: Zheng, Kaizhi, et al.
Published: (2025)
by: Zheng, Kaizhi, et al.
Published: (2025)
The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety
by: Jalilian, Laleh, et al.
Published: (2024)
by: Jalilian, Laleh, et al.
Published: (2024)
Solutions to Deepfakes: Can Camera Hardware, Cryptography, and Deep Learning Verify Real Images?
by: Vilesov, Alexander, et al.
Published: (2024)
by: Vilesov, Alexander, et al.
Published: (2024)
Bridging the Gap Between Multimodal Foundation Models and World Models
by: He, Xuehai
Published: (2025)
by: He, Xuehai
Published: (2025)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
by: Zheng, Kaizhi, et al.
Published: (2022)
by: Zheng, Kaizhi, et al.
Published: (2022)
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
by: Zheng, Kaizhi, et al.
Published: (2024)
by: Zheng, Kaizhi, et al.
Published: (2024)
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
by: Zhou, Shijie, et al.
Published: (2023)
by: Zhou, Shijie, et al.
Published: (2023)
3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
by: He, Yunhong, et al.
Published: (2025)
by: He, Yunhong, et al.
Published: (2025)
SparseGS: Real-Time 360° Sparse View Synthesis using Gaussian Splatting
by: Xiong, Haolin, et al.
Published: (2023)
by: Xiong, Haolin, et al.
Published: (2023)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
InstantRestore: Single-Step Personalized Face Restoration with Shared-Image Attention
by: Zhang, Howard, et al.
Published: (2024)
by: Zhang, Howard, et al.
Published: (2024)
4K4DGen: Panoramic 4D Generation at 4K Resolution
by: Li, Renjie, et al.
Published: (2024)
by: Li, Renjie, et al.
Published: (2024)
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models
by: Upadhyay, Rishi, et al.
Published: (2026)
by: Upadhyay, Rishi, et al.
Published: (2026)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
by: Hao, Jinkun, et al.
Published: (2026)
by: Hao, Jinkun, et al.
Published: (2026)
Multi-Exposure Image Fusion via Distilled 3D LUT Grid with Editable Mode
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
ComCLIP: Training-Free Compositional Image and Text Matching
by: Jiang, Kenan, et al.
Published: (2022)
by: Jiang, Kenan, et al.
Published: (2022)
2-Factor Retrieval for Improved Human-AI Decision Making in Radiology
by: Solomon, Jim, et al.
Published: (2024)
by: Solomon, Jim, et al.
Published: (2024)
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
PromptDecipher: Supporting AI Tutor Authoring Through Editable Simulated Interactions
by: Koyama, Miina, et al.
Published: (2026)
by: Koyama, Miina, et al.
Published: (2026)
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
by: Ma, Wufei, et al.
Published: (2026)
by: Ma, Wufei, et al.
Published: (2026)
HybridWorldSim: A Scalable and Controllable High-fidelity Simulator for Autonomous Driving
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes
by: Xiong, Yajiao, et al.
Published: (2025)
by: Xiong, Yajiao, et al.
Published: (2025)
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
by: Pruksachatkun, Yada, et al.
Published: (2026)
by: Pruksachatkun, Yada, et al.
Published: (2026)
Constructing a 3D Scene from a Single Image
by: Zheng, Kaizhi, et al.
Published: (2025)
by: Zheng, Kaizhi, et al.
Published: (2025)
Interactive Text-to-SQL Generation via Editable Step-by-Step Explanations
by: Tian, Yuan, et al.
Published: (2023)
by: Tian, Yuan, et al.
Published: (2023)
Editable Image Elements for Controllable Synthesis
by: Mu, Jiteng, et al.
Published: (2024)
by: Mu, Jiteng, et al.
Published: (2024)
Mojito: Motion Trajectory and Intensity Control for Video Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
Thermal Imaging and Radar for Remote Sleep Monitoring of Breathing and Apnea
by: Del Regno, Kai, et al.
Published: (2024)
by: Del Regno, Kai, et al.
Published: (2024)
SimSpark: Interactive Simulation of Social Media Behaviors
by: Lin, Ziyue, et al.
Published: (2025)
by: Lin, Ziyue, et al.
Published: (2025)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
Language Modeling with Editable External Knowledge
by: Li, Belinda Z., et al.
Published: (2024)
by: Li, Belinda Z., et al.
Published: (2024)
Detecting Activities of Daily Living in Egocentric Video to Contextualize Hand Use at Home in Outpatient Neurorehabilitation Settings
by: Kadambi, Adesh, et al.
Published: (2024)
by: Kadambi, Adesh, et al.
Published: (2024)
WeatherProof: Leveraging Language Guidance for Semantic Segmentation in Adverse Weather
by: Gella, Blake, et al.
Published: (2024)
by: Gella, Blake, et al.
Published: (2024)
Similar Items
-
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
by: Zhou, Shijie, et al.
Published: (2025) -
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
by: Zheng, Kaizhi, et al.
Published: (2023) -
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026) -
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
by: Jeon, Changwoo, et al.
Published: (2026) -
Self-Evolving 3D Scene Generation from a Single Image
by: Zheng, Kaizhi, et al.
Published: (2025)