Yume-1.5: A Text-Controlled Interactive World Generation Model
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Xiaofeng, Li, Zhen, Li, Chuanhao, Xu, Xiaojie, Ying, Kaining, He, Tong, Pang, Jiangmiao, Qiao, Yu, Zhang, Kaipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Yume: An Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026)
by: Mao, Xiaofeng, et al.
Published: (2026)
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
by: Xu, Xiaojie, et al.
Published: (2026)
by: Xu, Xiaojie, et al.
Published: (2026)
Sekai: A Video Dataset towards World Exploration
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
SVBench: Evaluation of Video Generation Models on Social Reasoning
by: Peng, Wenshuo, et al.
Published: (2025)
by: Peng, Wenshuo, et al.
Published: (2025)
WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
by: Li, Zhen, et al.
Published: (2026)
by: Li, Zhen, et al.
Published: (2026)
IA-T2I: Internet-Augmented Text-to-Image Generation
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
World Craft: Agentic Framework to Create Visualizable Worlds via Text
by: Sun, Jianwen, et al.
Published: (2026)
by: Sun, Jianwen, et al.
Published: (2026)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
by: Hao, Jinkun, et al.
Published: (2026)
by: Hao, Jinkun, et al.
Published: (2026)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
DeepVerse: 4D Autoregressive Video Generation as a World Model
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
by: Feng, Yukang, et al.
Published: (2025)
by: Feng, Yukang, et al.
Published: (2025)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
by: Zhou, Pengfei, et al.
Published: (2024)
by: Zhou, Pengfei, et al.
Published: (2024)
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos
by: Li, Kaining, et al.
Published: (2025)
by: Li, Kaining, et al.
Published: (2025)
CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation
by: Liu, Xi, et al.
Published: (2024)
by: Liu, Xi, et al.
Published: (2024)
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
by: Zhao, Shitian, et al.
Published: (2025)
by: Zhao, Shitian, et al.
Published: (2025)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
by: Chang, Yifan, et al.
Published: (2025)
by: Chang, Yifan, et al.
Published: (2025)
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
by: Chen, Jiahe, et al.
Published: (2026)
by: Chen, Jiahe, et al.
Published: (2026)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
Composition-Incremental Learning for Compositional Generalization
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
Multi-Object Tracking by Hierarchical Visual Representations
by: Cao, Jinkun, et al.
Published: (2024)
by: Cao, Jinkun, et al.
Published: (2024)
Aether: Geometric-Aware Unified World Modeling
by: Aether Team, et al.
Published: (2025)
by: Aether Team, et al.
Published: (2025)
MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow Generation
by: Ren, Kerui, et al.
Published: (2025)
by: Ren, Kerui, et al.
Published: (2025)
Generative World Renderer
by: Huang, Zheng-Hui, et al.
Published: (2026)
by: Huang, Zheng-Hui, et al.
Published: (2026)
STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System
by: Luo, Zhen, et al.
Published: (2026)
by: Luo, Zhen, et al.
Published: (2026)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
MOVE: Motion-Guided Few-Shot Video Object Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
Segment Anything Across Shots: A Method and Benchmark
by: Hu, Hengrui, et al.
Published: (2025)
by: Hu, Hengrui, et al.
Published: (2025)
Consistency of Compositional Generalization across Multiple Levels
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
WonderTurbo: Generating Interactive 3D World in 0.72 Seconds
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
by: Liu, Shuo, et al.
Published: (2024)
by: Liu, Shuo, et al.
Published: (2024)
Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation
by: Wu, Qingxuan, et al.
Published: (2025)
by: Wu, Qingxuan, et al.
Published: (2025)
Similar Items
-
Yume: An Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025) -
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026) -
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
by: Xu, Xiaojie, et al.
Published: (2026) -
Sekai: A Video Dataset towards World Exploration
by: Li, Zhen, et al.
Published: (2025) -
SVBench: Evaluation of Video Generation Models on Social Reasoning
by: Peng, Wenshuo, et al.
Published: (2025)