MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Enshen, Qin, Yiran, Yin, Zhenfei, Huang, Yuzhou, Zhang, Ruimao, Sheng, Lu, Qiao, Yu, Shao, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
by: Qin, Yiran, et al.
Published: (2023)
by: Qin, Yiran, et al.
Published: (2023)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
Story3D-Agent: Exploring 3D Storytelling Visualization with Large Language Models
by: Huang, Yuzhou, et al.
Published: (2024)
by: Huang, Yuzhou, et al.
Published: (2024)
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
by: Qin, Yiran, et al.
Published: (2026)
by: Qin, Yiran, et al.
Published: (2026)
RoboDreamer: Learning Compositional World Models for Robot Imagination
by: Zhou, Siyuan, et al.
Published: (2024)
by: Zhou, Siyuan, et al.
Published: (2024)
RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
by: Qin, Yiran, et al.
Published: (2025)
by: Qin, Yiran, et al.
Published: (2025)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Mind Dreamer: Untethering Imagination via Active Causal Intervention on Latent Manifolds
by: Xu, Shaojun, et al.
Published: (2026)
by: Xu, Shaojun, et al.
Published: (2026)
High-Dynamic Radar Sequence Prediction for Weather Nowcasting Using Spatiotemporal Coherent Gaussian Representation
by: Wang, Ziye, et al.
Published: (2025)
by: Wang, Ziye, et al.
Published: (2025)
Toward Accurate Camera-based 3D Object Detection via Cascade Depth Estimation and Calibration
by: Wang, Chaoqun, et al.
Published: (2024)
by: Wang, Chaoqun, et al.
Published: (2024)
ChronoDreamer: Action-Conditioned World Model as an Online Simulator for Robotic Planning
by: Zhou, Zhenhao, et al.
Published: (2025)
by: Zhou, Zhenhao, et al.
Published: (2025)
V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors
by: He, Songjia, et al.
Published: (2026)
by: He, Songjia, et al.
Published: (2026)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
by: Chen, Zeren, et al.
Published: (2023)
by: Chen, Zeren, et al.
Published: (2023)
CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion
by: Ma, Jiahua, et al.
Published: (2025)
by: Ma, Jiahua, et al.
Published: (2025)
HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
by: Lu, Xiaoya, et al.
Published: (2026)
by: Lu, Xiaoya, et al.
Published: (2026)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants
by: Qin, Yiran, et al.
Published: (2025)
by: Qin, Yiran, et al.
Published: (2025)
EditWorld: Simulating World Dynamics for Instruction-Following Image Editing
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
MobileDreamer: Generative Sketch World Model for GUI Agent
by: Cao, Yilin, et al.
Published: (2026)
by: Cao, Yilin, et al.
Published: (2026)
PrivilegedDreamer: Explicit Imagination of Privileged Information for Rapid Adaptation of Learned Policies
by: Byrd, Morgan, et al.
Published: (2025)
by: Byrd, Morgan, et al.
Published: (2025)
RH20T-P: A Primitive-Level Robotic Dataset Towards Composable Generalization Agents
by: Chen, Zeren, et al.
Published: (2024)
by: Chen, Zeren, et al.
Published: (2024)
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
by: Ni, Chaojun, et al.
Published: (2024)
by: Ni, Chaojun, et al.
Published: (2024)
VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
by: Kang, Li, et al.
Published: (2025)
by: Kang, Li, et al.
Published: (2025)
Assessment of Multimodal Large Language Models in Alignment with Human Values
by: Shi, Zhelun, et al.
Published: (2024)
by: Shi, Zhelun, et al.
Published: (2024)
SafeDreamer: Safe Reinforcement Learning with World Models
by: Huang, Weidong, et al.
Published: (2023)
by: Huang, Weidong, et al.
Published: (2023)
Counterfactual Visual Explanation via Causally-Guided Adversarial Steering
by: Qiao, Yiran, et al.
Published: (2025)
by: Qiao, Yiran, et al.
Published: (2025)
DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering
by: Qiao, Yiran, et al.
Published: (2026)
by: Qiao, Yiran, et al.
Published: (2026)
DyMoDreamer: World Modeling with Dynamic Modulation
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
by: Wang, Boyuan, et al.
Published: (2025)
by: Wang, Boyuan, et al.
Published: (2025)
ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
by: Zhou, Heng, et al.
Published: (2025)
by: Zhou, Heng, et al.
Published: (2025)
TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance
by: Zhang, Zhemeng, et al.
Published: (2026)
by: Zhang, Zhemeng, et al.
Published: (2026)
Front‐Footed Defense: Leveraging Early Counsel Intervention for Expedited Justice
by: Chengchen He, et al.
Published: (2026)
by: Chengchen He, et al.
Published: (2026)
DreamerV3 for Traffic Signal Control: Hyperparameter Tuning and Performance
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots
by: Wu, Sheng, et al.
Published: (2025)
by: Wu, Sheng, et al.
Published: (2025)
Dancing in Chains: Reconciling Instruction Following and Faithfulness in Language Models
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Prune4Web: DOM Tree Pruning Programming for Web Agent
by: Zhang, Jiayuan, et al.
Published: (2025)
by: Zhang, Jiayuan, et al.
Published: (2025)
Advancing Medical Radiograph Representation Learning: A Hybrid Pre-training Paradigm with Multilevel Semantic Granularity
by: Jiang, Hanqi, et al.
Published: (2024)
by: Jiang, Hanqi, et al.
Published: (2024)
Probing initial isocurvature perturbation with 21cm one-point statistics
by: Qin, Zhenfei, et al.
Published: (2025)
by: Qin, Zhenfei, et al.
Published: (2025)
Similar Items
-
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
by: Qin, Yiran, et al.
Published: (2023) -
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024) -
Story3D-Agent: Exploring 3D Storytelling Visualization with Large Language Models
by: Huang, Yuzhou, et al.
Published: (2024) -
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
by: Qin, Yiran, et al.
Published: (2026) -
RoboDreamer: Learning Compositional World Models for Robot Imagination
by: Zhou, Siyuan, et al.
Published: (2024)