MuDreamer: Learning Predictive World Models without Reconstruction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Burchi, Maxime, Timofte, Radu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Transformer-based World Models with Contrastive Predictive Coding
von: Burchi, Maxime, et al.
Veröffentlicht: (2025)
von: Burchi, Maxime, et al.
Veröffentlicht: (2025)
Accurate and Efficient World Modeling with Masked Latent Transformers
von: Burchi, Maxime, et al.
Veröffentlicht: (2025)
von: Burchi, Maxime, et al.
Veröffentlicht: (2025)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
Learned Lightweight Smartphone ISP with Unpaired Data
von: Arhire, Andrei, et al.
Veröffentlicht: (2025)
von: Arhire, Andrei, et al.
Veröffentlicht: (2025)
SafeDreamer: Safe Reinforcement Learning with World Models
von: Huang, Weidong, et al.
Veröffentlicht: (2023)
von: Huang, Weidong, et al.
Veröffentlicht: (2023)
MobileAgeNet: Lightweight Facial Age Estimation for Mobile Deployment
von: Kumar, Arun, et al.
Veröffentlicht: (2026)
von: Kumar, Arun, et al.
Veröffentlicht: (2026)
WaveHiT-SR: Hierarchical Wavelet Network for Efficient Image Super-Resolution
von: Ali, Fayaz, et al.
Veröffentlicht: (2025)
von: Ali, Fayaz, et al.
Veröffentlicht: (2025)
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
von: Ni, Chaojun, et al.
Veröffentlicht: (2024)
von: Ni, Chaojun, et al.
Veröffentlicht: (2024)
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
von: Jesani, Krunal, et al.
Veröffentlicht: (2025)
von: Jesani, Krunal, et al.
Veröffentlicht: (2025)
Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design
von: Duvvuri, Raghuvir, et al.
Veröffentlicht: (2025)
von: Duvvuri, Raghuvir, et al.
Veröffentlicht: (2025)
Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs
von: Adhikari, Santosh Premi, et al.
Veröffentlicht: (2026)
von: Adhikari, Santosh Premi, et al.
Veröffentlicht: (2026)
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
von: Zhou, Yang, et al.
Veröffentlicht: (2026)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
Latent Video Prediction Learns Better World Models
von: Alrasheed, Ali J, et al.
Veröffentlicht: (2026)
von: Alrasheed, Ali J, et al.
Veröffentlicht: (2026)
PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics
von: Luan, Xueyu, et al.
Veröffentlicht: (2026)
von: Luan, Xueyu, et al.
Veröffentlicht: (2026)
SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input
von: Lv, Zhen, et al.
Veröffentlicht: (2024)
von: Lv, Zhen, et al.
Veröffentlicht: (2024)
Motion Dreamer: Boundary Conditional Motion Reasoning for Physically Coherent Video Generation
von: Xu, Tianshuo, et al.
Veröffentlicht: (2024)
von: Xu, Tianshuo, et al.
Veröffentlicht: (2024)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
von: Gado, Mohamed, et al.
Veröffentlicht: (2025)
von: Gado, Mohamed, et al.
Veröffentlicht: (2025)
PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
von: Xu, Sirui, et al.
Veröffentlicht: (2024)
von: Xu, Sirui, et al.
Veröffentlicht: (2024)
3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
von: Zhang, Frank, et al.
Veröffentlicht: (2024)
von: Zhang, Frank, et al.
Veröffentlicht: (2024)
Cluster and Predict Latent Patches for Improved Masked Image Modeling
von: Darcet, Timothée, et al.
Veröffentlicht: (2025)
von: Darcet, Timothée, et al.
Veröffentlicht: (2025)
MuST: Multi-Scale Transformers for Surgical Phase Recognition
von: Pérez, Alejandra, et al.
Veröffentlicht: (2024)
von: Pérez, Alejandra, et al.
Veröffentlicht: (2024)
MuNet: A Mutualistic Network for Joint 3D Human Mesh Recovery and 3D Clothed Human Reconstruction from Single Images
von: Gao, Yunqi, et al.
Veröffentlicht: (2026)
von: Gao, Yunqi, et al.
Veröffentlicht: (2026)
GraphicsDreamer: Image to 3D Generation with Physical Consistency
von: Chen, Pei, et al.
Veröffentlicht: (2024)
von: Chen, Pei, et al.
Veröffentlicht: (2024)
KAN-Dreamer: Benchmarking Kolmogorov-Arnold Networks as Function Approximators in World Models
von: Shi, Chenwei, et al.
Veröffentlicht: (2025)
von: Shi, Chenwei, et al.
Veröffentlicht: (2025)
Factorized-Dreamer: Training A High-Quality Video Generator with Limited and Low-Quality Data
von: Yang, Tao, et al.
Veröffentlicht: (2024)
von: Yang, Tao, et al.
Veröffentlicht: (2024)
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
Exploring the Underwater World Segmentation without Extra Training
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
3D Reconstruction of Objects in Hands without Real World 3D Supervision
von: Prakash, Aditya, et al.
Veröffentlicht: (2023)
von: Prakash, Aditya, et al.
Veröffentlicht: (2023)
Optimizing Dense Visual Predictions Through Multi-Task Coherence and Prioritization
von: Fontana, Maxime, et al.
Veröffentlicht: (2024)
von: Fontana, Maxime, et al.
Veröffentlicht: (2024)
Temporal-Spatial Tubelet Embedding for Cloud-Robust MSI Reconstruction using MSI-SAR Fusion: A Multi-Head Self-Attention Video Vision Transformer Approach
von: Wang, Yiqun, et al.
Veröffentlicht: (2025)
von: Wang, Yiqun, et al.
Veröffentlicht: (2025)
Practical Manipulation Model for Robust Deepfake Detection
von: Hopf, Benedikt, et al.
Veröffentlicht: (2025)
von: Hopf, Benedikt, et al.
Veröffentlicht: (2025)
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2026)
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2026)
BrainDreamer: Reasoning-Coherent and Controllable Image Generation from EEG Brain Signals via Language Guidance
von: Wang, Ling, et al.
Veröffentlicht: (2024)
von: Wang, Ling, et al.
Veröffentlicht: (2024)
GARF: Learning Generalizable 3D Reassembly for Real-World Fractures
von: Li, Sihang, et al.
Veröffentlicht: (2025)
von: Li, Sihang, et al.
Veröffentlicht: (2025)
ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2024)
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
von: Wang, Yilin, et al.
Veröffentlicht: (2025)
von: Wang, Yilin, et al.
Veröffentlicht: (2025)
SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Learning Transformer-based World Models with Contrastive Predictive Coding
von: Burchi, Maxime, et al.
Veröffentlicht: (2025) -
Accurate and Efficient World Modeling with Masked Latent Transformers
von: Burchi, Maxime, et al.
Veröffentlicht: (2025) -
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
von: Burchi, Maxime, et al.
Veröffentlicht: (2024) -
Learned Lightweight Smartphone ISP with Unpaired Data
von: Arhire, Andrei, et al.
Veröffentlicht: (2025) -
SafeDreamer: Safe Reinforcement Learning with World Models
von: Huang, Weidong, et al.
Veröffentlicht: (2023)