DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Gaoyue, Pan, Hengkai, LeCun, Yann, Pinto, Lerrel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
by: Terver, Basile, et al.
Published: (2025)
by: Terver, Basile, et al.
Published: (2025)
OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
AdaWM: Adaptive World Model based Planning for Autonomous Driving
by: Wang, Hang, et al.
Published: (2025)
by: Wang, Hang, et al.
Published: (2025)
ResWM: Residual-Action World Model for Visual RL
by: Zhang, Jseen, et al.
Published: (2026)
by: Zhang, Jseen, et al.
Published: (2026)
Value-guided action planning with JEPA world models
by: Destrade, Matthieu, et al.
Published: (2025)
by: Destrade, Matthieu, et al.
Published: (2025)
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control
by: Cui, Zichen Jeff, et al.
Published: (2024)
by: Cui, Zichen Jeff, et al.
Published: (2024)
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
by: Jiang, Feng, et al.
Published: (2026)
by: Jiang, Feng, et al.
Published: (2026)
Parallel Stochastic Gradient-Based Planning for World Models
by: Psenka, Michael, et al.
Published: (2026)
by: Psenka, Michael, et al.
Published: (2026)
EgoZero: Robot Learning from Smart Glasses
by: Liu, Vincent, et al.
Published: (2025)
by: Liu, Vincent, et al.
Published: (2025)
Fast and Exact Enumeration of Deep Networks Partitions Regions
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields
by: Yang, Zhaoyang, et al.
Published: (2026)
by: Yang, Zhaoyang, et al.
Published: (2026)
DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
by: Azhari, Maulana Bisyir, et al.
Published: (2025)
by: Azhari, Maulana Bisyir, et al.
Published: (2025)
LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation
by: Huang, Yuhang, et al.
Published: (2025)
by: Huang, Yuhang, et al.
Published: (2025)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Causal-JEPA: Learning World Models through Object-Level Latent Masking
by: Nam, Heejeong, et al.
Published: (2026)
by: Nam, Heejeong, et al.
Published: (2026)
Learning Precise, Contact-Rich Manipulation through Uncalibrated Tactile Skins
by: Pattabiraman, Venkatesh, et al.
Published: (2024)
by: Pattabiraman, Venkatesh, et al.
Published: (2024)
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
by: Zha, Lihan, et al.
Published: (2026)
by: Zha, Lihan, et al.
Published: (2026)
Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
by: Dupoux, Emmanuel, et al.
Published: (2026)
by: Dupoux, Emmanuel, et al.
Published: (2026)
Reliable Semantic Understanding for Real World Zero-shot Object Goal Navigation
by: Unlu, Halil Utku, et al.
Published: (2024)
by: Unlu, Halil Utku, et al.
Published: (2024)
ContactGaussian-WM: Learning Physics-Grounded World Model from Videos
by: Wang, Meizhong, et al.
Published: (2026)
by: Wang, Meizhong, et al.
Published: (2026)
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024)
by: Garrido, Quentin, et al.
Published: (2024)
Leveraging Pre-trained Large Language Models with Refined Prompting for Online Task and Motion Planning
by: Guo, Huihui, et al.
Published: (2025)
by: Guo, Huihui, et al.
Published: (2025)
Zero-shot Interactive Perception
by: Sripada, Venkatesh, et al.
Published: (2026)
by: Sripada, Venkatesh, et al.
Published: (2026)
Policy-Guided World Model Planning for Language-Conditioned Visual Navigation
by: Chahe, Amirhosein, et al.
Published: (2026)
by: Chahe, Amirhosein, et al.
Published: (2026)
eFlesh: Highly customizable Magnetic Touch Sensing using Cut-Cell Microstructures
by: Pattabiraman, Venkatesh, et al.
Published: (2025)
by: Pattabiraman, Venkatesh, et al.
Published: (2025)
AnySkin: Plug-and-play Skin Sensing for Robotic Touch
by: Bhirangi, Raunaq, et al.
Published: (2024)
by: Bhirangi, Raunaq, et al.
Published: (2024)
Closing the Train-Test Gap in World Models for Gradient-Based Planning
by: Parthasarathy, Arjun, et al.
Published: (2025)
by: Parthasarathy, Arjun, et al.
Published: (2025)
LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
by: Maes, Lucas, et al.
Published: (2026)
by: Maes, Lucas, et al.
Published: (2026)
Zero-shot Object Navigation with Vision-Language Models Reasoning
by: Wen, Congcong, et al.
Published: (2024)
by: Wen, Congcong, et al.
Published: (2024)
Show and Grasp: Few-shot Semantic Segmentation for Robot Grasping through Zero-shot Foundation Models
by: Barcellona, Leonardo, et al.
Published: (2024)
by: Barcellona, Leonardo, et al.
Published: (2024)
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
by: Hansen, Nicklas, et al.
Published: (2024)
by: Hansen, Nicklas, et al.
Published: (2024)
Wonderful Team: Zero-Shot Physical Task Planning with Visual LLMs
by: Wang, Zidan, et al.
Published: (2024)
by: Wang, Zidan, et al.
Published: (2024)
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
by: Chi, Haozhuang, et al.
Published: (2026)
by: Chi, Haozhuang, et al.
Published: (2026)
AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models
by: Yao, Fanglong, et al.
Published: (2024)
by: Yao, Fanglong, et al.
Published: (2024)
RUKA: Rethinking the Design of Humanoid Hands with Learning
by: Zorin, Anya, et al.
Published: (2025)
by: Zorin, Anya, et al.
Published: (2025)
Similar Items
-
Navigation World Models
by: Bar, Amir, et al.
Published: (2024) -
What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
by: Terver, Basile, et al.
Published: (2025) -
OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
by: Goswami, Raktim Gautam, et al.
Published: (2025) -
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025) -
AdaWM: Adaptive World Model based Planning for Autonomous Driving
by: Wang, Hang, et al.
Published: (2025)