Learning Latent Action World Models In The Wild
Fuente:
arXiv
Saved in:
| Main Authors: | Garrido, Quentin, Nagarajan, Tushar, Terver, Basile, Ballas, Nicolas, LeCun, Yann, Rabbat, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
by: Terver, Basile, et al.
Published: (2026)
by: Terver, Basile, et al.
Published: (2026)
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024)
by: Garrido, Quentin, et al.
Published: (2024)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
by: Garrido, Quentin, et al.
Published: (2025)
by: Garrido, Quentin, et al.
Published: (2025)
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)
by: Bardes, Adrien, et al.
Published: (2024)
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Blockwise Self-Supervised Learning at Scale
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
Back to the Features: DINO as a Foundation for Video World Models
by: Baldassarre, Federico, et al.
Published: (2025)
by: Baldassarre, Federico, et al.
Published: (2025)
Interpreting Physics in Video World Models
by: Joseph, Sonia, et al.
Published: (2026)
by: Joseph, Sonia, et al.
Published: (2026)
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
by: Mur-Labadia, Lorenzo, et al.
Published: (2026)
by: Mur-Labadia, Lorenzo, et al.
Published: (2026)
What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
by: Terver, Basile, et al.
Published: (2025)
by: Terver, Basile, et al.
Published: (2025)
Variance-Covariance Regularization Improves Representation Learning
by: Zhu, Jiachen, et al.
Published: (2023)
by: Zhu, Jiachen, et al.
Published: (2023)
Scaling Language-Free Visual Representation Learning
by: Fan, David, et al.
Published: (2025)
by: Fan, David, et al.
Published: (2025)
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
by: Rudman, William, et al.
Published: (2025)
by: Rudman, William, et al.
Published: (2025)
Transformers without Normalization
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
Causal-JEPA: Learning World Models through Object-Level Latent Masking
by: Nam, Heejeong, et al.
Published: (2026)
by: Nam, Heejeong, et al.
Published: (2026)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
by: Assran, Mido, et al.
Published: (2025)
by: Assran, Mido, et al.
Published: (2025)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
Advancing human-centric AI for robust X-ray analysis through holistic self-supervised learning
by: Moutakanni, Théo, et al.
Published: (2024)
by: Moutakanni, Théo, et al.
Published: (2024)
DiLA: Disentangled Latent Action World Models
by: Zhang, Tianqiu, et al.
Published: (2026)
by: Zhang, Tianqiu, et al.
Published: (2026)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
A hierarchical loss and its problems when classifying non-hierarchically
by: Wu, Cinna, et al.
Published: (2017)
by: Wu, Cinna, et al.
Published: (2017)
Object-Centric Latent Action Learning
by: Klepach, Albina, et al.
Published: (2025)
by: Klepach, Albina, et al.
Published: (2025)
User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance
by: Verghese, Mrinal, et al.
Published: (2024)
by: Verghese, Mrinal, et al.
Published: (2024)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
by: Zhai, Yuexiang, et al.
Published: (2024)
by: Zhai, Yuexiang, et al.
Published: (2024)
Latent Video Prediction Learns Better World Models
by: Alrasheed, Ali J, et al.
Published: (2026)
by: Alrasheed, Ali J, et al.
Published: (2026)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
by: Yun, Zeyu, et al.
Published: (2023)
by: Yun, Zeyu, et al.
Published: (2023)
DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks
by: Deng, Yueci, et al.
Published: (2026)
by: Deng, Yueci, et al.
Published: (2026)
Learning Visual Feature-Based World Models via Residual Latent Action
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
Learning Additively Compositional Latent Actions for Embodied AI
by: Wei, Hangxing, et al.
Published: (2026)
by: Wei, Hangxing, et al.
Published: (2026)
Learning Vision-Language-Action World Models for Autonomous Driving
by: Wang, Guoqing, et al.
Published: (2026)
by: Wang, Guoqing, et al.
Published: (2026)
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
by: Hansen, Nicklas, et al.
Published: (2024)
by: Hansen, Nicklas, et al.
Published: (2024)
The Entropy Enigma: Success and Failure of Entropy Minimization
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Similar Items
-
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
by: Balestriero, Randall, et al.
Published: (2025) -
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
by: Terver, Basile, et al.
Published: (2026) -
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024) -
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
by: Garrido, Quentin, et al.
Published: (2025) -
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)