V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Assran, Mido, Bardes, Adrien, Fan, David, Garrido, Quentin, Howes, Russell, Mojtaba, Komeili, Muckley, Matthew, Rizvi, Ammar, Roberts, Claire, Sinha, Koustuv, Zholus, Artem, Arnaud, Sergio, Gejji, Abha, Martin, Ada, Hogan, Francois Robert, Dugas, Daniel, Bojanowski, Piotr, Khalidov, Vasil, Labatut, Patrick, Massa, Francisco, Szafraniec, Marc, Krishnakumar, Kapil, Li, Yong, Ma, Xiaodong, Chandar, Sarath, Meier, Franziska, LeCun, Yann, Rabbat, Michael, Ballas, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
by: Mur-Labadia, Lorenzo, et al.
Published: (2026)
by: Mur-Labadia, Lorenzo, et al.
Published: (2026)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
Back to the Features: DINO as a Foundation for Video World Models
by: Baldassarre, Federico, et al.
Published: (2025)
by: Baldassarre, Federico, et al.
Published: (2025)
Hierarchical Planning with Latent World Models
by: Zhang, Wancong, et al.
Published: (2026)
by: Zhang, Wancong, et al.
Published: (2026)
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)
by: Bardes, Adrien, et al.
Published: (2024)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
by: Garrido, Quentin, et al.
Published: (2025)
by: Garrido, Quentin, et al.
Published: (2025)
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024)
by: Garrido, Quentin, et al.
Published: (2024)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
by: Krojer, Benno, et al.
Published: (2025)
by: Krojer, Benno, et al.
Published: (2025)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
by: Nilaksh, et al.
Published: (2026)
by: Nilaksh, et al.
Published: (2026)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Mastering Memory Tasks with World Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
DINOv2: Learning Robust Visual Features without Supervision
by: Oquab, Maxime, et al.
Published: (2023)
by: Oquab, Maxime, et al.
Published: (2023)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
by: Arnaud, Sergio, et al.
Published: (2025)
by: Arnaud, Sergio, et al.
Published: (2025)
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026)
by: Garrido, Quentin, et al.
Published: (2026)
Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach
by: Vo, Huy V., et al.
Published: (2024)
by: Vo, Huy V., et al.
Published: (2024)
Scaling Language-Free Visual Representation Learning
by: Fan, David, et al.
Published: (2025)
by: Fan, David, et al.
Published: (2025)
BindGPT: A Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning
by: Zholus, Artem, et al.
Published: (2024)
by: Zholus, Artem, et al.
Published: (2024)
Disentangling the Factors of Convergence between Brains and Computer Vision Models
by: Raugel, Joséphine, et al.
Published: (2025)
by: Raugel, Joséphine, et al.
Published: (2025)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA
by: Huang, Hai, et al.
Published: (2026)
by: Huang, Hai, et al.
Published: (2026)
You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning
by: Moutakanni, Théo, et al.
Published: (2024)
by: Moutakanni, Théo, et al.
Published: (2024)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images
by: Raugel, Joséphine, et al.
Published: (2026)
by: Raugel, Joséphine, et al.
Published: (2026)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
TAPNext++: What's Next for Tracking Any Point (TAP)?
by: Jung, Sebastian, et al.
Published: (2026)
by: Jung, Sebastian, et al.
Published: (2026)
Value-guided action planning with JEPA world models
by: Destrade, Matthieu, et al.
Published: (2025)
by: Destrade, Matthieu, et al.
Published: (2025)
Efficient Universal Perception Encoder
by: Zhu, Chenchen, et al.
Published: (2026)
by: Zhu, Chenchen, et al.
Published: (2026)
Advancing human-centric AI for robust X-ray analysis through holistic self-supervised learning
by: Moutakanni, Théo, et al.
Published: (2024)
by: Moutakanni, Théo, et al.
Published: (2024)
What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
by: Terver, Basile, et al.
Published: (2025)
by: Terver, Basile, et al.
Published: (2025)
Pluralistic psychotherapists' and counsellors' experiences of working with actively suicidal clients: A qualitative interpretative phenomenological analysis
by: Leo Muckley
Published: (2024)
by: Leo Muckley
Published: (2024)
Video Technology: Conveying Information Visually.
by: Bardes, D'Ellen
Published: (1985)
by: Bardes, D'Ellen
Published: (1985)
Attention Novices: Friendly Intro to Shiny Disks.
by: Bardes, D'Ellen
Published: (1986)
by: Bardes, D'Ellen
Published: (1986)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
Causal-JEPA: Learning World Models through Object-Level Latent Masking
by: Nam, Heejeong, et al.
Published: (2026)
by: Nam, Heejeong, et al.
Published: (2026)
Parallel Stochastic Gradient-Based Planning for World Models
by: Psenka, Michael, et al.
Published: (2026)
by: Psenka, Michael, et al.
Published: (2026)
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
by: Baldelli, Davide, et al.
Published: (2026)
by: Baldelli, Davide, et al.
Published: (2026)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
by: Guiroy, Simon, et al.
Published: (2025)
by: Guiroy, Simon, et al.
Published: (2025)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
Similar Items
-
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
by: Mur-Labadia, Lorenzo, et al.
Published: (2026) -
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
by: Lin, Han, et al.
Published: (2024) -
Back to the Features: DINO as a Foundation for Video World Models
by: Baldassarre, Federico, et al.
Published: (2025) -
Hierarchical Planning with Latent World Models
by: Zhang, Wancong, et al.
Published: (2026) -
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)