Revisiting Feature Prediction for Learning Visual Representations from Video
Fuente:
arXiv
Salvato in:
| Autori principali: | Bardes, Adrien, Garrido, Quentin, Ponce, Jean, Chen, Xinlei, Rabbat, Michael, LeCun, Yann, Assran, Mahmoud, Ballas, Nicolas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning and Leveraging World Models in Visual Representation Learning
di: Garrido, Quentin, et al.
Pubblicazione: (2024)
di: Garrido, Quentin, et al.
Pubblicazione: (2024)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
di: Garrido, Quentin, et al.
Pubblicazione: (2025)
di: Garrido, Quentin, et al.
Pubblicazione: (2025)
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
di: Mur-Labadia, Lorenzo, et al.
Pubblicazione: (2026)
di: Mur-Labadia, Lorenzo, et al.
Pubblicazione: (2026)
Learning Latent Action World Models In The Wild
di: Garrido, Quentin, et al.
Pubblicazione: (2026)
di: Garrido, Quentin, et al.
Pubblicazione: (2026)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
di: Balestriero, Randall, et al.
Pubblicazione: (2025)
di: Balestriero, Randall, et al.
Pubblicazione: (2025)
Scaling Language-Free Visual Representation Learning
di: Fan, David, et al.
Pubblicazione: (2025)
di: Fan, David, et al.
Pubblicazione: (2025)
What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
di: Terver, Basile, et al.
Pubblicazione: (2025)
di: Terver, Basile, et al.
Pubblicazione: (2025)
Video Representation Learning with Joint-Embedding Predictive Architectures
di: Drozdov, Katrina, et al.
Pubblicazione: (2024)
di: Drozdov, Katrina, et al.
Pubblicazione: (2024)
Learning by Reconstruction Produces Uninformative Features For Perception
di: Balestriero, Randall, et al.
Pubblicazione: (2024)
di: Balestriero, Randall, et al.
Pubblicazione: (2024)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
di: Krojer, Benno, et al.
Pubblicazione: (2025)
di: Krojer, Benno, et al.
Pubblicazione: (2025)
Stochastic positional embeddings improve masked image modeling
di: Bar, Amir, et al.
Pubblicazione: (2023)
di: Bar, Amir, et al.
Pubblicazione: (2023)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
di: Lin, Han, et al.
Pubblicazione: (2024)
di: Lin, Han, et al.
Pubblicazione: (2024)
Hierarchical Planning with Latent World Models
di: Zhang, Wancong, et al.
Pubblicazione: (2026)
di: Zhang, Wancong, et al.
Pubblicazione: (2026)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
di: Balestriero, Randall, et al.
Pubblicazione: (2025)
di: Balestriero, Randall, et al.
Pubblicazione: (2025)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
di: Assran, Mido, et al.
Pubblicazione: (2025)
di: Assran, Mido, et al.
Pubblicazione: (2025)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
di: Yun, Zeyu, et al.
Pubblicazione: (2023)
di: Yun, Zeyu, et al.
Pubblicazione: (2023)
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
di: Terver, Basile, et al.
Pubblicazione: (2026)
di: Terver, Basile, et al.
Pubblicazione: (2026)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
di: Tong, Shengbang, et al.
Pubblicazione: (2024)
di: Tong, Shengbang, et al.
Pubblicazione: (2024)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
di: Goswami, Raktim Gautam, et al.
Pubblicazione: (2025)
di: Goswami, Raktim Gautam, et al.
Pubblicazione: (2025)
Transformers without Normalization
di: Zhu, Jiachen, et al.
Pubblicazione: (2025)
di: Zhu, Jiachen, et al.
Pubblicazione: (2025)
A hierarchical loss and its problems when classifying non-hierarchically
di: Wu, Cinna, et al.
Pubblicazione: (2017)
di: Wu, Cinna, et al.
Pubblicazione: (2017)
Variance-Covariance Regularization Improves Representation Learning
di: Zhu, Jiachen, et al.
Pubblicazione: (2023)
di: Zhu, Jiachen, et al.
Pubblicazione: (2023)
Parallel Stochastic Gradient-Based Planning for World Models
di: Psenka, Michael, et al.
Pubblicazione: (2026)
di: Psenka, Michael, et al.
Pubblicazione: (2026)
Whole-Body Conditioned Egocentric Video Prediction
di: Bai, Yutong, et al.
Pubblicazione: (2025)
di: Bai, Yutong, et al.
Pubblicazione: (2025)
Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
di: Kuang, Yilun, et al.
Pubblicazione: (2026)
di: Kuang, Yilun, et al.
Pubblicazione: (2026)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
di: Goswami, Raktim Gautam, et al.
Pubblicazione: (2024)
di: Goswami, Raktim Gautam, et al.
Pubblicazione: (2024)
Representation Learning for Spatiotemporal Physical Systems
di: Qu, Helen, et al.
Pubblicazione: (2026)
di: Qu, Helen, et al.
Pubblicazione: (2026)
Value-guided action planning with JEPA world models
di: Destrade, Matthieu, et al.
Pubblicazione: (2025)
di: Destrade, Matthieu, et al.
Pubblicazione: (2025)
Blockwise Self-Supervised Learning at Scale
di: Siddiqui, Shoaib Ahmed, et al.
Pubblicazione: (2023)
di: Siddiqui, Shoaib Ahmed, et al.
Pubblicazione: (2023)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
di: Tong, Shengbang, et al.
Pubblicazione: (2024)
di: Tong, Shengbang, et al.
Pubblicazione: (2024)
Back to the Features: DINO as a Foundation for Video World Models
di: Baldassarre, Federico, et al.
Pubblicazione: (2025)
di: Baldassarre, Federico, et al.
Pubblicazione: (2025)
The Entropy Enigma: Success and Failure of Entropy Minimization
di: Press, Ori, et al.
Pubblicazione: (2024)
di: Press, Ori, et al.
Pubblicazione: (2024)
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
di: Hansen, Nicklas, et al.
Pubblicazione: (2024)
di: Hansen, Nicklas, et al.
Pubblicazione: (2024)
Variance Covariance Regularization Enforces Pairwise Independence in Self-Supervised Representations
di: Mialon, Grégoire, et al.
Pubblicazione: (2022)
di: Mialon, Grégoire, et al.
Pubblicazione: (2022)
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
di: Zhou, Gaoyue, et al.
Pubblicazione: (2024)
di: Zhou, Gaoyue, et al.
Pubblicazione: (2024)
DINOv2: Learning Robust Visual Features without Supervision
di: Oquab, Maxime, et al.
Pubblicazione: (2023)
di: Oquab, Maxime, et al.
Pubblicazione: (2023)
Fast and Exact Enumeration of Deep Networks Partitions Regions
di: Balestriero, Randall, et al.
Pubblicazione: (2024)
di: Balestriero, Randall, et al.
Pubblicazione: (2024)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
di: Dawid, Anna, et al.
Pubblicazione: (2023)
di: Dawid, Anna, et al.
Pubblicazione: (2023)
Self-Supervised Learning with Lie Symmetries for Partial Differential Equations
di: Mialon, Grégoire, et al.
Pubblicazione: (2023)
di: Mialon, Grégoire, et al.
Pubblicazione: (2023)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
di: Lavoie, Samuel, et al.
Pubblicazione: (2024)
di: Lavoie, Samuel, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Learning and Leveraging World Models in Visual Representation Learning
di: Garrido, Quentin, et al.
Pubblicazione: (2024) -
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
di: Garrido, Quentin, et al.
Pubblicazione: (2025) -
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
di: Mur-Labadia, Lorenzo, et al.
Pubblicazione: (2026) -
Learning Latent Action World Models In The Wild
di: Garrido, Quentin, et al.
Pubblicazione: (2026) -
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
di: Balestriero, Randall, et al.
Pubblicazione: (2025)