Saved in:
| Main Authors: | Joseph, Sonia, Garrido, Quentin, Balestriero, Randall, Kowal, Matthew, Fel, Thomas, Bakhtiari, Shahab, Richards, Blake, Rabbat, Mike |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.07050 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
by: Terver, Basile, et al.
Published: (2026)
by: Terver, Basile, et al.
Published: (2026)
seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
by: Ghaemi, Hafez, et al.
Published: (2025)
by: Ghaemi, Hafez, et al.
Published: (2025)
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026)
by: Garrido, Quentin, et al.
Published: (2026)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
by: Thasarathan, Harrish, et al.
Published: (2025)
by: Thasarathan, Harrish, et al.
Published: (2025)
UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling
by: Al-Tahan, Haider, et al.
Published: (2024)
by: Al-Tahan, Haider, et al.
Published: (2024)
Sparks of Explainability: Recent Advancements in Explaining Large Vision Models
by: Fel, Thomas
Published: (2025)
by: Fel, Thomas
Published: (2025)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)
by: Bardes, Adrien, et al.
Published: (2024)
Self-Supervised Anomaly Detection in the Wild: Favor Joint Embeddings Methods
by: Otero, Daniel, et al.
Published: (2024)
by: Otero, Daniel, et al.
Published: (2024)
Deep Networks Always Grok and Here is Why
by: Humayun, Ahmed Imtiaz, et al.
Published: (2024)
by: Humayun, Ahmed Imtiaz, et al.
Published: (2024)
On the Geometry of Deep Learning
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
by: Bordes, Florian, et al.
Published: (2025)
by: Bordes, Florian, et al.
Published: (2025)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
by: Garrido, Quentin, et al.
Published: (2025)
by: Garrido, Quentin, et al.
Published: (2025)
Post-Hoc Guidance for Consistency Models by Joint Flow Distribution Learning
by: Hsu, Chia-Hong, et al.
Published: (2026)
by: Hsu, Chia-Hong, et al.
Published: (2026)
GPS-SSL: Guided Positive Sampling to Inject Prior Into Self-Supervised Learning
by: Feizi, Aarash, et al.
Published: (2024)
by: Feizi, Aarash, et al.
Published: (2024)
Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data Curriculum
by: Lu, Wenquan, et al.
Published: (2025)
by: Lu, Wenquan, et al.
Published: (2025)
Self-supervised Video Instance Segmentation Can Boost Geographic Entity Alignment in Historical Maps
by: Xia, Xue, et al.
Published: (2024)
by: Xia, Xue, et al.
Published: (2024)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training Speed
by: Zhang, Jiaqi, et al.
Published: (2025)
by: Zhang, Jiaqi, et al.
Published: (2025)
Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
by: Doshi, Fenil R., et al.
Published: (2025)
by: Doshi, Fenil R., et al.
Published: (2025)
VISReg: Variance-Invariance-Sketching Regularization for JEPA training
by: Wu, Haiyu, et al.
Published: (2026)
by: Wu, Haiyu, et al.
Published: (2026)
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
by: Du, Hongyang, et al.
Published: (2026)
by: Du, Hongyang, et al.
Published: (2026)
Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
by: Van Assel, Hugues, et al.
Published: (2025)
by: Van Assel, Hugues, et al.
Published: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024)
by: Garrido, Quentin, et al.
Published: (2024)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Bi-Orthogonal Factor Decomposition for Vision Transformers
by: Doshi, Fenil R., et al.
Published: (2026)
by: Doshi, Fenil R., et al.
Published: (2026)
A Geometric Unification of Concept Learning with Concept Cones
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
Understanding Visual Feature Reliance through the Lens of Complexity
by: Fel, Thomas, et al.
Published: (2024)
by: Fel, Thomas, et al.
Published: (2024)
Visual Concept Connectome (VCC): Open World Concept Discovery and their Interlayer Connections in Deep Models
by: Kowal, Matthew, et al.
Published: (2024)
by: Kowal, Matthew, et al.
Published: (2024)
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
Beyond and Free from Diffusion: Invertible Guided Consistency Training
by: Hsu, Chia-Hong, et al.
Published: (2025)
by: Hsu, Chia-Hong, et al.
Published: (2025)
From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
by: Bisulco, Anthony, et al.
Published: (2025)
by: Bisulco, Anthony, et al.
Published: (2025)
Understanding Video Transformers via Universal Concept Discovery
by: Kowal, Matthew, et al.
Published: (2024)
by: Kowal, Matthew, et al.
Published: (2024)
WorldModelBench: Judging Video Generation Models As World Models
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
Similar Items
-
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
by: Balestriero, Randall, et al.
Published: (2025) -
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
by: Fel, Thomas, et al.
Published: (2025) -
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
by: Terver, Basile, et al.
Published: (2026) -
seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
by: Ghaemi, Hafez, et al.
Published: (2025) -
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026)