DINOv2: Learning Robust Visual Features without Supervision
Fuente:
arXiv
Guardado en:
| Autores principales: | Oquab, Maxime, Darcet, Timothée, Moutakanni, Théo, Vo, Huy, Szafraniec, Marc, Khalidov, Vasil, Fernandez, Pierre, Haziza, Daniel, Massa, Francisco, El-Nouby, Alaaeldin, Assran, Mahmoud, Ballas, Nicolas, Galuba, Wojciech, Howes, Russell, Huang, Po-Yao, Li, Shang-Wen, Misra, Ishan, Rabbat, Michael, Sharma, Vasu, Synnaeve, Gabriel, Xu, Hu, Jegou, Hervé, Mairal, Julien, Labatut, Patrick, Joulin, Armand, Bojanowski, Piotr |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach
por: Vo, Huy V., et al.
Publicado: (2024)
por: Vo, Huy V., et al.
Publicado: (2024)
DINOv3
por: Siméoni, Oriane, et al.
Publicado: (2025)
por: Siméoni, Oriane, et al.
Publicado: (2025)
Vision Transformers Need Registers
por: Darcet, Timothée, et al.
Publicado: (2023)
por: Darcet, Timothée, et al.
Publicado: (2023)
Cluster and Predict Latent Patches for Improved Masked Image Modeling
por: Darcet, Timothée, et al.
Publicado: (2025)
por: Darcet, Timothée, et al.
Publicado: (2025)
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
por: Jose, Cijo, et al.
Publicado: (2024)
por: Jose, Cijo, et al.
Publicado: (2024)
You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning
por: Moutakanni, Théo, et al.
Publicado: (2024)
por: Moutakanni, Théo, et al.
Publicado: (2024)
Back to the Features: DINO as a Foundation for Video World Models
por: Baldassarre, Federico, et al.
Publicado: (2025)
por: Baldassarre, Federico, et al.
Publicado: (2025)
Advancing human-centric AI for robust X-ray analysis through holistic self-supervised learning
por: Moutakanni, Théo, et al.
Publicado: (2024)
por: Moutakanni, Théo, et al.
Publicado: (2024)
Scalable Pre-training of Large Autoregressive Image Models
por: El-Nouby, Alaaeldin, et al.
Publicado: (2024)
por: El-Nouby, Alaaeldin, et al.
Publicado: (2024)
Efficient Universal Perception Encoder
por: Zhu, Chenchen, et al.
Publicado: (2026)
por: Zhu, Chenchen, et al.
Publicado: (2026)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
por: Assran, Mido, et al.
Publicado: (2025)
por: Assran, Mido, et al.
Publicado: (2025)
Revisiting Feature Prediction for Learning Visual Representations from Video
por: Bardes, Adrien, et al.
Publicado: (2024)
por: Bardes, Adrien, et al.
Publicado: (2024)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
por: Garrido, Quentin, et al.
Publicado: (2025)
por: Garrido, Quentin, et al.
Publicado: (2025)
CHMv2: Improvements in Global Canopy Height Mapping using DINOv3
por: Brandt, John, et al.
Publicado: (2026)
por: Brandt, John, et al.
Publicado: (2026)
Disentangling the Factors of Convergence between Brains and Computer Vision Models
por: Raugel, Joséphine, et al.
Publicado: (2025)
por: Raugel, Joséphine, et al.
Publicado: (2025)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
por: Balestriero, Randall, et al.
Publicado: (2025)
por: Balestriero, Randall, et al.
Publicado: (2025)
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
por: Mur-Labadia, Lorenzo, et al.
Publicado: (2026)
por: Mur-Labadia, Lorenzo, et al.
Publicado: (2026)
Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images
por: Raugel, Joséphine, et al.
Publicado: (2026)
por: Raugel, Joséphine, et al.
Publicado: (2026)
Scaling Laws for Optimal Data Mixtures
por: Shukor, Mustafa, et al.
Publicado: (2025)
por: Shukor, Mustafa, et al.
Publicado: (2025)
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
por: Haziza, Daniel, et al.
Publicado: (2025)
por: Haziza, Daniel, et al.
Publicado: (2025)
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
por: Marouani, Alexis, et al.
Publicado: (2026)
por: Marouani, Alexis, et al.
Publicado: (2026)
Learning and Leveraging World Models in Visual Representation Learning
por: Garrido, Quentin, et al.
Publicado: (2024)
por: Garrido, Quentin, et al.
Publicado: (2024)
Scaling Laws for Native Multimodal Models
por: Shukor, Mustafa, et al.
Publicado: (2025)
por: Shukor, Mustafa, et al.
Publicado: (2025)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
por: Krojer, Benno, et al.
Publicado: (2025)
por: Krojer, Benno, et al.
Publicado: (2025)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
por: Lin, Han, et al.
Publicado: (2024)
por: Lin, Han, et al.
Publicado: (2024)
Learning Latent Action World Models In The Wild
por: Garrido, Quentin, et al.
Publicado: (2026)
por: Garrido, Quentin, et al.
Publicado: (2026)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
por: Lavoie, Samuel, et al.
Publicado: (2024)
por: Lavoie, Samuel, et al.
Publicado: (2024)
Machine learning methods for finite population parameter estimation in survey sampling
por: Dagdoug, Mehdi, et al.
Publicado: (2026)
por: Dagdoug, Mehdi, et al.
Publicado: (2026)
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
por: Bachmann, Roman, et al.
Publicado: (2025)
por: Bachmann, Roman, et al.
Publicado: (2025)
Convergence of a double step scheme for a class of second order Clarke subdifferential inclusions
por: Bartosz, Krzysztof, et al.
Publicado: (2023)
por: Bartosz, Krzysztof, et al.
Publicado: (2023)
A note on the spectral gap for log-concave probability measures on convex bodies
por: Bonnefont, Michel, et al.
Publicado: (2023)
por: Bonnefont, Michel, et al.
Publicado: (2023)
Variable Selection for Linear Regression Imputation in Surveys
por: An, Ziming, et al.
Publicado: (2026)
por: An, Ziming, et al.
Publicado: (2026)
Stochastic positional embeddings improve masked image modeling
por: Bar, Amir, et al.
Publicado: (2023)
por: Bar, Amir, et al.
Publicado: (2023)
Shruteek/Optimized-sgRNA-Design: v1.0.0
por: Shruteek Mairal
Publicado: (2025)
por: Shruteek Mairal
Publicado: (2025)
Antropología de las ciudades históricas
por: Gaspar Mairal
Publicado: (2001)
por: Gaspar Mairal
Publicado: (2001)
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
por: Suau, Xavier, et al.
Publicado: (2024)
por: Suau, Xavier, et al.
Publicado: (2024)
Set Block Decoding is a Language Model Inference Accelerator
por: Gat, Itai, et al.
Publicado: (2025)
por: Gat, Itai, et al.
Publicado: (2025)
Statistical inference in the presence of imputed survey data through regression trees and random forests
por: Mehdi Dagdoug, et al.
Publicado: (2025)
por: Mehdi Dagdoug, et al.
Publicado: (2025)
On high‐dimensional variance estimation in survey sampling
por: Esther Eustache, et al.
Publicado: (2025)
por: Esther Eustache, et al.
Publicado: (2025)
Gradient-Guided Annealing for Domain Generalization
por: Ballas, Aristotelis, et al.
Publicado: (2025)
por: Ballas, Aristotelis, et al.
Publicado: (2025)
Ejemplares similares
-
Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach
por: Vo, Huy V., et al.
Publicado: (2024) -
DINOv3
por: Siméoni, Oriane, et al.
Publicado: (2025) -
Vision Transformers Need Registers
por: Darcet, Timothée, et al.
Publicado: (2023) -
Cluster and Predict Latent Patches for Improved Masked Image Modeling
por: Darcet, Timothée, et al.
Publicado: (2025) -
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
por: Jose, Cijo, et al.
Publicado: (2024)