Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Scardecchia, Mattia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914073043533824
author Scardecchia, Mattia
author_facet Scardecchia, Mattia
contents Recent advances in self-supervised learning (SSL) have made it possible to learn general-purpose visual features that capture both the high-level semantics and the fine-grained spatial structure of images. Most notably, the recent DINOv2 has established a new state of the art by surpassing weakly supervised methods (WSL) like OpenCLIP on most benchmarks. In this survey, we examine the core ideas behind its approach, multi-crop view augmentation and self-distillation with a mean teacher, and trace their development in previous work. We then compare the performance of DINO and DINOv2 with other SSL and WSL methods across various downstream tasks, and highlight some remarkable emergent properties of their learned features with transformer backbones. We conclude by briefly discussing DINOv2's limitations, its impact, and future research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
Scardecchia, Mattia
Computer Vision and Pattern Recognition
Machine Learning
Image and Video Processing
Recent advances in self-supervised learning (SSL) have made it possible to learn general-purpose visual features that capture both the high-level semantics and the fine-grained spatial structure of images. Most notably, the recent DINOv2 has established a new state of the art by surpassing weakly supervised methods (WSL) like OpenCLIP on most benchmarks. In this survey, we examine the core ideas behind its approach, multi-crop view augmentation and self-distillation with a mean teacher, and trace their development in previous work. We then compare the performance of DINO and DINOv2 with other SSL and WSL methods across various downstream tasks, and highlight some remarkable emergent properties of their learned features with transformer backbones. We conclude by briefly discussing DINOv2's limitations, its impact, and future research directions.
title Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
topic Computer Vision and Pattern Recognition
Machine Learning
Image and Video Processing
url https://arxiv.org/abs/2510.03606