Distilling Vision Transformers for Distortion-Robust Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alexis, Konstantinos, Giannopoulos, Giorgos, Gunopulos, Dimitrios
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915955267862528
author Alexis, Konstantinos
Giannopoulos, Giorgos
Gunopulos, Dimitrios
author_facet Alexis, Konstantinos
Giannopoulos, Giorgos
Gunopulos, Dimitrios
contents Self-supervised learning has achieved remarkable success in learning visual representations from clean data, yet remains challenging when clean observations are sparse or not available at all. In this paper, we demonstrate that pretrained vision models can be leveraged to learn distortion-robust representations, which can then be effectively applied to downstream tasks operating on distorted observations. In particular, we propose an asymmetric knowledge distillation framework in which both teacher and student are initialized from the same pretrained Vision Transformer but receive different views of each image: the teacher processes clean images, while the student sees their distorted versions. We introduce multi-level distillation that aligns global embeddings, patch-level features, and attention maps and show that the student is able to approximate clean-image representations despite never directly accessing clean data. We evaluate our approach on image classification tasks across several datasets and under various distortions, consistently outperforming existing alternatives for the same amount of human supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22529
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Distilling Vision Transformers for Distortion-Robust Representation Learning
Alexis, Konstantinos
Giannopoulos, Giorgos
Gunopulos, Dimitrios
Computer Vision and Pattern Recognition
Self-supervised learning has achieved remarkable success in learning visual representations from clean data, yet remains challenging when clean observations are sparse or not available at all. In this paper, we demonstrate that pretrained vision models can be leveraged to learn distortion-robust representations, which can then be effectively applied to downstream tasks operating on distorted observations. In particular, we propose an asymmetric knowledge distillation framework in which both teacher and student are initialized from the same pretrained Vision Transformer but receive different views of each image: the teacher processes clean images, while the student sees their distorted versions. We introduce multi-level distillation that aligns global embeddings, patch-level features, and attention maps and show that the student is able to approximate clean-image representations despite never directly accessing clean data. We evaluate our approach on image classification tasks across several datasets and under various distortions, consistently outperforming existing alternatives for the same amount of human supervision.
title Distilling Vision Transformers for Distortion-Robust Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.22529