Unsupervised Training of Vision Transformers with Synthetic Negatives

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Giakoumoglou, Nikolaos, Floros, Andreas, Papadopoulos, Kleanthis Marios, Stathaki, Tania
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916929983217664
author Giakoumoglou, Nikolaos
Floros, Andreas
Papadopoulos, Kleanthis Marios
Stathaki, Tania
author_facet Giakoumoglou, Nikolaos
Floros, Andreas
Papadopoulos, Kleanthis Marios
Stathaki, Tania
contents This paper does not introduce a novel method per se. Instead, we address the neglected potential of hard negative samples in self-supervised learning. Previous works explored synthetic hard negatives but rarely in the context of vision transformers. We build on this observation and integrate synthetic hard negatives to improve vision transformer representation learning. This simple yet effective technique notably improves the discriminative power of learned representations. Our experiments show performance improvements for both DeiT-S and Swin-T architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Training of Vision Transformers with Synthetic Negatives
Giakoumoglou, Nikolaos
Floros, Andreas
Papadopoulos, Kleanthis Marios
Stathaki, Tania
Computer Vision and Pattern Recognition
Artificial Intelligence
This paper does not introduce a novel method per se. Instead, we address the neglected potential of hard negative samples in self-supervised learning. Previous works explored synthetic hard negatives but rarely in the context of vision transformers. We build on this observation and integrate synthetic hard negatives to improve vision transformer representation learning. This simple yet effective technique notably improves the discriminative power of learned representations. Our experiments show performance improvements for both DeiT-S and Swin-T architectures.
title Unsupervised Training of Vision Transformers with Synthetic Negatives
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.02024