Unsupervised Training of Vision Transformers with Synthetic Negatives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916929983217664 |
|---|---|
| author | Giakoumoglou, Nikolaos Floros, Andreas Papadopoulos, Kleanthis Marios Stathaki, Tania |
| author_facet | Giakoumoglou, Nikolaos Floros, Andreas Papadopoulos, Kleanthis Marios Stathaki, Tania |
| contents | This paper does not introduce a novel method per se. Instead, we address the neglected potential of hard negative samples in self-supervised learning. Previous works explored synthetic hard negatives but rarely in the context of vision transformers. We build on this observation and integrate synthetic hard negatives to improve vision transformer representation learning. This simple yet effective technique notably improves the discriminative power of learned representations. Our experiments show performance improvements for both DeiT-S and Swin-T architectures. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_02024 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Unsupervised Training of Vision Transformers with Synthetic Negatives Giakoumoglou, Nikolaos Floros, Andreas Papadopoulos, Kleanthis Marios Stathaki, Tania Computer Vision and Pattern Recognition Artificial Intelligence This paper does not introduce a novel method per se. Instead, we address the neglected potential of hard negative samples in self-supervised learning. Previous works explored synthetic hard negatives but rarely in the context of vision transformers. We build on this observation and integrate synthetic hard negatives to improve vision transformer representation learning. This simple yet effective technique notably improves the discriminative power of learned representations. Our experiments show performance improvements for both DeiT-S and Swin-T architectures. |
| title | Unsupervised Training of Vision Transformers with Synthetic Negatives |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2509.02024 |