Analysis of Spatial augmentation in Self-supervised models in the purview of training and test distributions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jha, Abhishek, Tuytelaars, Tinne
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909327581773824
author Jha, Abhishek
Tuytelaars, Tinne
author_facet Jha, Abhishek
Tuytelaars, Tinne
contents In this paper, we present an empirical study of typical spatial augmentation techniques used in self-supervised representation learning methods (both contrastive and non-contrastive), namely random crop and cutout. Our contributions are: (a) we dissociate random cropping into two separate augmentations, overlap and patch, and provide a detailed analysis on the effect of area of overlap and patch size to the accuracy on down stream tasks. (b) We offer an insight into why cutout augmentation does not learn good representation, as reported in earlier literature. Finally, based on these analysis, (c) we propose a distance-based margin to the invariance loss for learning scene-centric representations for the downstream task on object-centric distribution, showing that as simple as a margin proportional to the pixel distance between the two spatial views in the scence-centric images can improve the learned representation. Our study furthers the understanding of the spatial augmentations, and the effect of the domain-gap between the training augmentations and the test distribution.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18228
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Analysis of Spatial augmentation in Self-supervised models in the purview of training and test distributions
Jha, Abhishek
Tuytelaars, Tinne
Computer Vision and Pattern Recognition
In this paper, we present an empirical study of typical spatial augmentation techniques used in self-supervised representation learning methods (both contrastive and non-contrastive), namely random crop and cutout. Our contributions are: (a) we dissociate random cropping into two separate augmentations, overlap and patch, and provide a detailed analysis on the effect of area of overlap and patch size to the accuracy on down stream tasks. (b) We offer an insight into why cutout augmentation does not learn good representation, as reported in earlier literature. Finally, based on these analysis, (c) we propose a distance-based margin to the invariance loss for learning scene-centric representations for the downstream task on object-centric distribution, showing that as simple as a margin proportional to the pixel distance between the two spatial views in the scence-centric images can improve the learned representation. Our study furthers the understanding of the spatial augmentations, and the effect of the domain-gap between the training augmentations and the test distribution.
title Analysis of Spatial augmentation in Self-supervised models in the purview of training and test distributions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.18228