A Contrastive Learning Scheme with Transformer Innate Patches

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jyhne, Sander Riisøen, Andersen, Per-Arne, Goodwin, Morten
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929201604460544
author Jyhne, Sander Riisøen
Andersen, Per-Arne
Goodwin, Morten
author_facet Jyhne, Sander Riisøen
Andersen, Per-Arne
Goodwin, Morten
contents This paper presents Contrastive Transformer, a contrastive learning scheme using the Transformer innate patches. Contrastive Transformer enables existing contrastive learning techniques, often used for image classification, to benefit dense downstream prediction tasks such as semantic segmentation. The scheme performs supervised patch-level contrastive learning, selecting the patches based on the ground truth mask, subsequently used for hard-negative and hard-positive sampling. The scheme applies to all vision-transformer architectures, is easy to implement, and introduces minimal additional memory footprint. Additionally, the scheme removes the need for huge batch sizes, as each patch is treated as an image. We apply and test Contrastive Transformer for the case of aerial image segmentation, known for low-resolution data, large class imbalance, and similar semantic classes. We perform extensive experiments to show the efficacy of the Contrastive Transformer scheme on the ISPRS Potsdam aerial image segmentation dataset. Additionally, we show the generalizability of our scheme by applying it to multiple inherently different Transformer architectures. Ultimately, the results show a consistent increase in mean IoU across all classes.
format Preprint
id arxiv_https___arxiv_org_abs_2303_14806
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Contrastive Learning Scheme with Transformer Innate Patches
Jyhne, Sander Riisøen
Andersen, Per-Arne
Goodwin, Morten
Computer Vision and Pattern Recognition
Artificial Intelligence
This paper presents Contrastive Transformer, a contrastive learning scheme using the Transformer innate patches. Contrastive Transformer enables existing contrastive learning techniques, often used for image classification, to benefit dense downstream prediction tasks such as semantic segmentation. The scheme performs supervised patch-level contrastive learning, selecting the patches based on the ground truth mask, subsequently used for hard-negative and hard-positive sampling. The scheme applies to all vision-transformer architectures, is easy to implement, and introduces minimal additional memory footprint. Additionally, the scheme removes the need for huge batch sizes, as each patch is treated as an image. We apply and test Contrastive Transformer for the case of aerial image segmentation, known for low-resolution data, large class imbalance, and similar semantic classes. We perform extensive experiments to show the efficacy of the Contrastive Transformer scheme on the ISPRS Potsdam aerial image segmentation dataset. Additionally, we show the generalizability of our scheme by applying it to multiple inherently different Transformer architectures. Ultimately, the results show a consistent increase in mean IoU across all classes.
title A Contrastive Learning Scheme with Transformer Innate Patches
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2303.14806