Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haftlang, Morteza Kiani, Malmir, Mohammadhossein, Parand, Foroutan, Michelucci, Umberto, Ghazouali, Safouane El
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915484375449600
author Haftlang, Morteza Kiani
Malmir, Mohammadhossein
Parand, Foroutan
Michelucci, Umberto
Ghazouali, Safouane El
author_facet Haftlang, Morteza Kiani
Malmir, Mohammadhossein
Parand, Foroutan
Michelucci, Umberto
Ghazouali, Safouane El
contents Medical image segmentation is a critical task in clinical workflows, particularly for the detection and delineation of pathological regions. While convolutional architectures like U-Net have become standard for such tasks, their limited receptive field restricts global context modeling. Recent efforts integrating transformers have addressed this, but often result in deep, computationally expensive models unsuitable for real-time use. In this work, we present a novel end-to-end lightweight architecture designed specifically for real-time binary medical image segmentation. Our model combines a Swin Transformer-like encoder with a U-Net-like decoder, connected via skip pathways to preserve spatial detail while capturing contextual information. Unlike existing designs such as Swin Transformer or U-Net, our architecture is significantly shallower and competitively efficient. To improve the encoder's ability to learn meaningful features without relying on large amounts of labeled data, we first train it using Barlow Twins, a self-supervised learning method that helps the model focus on important patterns by reducing unnecessary repetition in the learned features. After this pretraining, we fine-tune the entire model for our specific task. Experiments on benchmark binary segmentation tasks demonstrate that our model achieves competitive accuracy with substantially reduced parameter count and faster inference, positioning it as a practical alternative for deployment in real-time and resource-limited clinical environments. The code for our method is available at Github repository: https://github.com/mkianih/Barlow-Swin.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06885
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
Haftlang, Morteza Kiani
Malmir, Mohammadhossein
Parand, Foroutan
Michelucci, Umberto
Ghazouali, Safouane El
Computer Vision and Pattern Recognition
Artificial Intelligence
Medical image segmentation is a critical task in clinical workflows, particularly for the detection and delineation of pathological regions. While convolutional architectures like U-Net have become standard for such tasks, their limited receptive field restricts global context modeling. Recent efforts integrating transformers have addressed this, but often result in deep, computationally expensive models unsuitable for real-time use. In this work, we present a novel end-to-end lightweight architecture designed specifically for real-time binary medical image segmentation. Our model combines a Swin Transformer-like encoder with a U-Net-like decoder, connected via skip pathways to preserve spatial detail while capturing contextual information. Unlike existing designs such as Swin Transformer or U-Net, our architecture is significantly shallower and competitively efficient. To improve the encoder's ability to learn meaningful features without relying on large amounts of labeled data, we first train it using Barlow Twins, a self-supervised learning method that helps the model focus on important patterns by reducing unnecessary repetition in the learned features. After this pretraining, we fine-tune the entire model for our specific task. Experiments on benchmark binary segmentation tasks demonstrate that our model achieves competitive accuracy with substantially reduced parameter count and faster inference, positioning it as a practical alternative for deployment in real-time and resource-limited clinical environments. The code for our method is available at Github repository: https://github.com/mkianih/Barlow-Swin.
title Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.06885