When Swin Transformer Meets KANs: An Improved Transformer Architecture for Medical Image Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sapkota, Nishchal, Shi, Haoyan, Zhang, Yejia, Ma, Xianshi, Zheng, Bofang, Vazquez, Fabian, Gu, Pengfei, Chen, Danny Z.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918304569884672
author Sapkota, Nishchal
Shi, Haoyan
Zhang, Yejia
Ma, Xianshi
Zheng, Bofang
Vazquez, Fabian
Gu, Pengfei
Chen, Danny Z.
author_facet Sapkota, Nishchal
Shi, Haoyan
Zhang, Yejia
Ma, Xianshi
Zheng, Bofang
Vazquez, Fabian
Gu, Pengfei
Chen, Danny Z.
contents Medical image segmentation is critical for accurate diagnostics and treatment planning, but remains challenging due to complex anatomical structures and limited annotated training data. CNN-based segmentation methods excel at local feature extraction, but struggle with modeling long-range dependencies. Transformers, on the other hand, capture global context more effectively, but are inherently data-hungry and computationally expensive. In this work, we introduce UKAST, a U-Net like architecture that integrates rational-function based Kolmogorov-Arnold Networks (KANs) into Swin Transformer encoders. By leveraging rational base functions and Group Rational KANs (GR-KANs) from the Kolmogorov-Arnold Transformer (KAT), our architecture addresses the inefficiencies of vanilla spline-based KANs, yielding a more expressive and data-efficient framework with reduced FLOPs and only a very small increase in parameter count compared to SwinUNETR. UKAST achieves state-of-the-art performance on four diverse 2D and 3D medical image segmentation benchmarks, consistently surpassing both CNN- and Transformer-based baselines. Notably, it attains superior accuracy in data-scarce settings, alleviating the data-hungry limitations of standard Vision Transformers. These results show the potential of KAN-enhanced Transformers to advance data-efficient medical image segmentation. Code is available at: https://github.com/nsapkota417/UKAST
format Preprint
id arxiv_https___arxiv_org_abs_2511_04084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Swin Transformer Meets KANs: An Improved Transformer Architecture for Medical Image Segmentation
Sapkota, Nishchal
Shi, Haoyan
Zhang, Yejia
Ma, Xianshi
Zheng, Bofang
Vazquez, Fabian
Gu, Pengfei
Chen, Danny Z.
Computer Vision and Pattern Recognition
Medical image segmentation is critical for accurate diagnostics and treatment planning, but remains challenging due to complex anatomical structures and limited annotated training data. CNN-based segmentation methods excel at local feature extraction, but struggle with modeling long-range dependencies. Transformers, on the other hand, capture global context more effectively, but are inherently data-hungry and computationally expensive. In this work, we introduce UKAST, a U-Net like architecture that integrates rational-function based Kolmogorov-Arnold Networks (KANs) into Swin Transformer encoders. By leveraging rational base functions and Group Rational KANs (GR-KANs) from the Kolmogorov-Arnold Transformer (KAT), our architecture addresses the inefficiencies of vanilla spline-based KANs, yielding a more expressive and data-efficient framework with reduced FLOPs and only a very small increase in parameter count compared to SwinUNETR. UKAST achieves state-of-the-art performance on four diverse 2D and 3D medical image segmentation benchmarks, consistently surpassing both CNN- and Transformer-based baselines. Notably, it attains superior accuracy in data-scarce settings, alleviating the data-hungry limitations of standard Vision Transformers. These results show the potential of KAN-enhanced Transformers to advance data-efficient medical image segmentation. Code is available at: https://github.com/nsapkota417/UKAST
title When Swin Transformer Meets KANs: An Improved Transformer Architecture for Medical Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.04084