Segformer++: Efficient Token-Merging Strategies for High-Resolution Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kienzle, Daniel, Kantonis, Marco, Schön, Robin, Lienhart, Rainer
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911885014597632
author Kienzle, Daniel
Kantonis, Marco
Schön, Robin
Lienhart, Rainer
author_facet Kienzle, Daniel
Kantonis, Marco
Schön, Robin
Lienhart, Rainer
contents Utilizing transformer architectures for semantic segmentation of high-resolution images is hindered by the attention's quadratic computational complexity in the number of tokens. A solution to this challenge involves decreasing the number of tokens through token merging, which has exhibited remarkable enhancements in inference speed, training efficiency, and memory utilization for image classification tasks. In this paper, we explore various token merging strategies within the framework of the Segformer architecture and perform experiments on multiple semantic segmentation and human pose estimation datasets. Notably, without model re-training, we, for example, achieve an inference acceleration of 61% on the Cityscapes dataset while maintaining the mIoU performance. Consequently, this paper facilitates the deployment of transformer-based architectures on resource-constrained devices and in real-time applications.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14467
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Segformer++: Efficient Token-Merging Strategies for High-Resolution Semantic Segmentation
Kienzle, Daniel
Kantonis, Marco
Schön, Robin
Lienhart, Rainer
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Utilizing transformer architectures for semantic segmentation of high-resolution images is hindered by the attention's quadratic computational complexity in the number of tokens. A solution to this challenge involves decreasing the number of tokens through token merging, which has exhibited remarkable enhancements in inference speed, training efficiency, and memory utilization for image classification tasks. In this paper, we explore various token merging strategies within the framework of the Segformer architecture and perform experiments on multiple semantic segmentation and human pose estimation datasets. Notably, without model re-training, we, for example, achieve an inference acceleration of 61% on the Cityscapes dataset while maintaining the mIoU performance. Consequently, this paper facilitates the deployment of transformer-based architectures on resource-constrained devices and in real-time applications.
title Segformer++: Efficient Token-Merging Strategies for High-Resolution Semantic Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.14467