Low-Resolution Self-Attention for Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yu-Huan, Zhang, Shi-Chen, Liu, Yun, Zhang, Le, Zhan, Xin, Zhou, Daquan, Feng, Jiashi, Cheng, Ming-Ming, Zhen, Liangli
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908408141053952
author Wu, Yu-Huan
Zhang, Shi-Chen
Liu, Yun
Zhang, Le
Zhan, Xin
Zhou, Daquan
Feng, Jiashi
Cheng, Ming-Ming
Zhen, Liangli
author_facet Wu, Yu-Huan
Zhang, Shi-Chen
Liu, Yun
Zhang, Le
Zhan, Xin
Zhou, Daquan
Feng, Jiashi
Cheng, Ming-Ming
Zhen, Liangli
contents Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction. While existing vision transformers demonstrate promising performance, they often utilize high-resolution context modeling, resulting in a computational bottleneck. In this work, we challenge conventional wisdom and introduce the Low-Resolution Self-Attention (LRSA) mechanism to capture global context at a significantly reduced computational cost, i.e., FLOPs. Our approach involves computing self-attention in a fixed low-resolution space regardless of the input image's resolution, with additional 3x3 depth-wise convolutions to capture fine details in the high-resolution space. We demonstrate the effectiveness of our LRSA approach by building the LRFormer, a vision transformer with an encoder-decoder structure. Extensive experiments on the ADE20K, COCO-Stuff, and Cityscapes datasets demonstrate that LRFormer outperforms state-of-the-art models. Code is available at https://github.com/yuhuan-wu/LRFormer.
format Preprint
id arxiv_https___arxiv_org_abs_2310_05026
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Low-Resolution Self-Attention for Semantic Segmentation
Wu, Yu-Huan
Zhang, Shi-Chen
Liu, Yun
Zhang, Le
Zhan, Xin
Zhou, Daquan
Feng, Jiashi
Cheng, Ming-Ming
Zhen, Liangli
Computer Vision and Pattern Recognition
Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction. While existing vision transformers demonstrate promising performance, they often utilize high-resolution context modeling, resulting in a computational bottleneck. In this work, we challenge conventional wisdom and introduce the Low-Resolution Self-Attention (LRSA) mechanism to capture global context at a significantly reduced computational cost, i.e., FLOPs. Our approach involves computing self-attention in a fixed low-resolution space regardless of the input image's resolution, with additional 3x3 depth-wise convolutions to capture fine details in the high-resolution space. We demonstrate the effectiveness of our LRSA approach by building the LRFormer, a vision transformer with an encoder-decoder structure. Extensive experiments on the ADE20K, COCO-Stuff, and Cityscapes datasets demonstrate that LRFormer outperforms state-of-the-art models. Code is available at https://github.com/yuhuan-wu/LRFormer.
title Low-Resolution Self-Attention for Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.05026