Low-Resolution Self-Attention for Semantic Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908408141053952 |
|---|---|
| author | Wu, Yu-Huan Zhang, Shi-Chen Liu, Yun Zhang, Le Zhan, Xin Zhou, Daquan Feng, Jiashi Cheng, Ming-Ming Zhen, Liangli |
| author_facet | Wu, Yu-Huan Zhang, Shi-Chen Liu, Yun Zhang, Le Zhan, Xin Zhou, Daquan Feng, Jiashi Cheng, Ming-Ming Zhen, Liangli |
| contents | Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction. While existing vision transformers demonstrate promising performance, they often utilize high-resolution context modeling, resulting in a computational bottleneck. In this work, we challenge conventional wisdom and introduce the Low-Resolution Self-Attention (LRSA) mechanism to capture global context at a significantly reduced computational cost, i.e., FLOPs. Our approach involves computing self-attention in a fixed low-resolution space regardless of the input image's resolution, with additional 3x3 depth-wise convolutions to capture fine details in the high-resolution space. We demonstrate the effectiveness of our LRSA approach by building the LRFormer, a vision transformer with an encoder-decoder structure. Extensive experiments on the ADE20K, COCO-Stuff, and Cityscapes datasets demonstrate that LRFormer outperforms state-of-the-art models. Code is available at https://github.com/yuhuan-wu/LRFormer. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_05026 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Low-Resolution Self-Attention for Semantic Segmentation Wu, Yu-Huan Zhang, Shi-Chen Liu, Yun Zhang, Le Zhan, Xin Zhou, Daquan Feng, Jiashi Cheng, Ming-Ming Zhen, Liangli Computer Vision and Pattern Recognition Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction. While existing vision transformers demonstrate promising performance, they often utilize high-resolution context modeling, resulting in a computational bottleneck. In this work, we challenge conventional wisdom and introduce the Low-Resolution Self-Attention (LRSA) mechanism to capture global context at a significantly reduced computational cost, i.e., FLOPs. Our approach involves computing self-attention in a fixed low-resolution space regardless of the input image's resolution, with additional 3x3 depth-wise convolutions to capture fine details in the high-resolution space. We demonstrate the effectiveness of our LRSA approach by building the LRFormer, a vision transformer with an encoder-decoder structure. Extensive experiments on the ADE20K, COCO-Stuff, and Cityscapes datasets demonstrate that LRFormer outperforms state-of-the-art models. Code is available at https://github.com/yuhuan-wu/LRFormer. |
| title | Low-Resolution Self-Attention for Semantic Segmentation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2310.05026 |