WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Lianghui, Li, Yingyue, Fang, Jiemin, Liu, Yan, Xin, Hao, Liu, Wenyu, Wang, Xinggang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914419131285504
author Zhu, Lianghui
Li, Yingyue
Fang, Jiemin
Liu, Yan
Xin, Hao
Liu, Wenyu
Wang, Xinggang
author_facet Zhu, Lianghui
Li, Yingyue
Fang, Jiemin
Liu, Yan
Xin, Hao
Liu, Wenyu
Wang, Xinggang
contents Transformer has been very successful in various computer vision tasks and understanding the working mechanism of transformer is important. As touchstones, weakly-supervised semantic segmentation (WSSS) and class activation map (CAM) are useful tasks for analyzing vision transformers (ViT). Based on the plain ViT pre-trained with ImageNet classification, we find that multi-layer, multi-head self-attention maps can provide rich and diverse information for weakly-supervised semantic segmentation and CAM generation, e.g., different attention heads of ViT focus on different image areas and object categories. Thus we propose a novel method to end-to-end estimate the importance of attention heads, where the self-attention maps are adaptively fused for high-quality CAM results that tend to have more complete objects. Besides, we propose a ViT-based gradient clipping decoder for online retraining with the CAM results efficiently and effectively. Furthermore, the gradient clipping decoder can make good use of the knowledge in large-scale pre-trained ViT and has a scalable ability. The proposed plain Transformer-based Weakly-supervised learning method (WeakTr) obtains the superior WSSS performance on standard benchmarks, i.e., 78.5% mIoU on the val set of PASCAL VOC 2012 and 51.1% mIoU on the val set of COCO 2014. Source code and checkpoints are available at https://github.com/hustvl/WeakTr.
format Preprint
id arxiv_https___arxiv_org_abs_2304_01184
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
Zhu, Lianghui
Li, Yingyue
Fang, Jiemin
Liu, Yan
Xin, Hao
Liu, Wenyu
Wang, Xinggang
Computer Vision and Pattern Recognition
Transformer has been very successful in various computer vision tasks and understanding the working mechanism of transformer is important. As touchstones, weakly-supervised semantic segmentation (WSSS) and class activation map (CAM) are useful tasks for analyzing vision transformers (ViT). Based on the plain ViT pre-trained with ImageNet classification, we find that multi-layer, multi-head self-attention maps can provide rich and diverse information for weakly-supervised semantic segmentation and CAM generation, e.g., different attention heads of ViT focus on different image areas and object categories. Thus we propose a novel method to end-to-end estimate the importance of attention heads, where the self-attention maps are adaptively fused for high-quality CAM results that tend to have more complete objects. Besides, we propose a ViT-based gradient clipping decoder for online retraining with the CAM results efficiently and effectively. Furthermore, the gradient clipping decoder can make good use of the knowledge in large-scale pre-trained ViT and has a scalable ability. The proposed plain Transformer-based Weakly-supervised learning method (WeakTr) obtains the superior WSSS performance on standard benchmarks, i.e., 78.5% mIoU on the val set of PASCAL VOC 2012 and 51.1% mIoU on the val set of COCO 2014. Source code and checkpoints are available at https://github.com/hustvl/WeakTr.
title WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2304.01184