TCFormer: Visual Recognition via Token Clustering Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Wang, Jin, Sheng, Xu, Lumin, Liu, Wentao, Qian, Chen, Ouyang, Wanli, Luo, Ping, Wang, Xiaogang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916325962547200
author Zeng, Wang
Jin, Sheng
Xu, Lumin
Liu, Wentao
Qian, Chen
Ouyang, Wanli
Luo, Ping
Wang, Xiaogang
author_facet Zeng, Wang
Jin, Sheng
Xu, Lumin
Liu, Wentao
Qian, Chen
Ouyang, Wanli
Luo, Ping
Wang, Xiaogang
contents Transformers are widely used in computer vision areas and have achieved remarkable success. Most state-of-the-art approaches split images into regular grids and represent each grid region with a vision token. However, fixed token distribution disregards the semantic meaning of different image regions, resulting in sub-optimal performance. To address this issue, we propose the Token Clustering Transformer (TCFormer), which generates dynamic vision tokens based on semantic meaning. Our dynamic tokens possess two crucial characteristics: (1) Representing image regions with similar semantic meanings using the same vision token, even if those regions are not adjacent, and (2) concentrating on regions with valuable details and represent them using fine tokens. Through extensive experimentation across various applications, including image classification, human pose estimation, semantic segmentation, and object detection, we demonstrate the effectiveness of our TCFormer. The code and models for this work are available at https://github.com/zengwang430521/TCFormer.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TCFormer: Visual Recognition via Token Clustering Transformer
Zeng, Wang
Jin, Sheng
Xu, Lumin
Liu, Wentao
Qian, Chen
Ouyang, Wanli
Luo, Ping
Wang, Xiaogang
Computer Vision and Pattern Recognition
Transformers are widely used in computer vision areas and have achieved remarkable success. Most state-of-the-art approaches split images into regular grids and represent each grid region with a vision token. However, fixed token distribution disregards the semantic meaning of different image regions, resulting in sub-optimal performance. To address this issue, we propose the Token Clustering Transformer (TCFormer), which generates dynamic vision tokens based on semantic meaning. Our dynamic tokens possess two crucial characteristics: (1) Representing image regions with similar semantic meanings using the same vision token, even if those regions are not adjacent, and (2) concentrating on regions with valuable details and represent them using fine tokens. Through extensive experimentation across various applications, including image classification, human pose estimation, semantic segmentation, and object detection, we demonstrate the effectiveness of our TCFormer. The code and models for this work are available at https://github.com/zengwang430521/TCFormer.
title TCFormer: Visual Recognition via Token Clustering Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.11321