GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Haonan, Liu, Jie, Tang, Jie, Wu, Gangshan, Xu, Bo, Chou, Yanbing, Wang, Yong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910528402620416
author Wang, Haonan
Liu, Jie
Tang, Jie
Wu, Gangshan
Xu, Bo
Chou, Yanbing
Wang, Yong
author_facet Wang, Haonan
Liu, Jie
Tang, Jie
Wu, Gangshan
Xu, Bo
Chou, Yanbing
Wang, Yong
contents In recent years, 2D human pose estimation has made significant progress on public benchmarks. However, many of these approaches face challenges of less applicability in the industrial community due to the large number of parametric quantities and computational overhead. Efficient human pose estimation remains a hurdle, especially for whole-body pose estimation with numerous keypoints. While most current methods for efficient human pose estimation primarily rely on CNNs, we propose the Group-based Token Pruning Transformer (GTPT) that fully harnesses the advantages of the Transformer. GTPT alleviates the computational burden by gradually introducing keypoints in a coarse-to-fine manner. It minimizes the computation overhead while ensuring high performance. Besides, GTPT groups keypoint tokens and prunes visual tokens to improve model performance while reducing redundancy. We propose the Multi-Head Group Attention (MHGA) between different groups to achieve global interaction with little computational overhead. We conducted experiments on COCO and COCO-WholeBody. Compared to other methods, the experimental results show that GTPT can achieve higher performance with less computation, especially in whole-body with numerous keypoints.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10756
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation
Wang, Haonan
Liu, Jie
Tang, Jie
Wu, Gangshan
Xu, Bo
Chou, Yanbing
Wang, Yong
Computer Vision and Pattern Recognition
In recent years, 2D human pose estimation has made significant progress on public benchmarks. However, many of these approaches face challenges of less applicability in the industrial community due to the large number of parametric quantities and computational overhead. Efficient human pose estimation remains a hurdle, especially for whole-body pose estimation with numerous keypoints. While most current methods for efficient human pose estimation primarily rely on CNNs, we propose the Group-based Token Pruning Transformer (GTPT) that fully harnesses the advantages of the Transformer. GTPT alleviates the computational burden by gradually introducing keypoints in a coarse-to-fine manner. It minimizes the computation overhead while ensuring high performance. Besides, GTPT groups keypoint tokens and prunes visual tokens to improve model performance while reducing redundancy. We propose the Multi-Head Group Attention (MHGA) between different groups to achieve global interaction with little computational overhead. We conducted experiments on COCO and COCO-WholeBody. Compared to other methods, the experimental results show that GTPT can achieve higher performance with less computation, especially in whole-body with numerous keypoints.
title GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.10756